Albert J. Zhai

dblp:292/3136 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-0647-1730ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 72% Robot navigation and mapping · 17% Trustworthy machine learning · 8%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.522024
On the Overconfidence Problem in Semantic 3D Mapping · ICRA 2024
Physical Property Understanding from Language-Embedded Feature Fields · CVPR 2024
Computer vision › 3D vision
camera pose estimation
0.912025
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video · CVPR 2025
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction
0.912025
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video · CVPR 2025
Computer vision › 3D vision › 3d scene understanding
3d semantic mapping
0.812024
On the Overconfidence Problem in Semantic 3D Mapping · ICRA 2024
Computer vision › 3D vision › inverse rendering
material property estimation
0.812024
Physical Property Understanding from Language-Embedded Feature Fields · CVPR 2024
Robotics › Robot navigation and mapping
object search
0.812024
On the Overconfidence Problem in Semantic 3D Mapping · ICRA 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration
0.812024
On the Overconfidence Problem in Semantic 3D Mapping · ICRA 2024
Robotics › Robot navigation and mapping
object goal navigation
0.712023
PEANUT: Predicting and Navigating to Unseen Targets · ICCV 2023
Computer vision › 3D vision › 3d scene understanding
dynamic scene understanding
0.312025
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video · CVPR 2025
Robotics › Robot navigation and mapping
semantic map
0.212023
PEANUT: Predicting and Navigating to Unseen Targets · ICCV 2023

Methods — techniques the papers use, named apart from their topics

vision foundation model · 0.9multi-stage optimization · 0.9zero-shot kernel regression · 0.8neural radiance field · 0.8learned calibration · 0.8large language model · 0.8bayesian fusion · 0.8supervised learning · 0.7object location prediction · 0.7
YearPublicationVenuePosition
2026 CropCraft: Complete Structural Characterization of Crop Plants from Images
abstract
The ability to automatically build 3D digital twins of plants from images has countless applications in agriculture, environmental science, robotics, and other fields. However, current 3D reconstruction methods fail to recover complete shapes of plants due to heavy occlusion and complex geometries. In this work, we present a novel method for 3D modeling of agricultural crops based on optimizing a parametric model of plant morphology via inverse procedural modeling. Our method first estimates depth maps by fitting a neural radiance field and then optimizes a specialized loss to estimate morphological parameters that result in consistent depth renderings. The resulting 3D model is complete and biologically plausible. We validate our method on a dataset of real images of agricultural fields, and demonstrate that the reconstructed canopies can be used for a variety of monitoring and simulation applications. Project page: https://ajzhai.github.io/CropCraft
Albert J. Zhai, Zhao Jiang, Sheng Wang 0020, Zhenong Jin, Kaiyu Guan, Shenlong Wang
3DV1
2025 AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
abstract
Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility.
Hao-Yu Hsu, Chih-Hao Lin, Albert J. Zhai, Hongchi Xia, Shenlong Wang
3DV3
2025 Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
abstract
This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehensive 4D understanding remains challenging. We introduce Uni4D, a multi-stage optimization framework that harnesses multiple pretrained models to advance dynamic 3D modeling, including static/dynamic reconstruction, camera pose estimation, and dense 3D motion tracking. Our results show state-of-the-art performance in dynamic 4D modeling with superior visual quality. Notably, Uni4D requires no retraining or fine- tuning, highlighting the effectiveness of repurposing visual foundation models for 4D understanding. Code and more results are available at: https://davidyao99.github.io/uni4d.
David Yifan Yao, Albert J. Zhai, Shenlong Wang
CVPR2
2024 Physical Property Understanding from Language-Embedded Feature Fields
abstract
Can computers perceive the physical properties of objects solely through vision? Research in cognitive science and vision science has shown that humans excel at identifying materials and estimating their physical properties based purely on visual appearance. In this paper, we present a novel approach for dense prediction of the physical properties of objects using a collection of images. Inspired by how humans reason about physics through vision, we leverage large language models to propose candidate materials for each object. We then construct a language-embedded point cloud and estimate the physical properties of each 3D point using a zero-shot kernel regression approach. Our method is accurate, annotation-free, and applicable to any object in the open world. Experiments demonstrate the effectiveness of the proposed approach in various physical property reasoning tasks, such as estimating the mass of common objects, as well as other properties like friction and hardness. Code is available at https://ajzhai.github.io/NeRF2Physics.
Albert J. Zhai, Emily Y. Chen, Gloria X. Wang, Sheng Wang 0020, Kaiyu Guan, Shenlong Wang
CVPR1
2024 On the Overconfidence Problem in Semantic 3D Mapping
abstract
Semantic 3D mapping, the process of fusing depth and image segmentation information between multiple views to build 3D maps annotated with object classes in real-time, is a recent topic of interest. This paper highlights the fusion overconfidence problem, in which conventional mapping methods assign high confidence to the entire map even when they are incorrect, leading to miscalibrated outputs. Several methods to improve uncertainty calibration at different stages in the fusion pipeline are presented and compared on the ScanNet dataset. We show that the most widely used Bayesian fusion strategy is among the worst calibrated, and propose a learned pipeline that combines fusion and calibration, GLFS, which achieves simultaneously higher accuracy and 3D map calibration while retaining real-time capability and adding only 525 learned parameters to the pipeline. We further illustrate the importance of map calibration on a downstream task by showing that incorporating proper semantic fusion to an indoor object search agent improves its success rates.
João Marcos Correia Marques, Albert J. Zhai, Shenlong Wang, Kris Hauser
ICRA2
2023 PEANUT: Predicting and Navigating to Unseen Targets
abstract
Efficient ObjectGoal navigation (ObjectNav) in novel environments requires an understanding of the spatial and semantic regularities in environment layouts. In this work, we present a straightforward method for learning these regularities by predicting the locations of unobserved objects from incomplete semantic maps. Our method differs from previous prediction-based navigation methods, such as frontier potential prediction or egocentric map completion, by directly predicting unseen targets while leveraging the global context from all previously explored areas. Our prediction model is lightweight and can be trained in a supervised manner using a relatively small amount of passively collected data. Once trained, the model can be incorporated into a modular pipeline for ObjectNav without the need for any reinforcement learning. We validate the effectiveness of our method on the HM3D and MP3D ObjectNav datasets. We find that it achieves the state-of-the-art on both datasets, despite not using any additional data for training. Code is available at https://ajzhai.github.io/PEANUT.
Albert J. Zhai, Shenlong Wang
ICCV1