Hongchi Xia

dblp:355/1029 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 79% Representation and self-supervised learning · 18% Autonomous driving · 3%
Computer graphics and multimedia
2 papers
Virtual and augmented reality · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › virtual environment
interactive virtual environments
1.622025
DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025
Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video · CVPR 2024
Computer vision › 3D vision
3d scene reconstruction
0.912025
DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction
0.912025
DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025
Computer vision › 3D vision › object modeling
3d object learning
0.812024
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024
Computer vision › 3D vision
neural radiance field
0.812024
Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video · CVPR 2024
Computer vision › 3D vision › geometric deep learning
3d representation learning
0.712023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.712023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023
Computer vision › 3D vision
point cloud processing
0.712023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023
Computer vision › 3D vision › object pose estimation
6d object pose estimation
0.212024
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024
Computer vision › 3D vision
camera pose estimation
0.212024
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024
Computer vision › 3D vision
novel view synthesis
0.212024
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024
Computer vision › 3D vision › range sensing
LiDAR
0.212023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023
Robotics › Autonomous driving
perception
0.212023
Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023

Methods — techniques the papers use, named apart from their topics

real-to-sim transfer · 1.7dual scene representation · 1.7articulation estimation · 1.7physics simulation · 1.5neural radiance field · 1.5mesh distillation · 1.5point cloud reconstruction · 0.8RGB-D video capture · 0.8multimodal learning · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2025 AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
abstract
Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility.
Hao-Yu Hsu, Chih-Hao Lin, Albert J. Zhai, Hongchi Xia, Shenlong Wang
3DV4
2025 DRAWER: Digital Reconstruction and Articulation With Environment Realism
abstract
Creating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach centers on two main contributions: (i) a reconstruction module based on a dual scene representation that reconstructs the scene with fine-grained geometric details, and (ii) an articulation module that identifies articulation types and hinge positions, reconstructs simulatable shapes and appearances and integrates them into the scene. The resulting virtual environment is photorealistic, interactive, and runs in real time, with compatibility for game engines and robotic simulation platforms. We demonstrate the potential of DRAWER by using it to automatically create an interactive game in Unreal Engine and to enable real-to-sim-to-real transfer for robotics applications. Project page: here.
Hongchi Xia, Entong Su, Marius Memmel, Arhan Jain, Raymond Yu, Numfor Mbiziwo-Tiapo, Ali Farhadi, Abhishek Gupta 0004, Shenlong Wang, Wei-Chiu Ma
CVPR1
2025 HoloScene: Simulation-Ready Interactive 3D Worlds from a Single Video
abstract
Digitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more critical aspects, such as geometry completeness, object interactivity, physical plausibility, photorealistic rendering, or realistic physical properties for reliable dynamic simulation. To address these limitations, we introduce HoloScene, a novel interactive 3D reconstruction framework that simultaneously achieves these requirements. HoloScene leverages a comprehensive interactive scene-graph representation, encoding object geometry, appearance, and physical properties alongside hierarchical and inter-object relationships. Reconstruction is formulated as an energy-based optimization problem, integrating observational data, physical constraints, and generative priors into a unified, coherent objective. Optimization is efficiently performed via a hybrid approach combining sampling-based exploration with gradient-based refinement. The resulting digital twins exhibit complete and precise geometry, physical stability, and realistic rendering from novel viewpoints. Evaluations conducted on multiple benchmark datasets demonstrate superior performance, while practical use-cases in interactive gaming and real-time digital-twin manipulation illustrate HoloScene's broad applicability and effectiveness.
Hongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet, Katelyn Gao, Michael Paulitsch, Benjamin Ummenhofer, Shenlong Wang
NeurIPS1
2024 RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
abstract
We introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing, the direct capture of the depth channel allows better 3D annotations and broader downstream applications. WildRGB-D comprises large-scale category-level RGB-D object videos, which are taken using an iPhone to go around the objects in 360 degrees. It contains around 8500 recorded objects and nearly 20000 RGB-D videos across 46 common object categories. These videos are taken with diverse cluttered backgrounds with three setups to cover as many real-world scenarios as possible: (i) a single object in one video; (ii) multiple objects in one video; and (iii) an object with a static hand in one video. The dataset is annotated with object masks, real-world scale camera poses, and reconstructed aggregated point clouds from RGBD videos. We benchmark four tasks with WildRGB-D including novel view synthesis, camera pose estimation, object 6d pose estimation, and object surface reconstruction. Our experiments show that the large-scale capture of RGB-D objects provides a large potential to advance 3D object learning. Our project page is https://wildrgbd.github.io/.
Hongchi Xia, Sifei Liu, Xiaolong Wang 0004
CVPR1
2024 Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video
abstract
Creating high-quality and interactive virtual environments, such as games and simulators, often involves complex and costly manual modeling processes. In this paper, we present Video2Game, a novel approach that automatically converts videos of real-world scenes into realistic and interactive game environments. At the heart of our system are three core components: (i) a neural radiance fields (NeRF) module that effectively captures the geometry and visual appearance of the scene; (ii) a mesh module that distills the knowledge from NeRF for faster rendering; and (iii) a physics module that models the interactions and physical dynamics among the objects. By following the carefully designed pipeline, one can construct an interactable and actionable digital replica of the real world. We benchmark our system on both indoor and large-scale outdoor scenes. We show that we can not only produce highly-realistic renderings in real-time, but also build interactive games on top.
Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong Wang
CVPR1
2023 Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving
abstract
Due to the difficulty of annotating the 3D LiDAR data of autonomous driving, an efficient unsupervised 3D representation learning method is important. In this paper, we design the Triangle Constrained Contrast (TriCC) framework tailored for autonomous driving scenes which learns 3D unsupervised representations through both the multimodal information and dynamic of temporal sequences. We treat one camera image and two LiDAR point clouds with different timestamps as a triplet. And our key design is the consistent constraint that automatically finds matching relationships among the triplet through “self-cycle” and learns representations from it. With the matching relations across the temporal dimension and modalities, we can further conduct a triplet contrast to improve learning efficiency. To the best of our knowledge, TriCC is the first framework that unifies both the temporal and multimodal semantics, which means it utilizes almost all the information in autonomous driving scenes. And compared with previous contrastive methods, it can automatically dig out contrasting pairs with higher difficulty, instead of relying on handcrafted ones. Extensive experiments are conducted with Minkowski-UNet and VoxelNet on several semantic segmentation and 3D detection datasets. Results show that TriCC learns effective representations with much fewer training iterations and improves the SOTA results greatly on all the downstream tasks. Code and models can be found at https://bopang1996.github.io/.
Bo Pang 0003, Hongchi Xia, Cewu Lu
CVPR2