EDBT 2026 Demo / reviewers in the wild / expert
Hongchi Xia
dblp:355/1029
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 79% Representation and self-supervised learning · 18% Autonomous driving · 3% | |
| Computer graphics and multimedia
2 papers |
Virtual and augmented reality · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality › virtual environment
interactive virtual environments |
1.6 | 2 | 2025 | DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025 Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video · CVPR 2024 |
Computer vision › 3D vision
3d scene reconstruction |
0.9 | 1 | 2025 | DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025 |
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction |
0.9 | 1 | 2025 | DRAWER: Digital Reconstruction and Articulation With Environment Realism · CVPR 2025 |
Computer vision › 3D vision › object modeling
3d object learning |
0.8 | 1 | 2024 | RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024 |
Computer vision › 3D vision
neural radiance field |
0.8 | 1 | 2024 | Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video · CVPR 2024 |
Computer vision › 3D vision › geometric deep learning
3d representation learning |
0.7 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.7 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning |
0.7 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Computer vision › 3D vision
point cloud processing |
0.7 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Computer vision › 3D vision › object pose estimation
6d object pose estimation |
0.2 | 1 | 2024 | RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024 |
Computer vision › 3D vision
camera pose estimation |
0.2 | 1 | 2024 | RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024 |
Computer vision › 3D vision
novel view synthesis |
0.2 | 1 | 2024 | RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos · CVPR 2024 |
Computer vision › 3D vision › range sensing
LiDAR |
0.2 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Robotics › Autonomous driving
perception |
0.2 | 1 | 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous Driving · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
real-to-sim transfer · 1.7dual scene representation · 1.7articulation estimation · 1.7physics simulation · 1.5neural radiance field · 1.5mesh distillation · 1.5point cloud reconstruction · 0.8RGB-D video capture · 0.8multimodal learning · 0.7contrastive learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoVFX: Physically Realistic Video Editing from Natural Language InstructionsabstractModern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility. Hao-Yu Hsu, Chih-Hao Lin, Albert J. Zhai, Hongchi Xia, Shenlong Wang |
3DV | 4 |
| 2025 | DRAWER: Digital Reconstruction and Articulation With Environment RealismabstractCreating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach centers on two main contributions: (i) a reconstruction module based on a dual scene representation that reconstructs the scene with fine-grained geometric details, and (ii) an articulation module that identifies articulation types and hinge positions, reconstructs simulatable shapes and appearances and integrates them into the scene. The resulting virtual environment is photorealistic, interactive, and runs in real time, with compatibility for game engines and robotic simulation platforms. We demonstrate the potential of DRAWER by using it to automatically create an interactive game in Unreal Engine and to enable real-to-sim-to-real transfer for robotics applications. Project page: here. Hongchi Xia, Entong Su, Marius Memmel, Arhan Jain, Raymond Yu, Numfor Mbiziwo-Tiapo, Ali Farhadi, Abhishek Gupta 0004, Shenlong Wang, Wei-Chiu Ma |
CVPR | 1 |
| 2025 | HoloScene: Simulation-Ready Interactive 3D Worlds from a Single VideoabstractDigitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more critical aspects, such as geometry completeness, object interactivity, physical plausibility, photorealistic rendering, or realistic physical properties for reliable dynamic simulation. To address these limitations, we introduce HoloScene, a novel interactive 3D reconstruction framework that simultaneously achieves these requirements. HoloScene leverages a comprehensive interactive scene-graph representation, encoding object geometry, appearance, and physical properties alongside hierarchical and inter-object relationships. Reconstruction is formulated as an energy-based optimization problem, integrating observational data, physical constraints, and generative priors into a unified, coherent objective. Optimization is efficiently performed via a hybrid approach combining sampling-based exploration with gradient-based refinement. The resulting digital twins exhibit complete and precise geometry, physical stability, and realistic rendering from novel viewpoints. Evaluations conducted on multiple benchmark datasets demonstrate superior performance, while practical use-cases in interactive gaming and real-time digital-twin manipulation illustrate HoloScene's broad applicability and effectiveness. Hongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet, Katelyn Gao, Michael Paulitsch, Benjamin Ummenhofer, Shenlong Wang |
NeurIPS | 1 |
| 2024 | RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D VideosabstractWe introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing, the direct capture of the depth channel allows better 3D annotations and broader downstream applications. WildRGB-D comprises large-scale category-level RGB-D object videos, which are taken using an iPhone to go around the objects in 360 degrees. It contains around 8500 recorded objects and nearly 20000 RGB-D videos across 46 common object categories. These videos are taken with diverse cluttered backgrounds with three setups to cover as many real-world scenarios as possible: (i) a single object in one video; (ii) multiple objects in one video; and (iii) an object with a static hand in one video. The dataset is annotated with object masks, real-world scale camera poses, and reconstructed aggregated point clouds from RGBD videos. We benchmark four tasks with WildRGB-D including novel view synthesis, camera pose estimation, object 6d pose estimation, and object surface reconstruction. Our experiments show that the large-scale capture of RGB-D objects provides a large potential to advance 3D object learning. Our project page is https://wildrgbd.github.io/. Hongchi Xia, Sifei Liu, Xiaolong Wang 0004 |
CVPR | 1 |
| 2024 | Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single VideoabstractCreating high-quality and interactive virtual environments, such as games and simulators, often involves complex and costly manual modeling processes. In this paper, we present Video2Game, a novel approach that automatically converts videos of real-world scenes into realistic and interactive game environments. At the heart of our system are three core components: (i) a neural radiance fields (NeRF) module that effectively captures the geometry and visual appearance of the scene; (ii) a mesh module that distills the knowledge from NeRF for faster rendering; and (iii) a physics module that models the interactions and physical dynamics among the objects. By following the carefully designed pipeline, one can construct an interactable and actionable digital replica of the real world. We benchmark our system on both indoor and large-scale outdoor scenes. We show that we can not only produce highly-realistic renderings in real-time, but also build interactive games on top. Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong Wang |
CVPR | 1 |
| 2023 | Unsupervised 3D Point Cloud Representation Learning by Triangle Constrained Contrast for Autonomous DrivingabstractDue to the difficulty of annotating the 3D LiDAR data of autonomous driving, an efficient unsupervised 3D representation learning method is important. In this paper, we design the Triangle Constrained Contrast (TriCC) framework tailored for autonomous driving scenes which learns 3D unsupervised representations through both the multimodal information and dynamic of temporal sequences. We treat one camera image and two LiDAR point clouds with different timestamps as a triplet. And our key design is the consistent constraint that automatically finds matching relationships among the triplet through “self-cycle” and learns representations from it. With the matching relations across the temporal dimension and modalities, we can further conduct a triplet contrast to improve learning efficiency. To the best of our knowledge, TriCC is the first framework that unifies both the temporal and multimodal semantics, which means it utilizes almost all the information in autonomous driving scenes. And compared with previous contrastive methods, it can automatically dig out contrasting pairs with higher difficulty, instead of relying on handcrafted ones. Extensive experiments are conducted with Minkowski-UNet and VoxelNet on several semantic segmentation and 3D detection datasets. Results show that TriCC learns effective representations with much fewer training iterations and improves the SOTA results greatly on all the downstream tasks. Code and models can be found at https://bopang1996.github.io/. Bo Pang 0003, Hongchi Xia, Cewu Lu |
CVPR | 2 |