EDBT 2026 Demo / reviewers in the wild / expert
Huachen Gao
dblp:353/7350
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-5981-8566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | You See it, You Got it: Learning 3D Creation on Pose-Free Videos at ScaleabstractRecent 3D generation models typically rely on limited-scale 3D ‘gold-labels’ or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning paradigms. In this work, we present See3D, a visual-conditional multi-view diffusion model trained on large-scale Internet videos for open-world 3D creation. The model aims to Get 3D knowledge by solely Seeing the visual contents from the vast and rapidly growing video data — You See it, You Got it. To achieve this, we first scale up the training data using a proposed data curation pipeline that automatically filters out multi-view inconsistencies and insufficient observations from source videos. This results in a high-quality, richly diverse, large-scale dataset of multi-view images, termed WebVi3D, containing 320M frames from 16M video clips. Nevertheless, learning generic 3D priors from videos without explicit 3D geometry or camera pose annotations is nontrivial, and annotating poses for web-scale videos is prohibitively expensive. To eliminate the need for pose conditions, we introduce an innovative visual-condition - a purely 2D-inductive visual signal generated by adding time-dependent noise to the masked video data. Finally, we introduce a novel visual-conditional 3D generation framework by integrating See3D into a warping-based pipeline for high-fidelity 3D generation. Our numerical and visual comparisons on single and sparse reconstruction benchmarks show that See3D, trained on cost- effective and scalable video data, achieves notable zero-shot and open-world generation capabilities, markedly outperforming models trained on costly and constrained 3D datasets. Additionally, our model naturally supports other image-conditioned 3D creation tasks, such as 3D editing, without further fine-tuning. Please refer to our project page at: https://vision.baai.ac.cn/see3d. Baorui Ma, Huachen Gao, Haoge Deng, Tiejun Huang 0001, Lulu Tang |
CVPR | 2 |
| 2025 | MVD-HuGaS: Human Gaussians from a Single Image via 3D Human Multi-View Diffusion Prior
Kaiqiang Xiong, Jianbo Jiao, Huachen Gao, Ronggang Wang |
PRCV (10) | 7 |
| 2024 | Surface-Centric Modeling for High-Fidelity Generalizable Neural Surface Reconstruction
Shihe Shen, Kaiqiang Xiong, Huachen Gao, Jianbo Jiao, Ronggang Wang |
ECCV (32) | 4 |
| 2024 | Disentangled Generation and Aggregation for Robust Radiance Fields
Shihe Shen, Huachen Gao, Wangze Xu, Luyang Tang, Kaiqiang Xiong, Jianbo Jiao, Ronggang Wang |
ECCV (49) | 2 |
| 2024 | MVPGS: Excavating Multi-view Priors for Gaussian Splatting from Sparse Input Views
Wangze Xu, Huachen Gao, Shihe Shen, Jianbo Jiao, Ronggang Wang |
ECCV (47) | 2 |
| 2024 | FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth ConsistencyabstractLearning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving large camera movements. In this paper, we propose FDCNeRF that leverages the direction information contained in the RGB-based optical flow and depth-based virtual flow as a direct guidance for camera pose optimization to reduce pose-geometry ambiguity. Additionally, we introduce Adaptive Pose-Aware Sampling (APAS) to replace the previous random ray sampling strategy, which reduces the difficulty of pose learning in early stages and preserves the diversity of rays in later stages. Experiments on the challenging Tanks and Temples dataset demonstrate that our method achieves state-of-the-art results in both novel view synthesis quality and pose estimation accuracy. Huachen Gao, Shihe Shen, Kaiqiang Xiong, Zhirui Gao, Yugui Xie, Ronggang Wang |
ICASSP | 1 |
| 2024 | Large Point-to-Gaussian Model for Image-to-3D Generation
Longfei Lu, Huachen Gao, Tao Dai 0001, Yaohua Zha, Zhi Hou, Junta Wu, Shutao Xia |
ACM Multimedia | 2 |
| 2023 | N2MVSNet: Non-Local Neighbors Aware Multi-View Stereo NetworkabstractLearning-based multi-view stereo (MVS) methods have been widely studied recently. However, current works are limited to using fixed-size convolution kernels, leading to suboptimal features that lack anisotropy in low-textured regions and tend to produce invalid depth blending at the edge of the foreground and background. In this paper, we propose N2MVSNet, which learns adaptive non-local neighbors matching (ANNM) and their spatial impact to overcome these deficiencies. Furthermore, we explore the ability of spatial perception to depth dimension and propose 3D ANNM. Besides, following the coarse-to-fine scheme, severe mismatchings in coarse stages will result in error accumulation and propagation in finer stages. To this end, we adopt the pretrained RGB guided depth refinement for depth hypothesis repolish. The robustness of the training process is further elevated by the energy aggregation loss. Extensive experiments on the DTU and Tanks and Temples datasets demonstrate that the proposed network achieves state-of-the-art results. Huachen Gao, Ronggang Wang |
ICASSP | 2 |
| 2023 | Bi-ClueMVSNet: Learning Bidirectional Occlusion Clues for Multi-View StereoabstractDeep learning-based multi-view stereo (MVS) meth-ods have achieved promising results in recent years. However, very few existing works take the occlusion issues into consideration, leading to poor reconstruction results on the boundaries and occluded areas. In this paper, the Bidirectional Occlusion Clues-based Multi-View Stereo Network (Bi-ClueMVSNet) is proposed as an end-to-end MVS framework that explicitly models the occlusion obstacle for depth map inference and 3D modeling. To this end, we use bidirectional projection for the first time to reduce the propagation and accumulation of incorrect matches and build the occlusion-enhanced network to further advance the representational ability from 2D visibility maps to 3D occlusion clues. As for depth map estimation, we combine the characteristics of both regression and classification approaches to propose the adaptive depth map inference strategy. Besides, the robustness of the training process is further guaranteed and elevated by the occlusion clues-based loss function. The proposed method significantly improves the accuracy of depth map inference in boundaries and heavily occluded areas and brings the overall quality of the reconstructed point cloud to a new altitude. Extensive experiments are performed on DTU, Tanks and Temples, and BlendedMVS datasets to demonstrate the persuasiveness of the proposed framework. Huachen Gao, Ronggang Wang |
IJCNN | 3 |