EDBT 2026 Demo / reviewers in the wild / expert
Haoxiang Chen 0004
dblp:307/5448 · also Hao-Xiang Chen 0004
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0003-4357-4226ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors
Sicong Du, Jiarun Liu, Haoxiang Chen 0004, Tai-Jiang Mu, Sheng Yang 0007 |
ICCV | 4 |
| 2025 | SDLKF: Signed Distance Linear Kernel Function for surface reconstruction
Haoxiang Chen 0004, Xiao-Lei Li, Tai-Jiang Mu, Qun-Ce Xu, Shi-Min Hu 0001 |
Comput. Graph. | 1 |
| 2025 | RS-SpecSDF: Reflection-supervised surface reconstruction and material estimation for specular indoor scenesabstractNeural Radiance Field (NeRF) has achieved impressive 3D reconstruction quality using implicit scene representations. However, planar specular reflections pose significant challenges in the 3D reconstruction task. It is a common practice to decompose the scene into physically real geometries and virtual images produced by the reflections. However, current methods struggle to resolve the ambiguities in the decomposition process , because they mostly rely on mirror masks as external cues. They also fail to acquire accurate surface materials, which is essential for downstream applications of the recovered geometries. In this paper, we present RS-SpecSDF, a novel framework for indoor scene surface reconstruction that can faithfully reconstruct specular reflectors while accurately decomposing the reflection from the scene geometries and recovering the accurate specular fraction and diffuse appearance of the surface without requiring mirror masks. Our key idea is to perform reflection ray-casting and use it as supervision for the decomposition of reflection and surface material. Our method is based on an observation that the virtual image seen by the camera ray should be consistent with the object that the ray hits after reflecting off the specular surface. To leverage this constraint, we propose the Reflection Consistency Loss and Reflection Certainty Loss to regularize the decomposition. Experiments conducted on both our newly-proposed synthetic dataset and a real-captured dataset demonstrate that our method achieves high-quality surface reconstruction and accurate material decomposition results without the need of mirror masks. Dong-Yu Chen, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu |
Graph. Model. | 2 |
| 2025 | SN$^{2}$2eRF: A Framework for Neural Radiance Fields Given Sparse and Noisy PosesabstractNeural Radiance Fields (NeRFs) have shown impressive capabilities in synthesizing photorealistic novel views. However, their application to room-size scenes is limited by the requirement of several hundred views with accurate poses for training. To address this challenge, we propose SN$^{2}$2eRF, a framework which can reconstruct the neural radiance field with significantly fewer views and noisy poses by exploiting multiple priors. Our key insight is to leverage both multi-view and monocular priors to constrain the optimization of NeRF in the setting of sparse and noisy pose inputs. Specifically, we extract and match key points to constrain pose optimization and use Ray Transformer with a monocular depth estimator to provide dense depth prior for geometry optimization. Benefiting from these priors, our approach achieves state-of-the-art accuracy in novel view synthesis for indoor room scenarios. Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | SLS4D: Sparse Latent Space for 4D Novel View SynthesisabstractNeural radiance fields (NeRF) have achieved great success in novel view synthesis and 3D representation for static scenarios. Existing dynamic NeRFs usually exploit a locally dense grid to fit the deformation fields; however, they fail to capture the global dynamics and concomitantly yield models of heavy parameters. We observe that the 4D space is inherently sparse. First, the deformation fields are sparse in spatial but dense in temporal due to the continuity of motion. Second, the radiance fields are only valid on the surface of the underlying scene, usually occupying a small fraction of the whole space. We thus represent the 4D scene using a learnable sparse latent space, a.k.a. SLS4D. Specifically, SLS4D first uses dense learnable time slot features to depict the temporal space, from which the deformation fields are fitted with linear multi-layer perceptions (MLP) to predict the displacement of a 3D position at any time. It then learns the spatial features of a 3D position using another sparse latent space. This is achieved by learning the adaptive weights of each latent feature with the attention mechanism. Extensive experiments demonstrate the effectiveness of our SLS4D: It achieves the best 4D novel view synthesis using only about 6% parameters of the most recent work. Qi-Yuan Feng, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | EVSplitting: An Efficient and Visually Consistent Splitting Algorithm for 3D Gaussian SplattingabstractThis paper presents EVSplitting, an efficient and visually consistent splitting algorithm for 3D Gaussian Splatting (3DGS). It is designed to make operating 3DGS as easy and effective as other 3D explicit representations, readily for industrial productions. The challenges of above target are: 1) The huge number and complex attributes of 3DGS make it tough to explicitly operate on 3DGS in a real-time and learning-free manner; 2) The visual effect of 3DGS is very difficult to maintain during explicit operations and 3) The anisotropism of Gaussian always leads to blurs and artifacts. As far as we know, no prior work can address these challenges well. In this work, we introduce a direct and efficient 3DGS splitting algorithm to solve them. Specifically, we formulate the 3DGS splitting as two minimization problems that aim to ensure visual consistency and reduce Gaussian overflow across boundary (splitting plane), respectively. Firstly, we impose conservations on the zero-, first- and second-order moments of the weighted Gaussian distribution to guarantee visual consistency. Secondly, we reduce the boundary overflow with a special constraint on the aforementioned conservations. With these conservations and constraints, we derive a closed-form solution for the 3DGS splitting problem. This yields an easy-to-implement, plug-and-play, efficient and fundamental tool, benefiting various downstream applications of 3DGS. Qi-Yuan Feng, Geng-Chen Cao, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
SIGGRAPH Asia | 3 |
| 2024 | DIScene: Object Decoupling and Interaction Modeling for Complex Scene Generation
Xiao-Lei Li, Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
SIGGRAPH Asia | 3 |
| 2024 | FragmentDiff: A Diffusion Model for Fractured Object Assembly
Qun-Ce Xu, Haoxiang Chen 0004, Jiacheng Hua, Xiaohua Zhan, Yongliang Yang 0002, Tai-Jiang Mu |
SIGGRAPH Asia | 2 |
| 2023 | Neural 3D reconstruction from sparse views using geometric priorsabstractSparse view 3D reconstruction has attracted increasing attention with the development of neural implicit 3D representation. Existing methods usually only make use of 2D views, requiring a dense set of input views for accurate 3D reconstruction. In this paper, we show that accurate 3D reconstruction can be achieved by incorporating geometric priors into neural implicit 3D reconstruction. Our method adopts the signed distance function as the 3D representation, and learns a generalizable 3D surface reconstruction model from sparse views. Specifically, we build a more effective and sparse feature volume from the input views by using corresponding depth maps, which can be provided by depth sensors or directly predicted from the input views. We recover better geometric details by imposing both depth and surface normal constraints in addition to the color loss when training the neural implicit 3D representation. Experiments demonstrate that our method both outperforms state-of-the-art approaches, and achieves good generalizability. Tai-Jiang Mu, Haoxiang Chen 0004, Junxiong Cai |
Comput. Vis. Media | 2 |
| 2023 | SPS: Accurate and Real-Time Semantic Positioning System Based on Low-Cost DEM MapsabstractThis paper presents a Semantic Positioning System (SPS) to enhance the accuracy of mobile device geo-localization in outdoor urban environments. Although the traditional Global Positioning System (GPS) can offer a rough localization, it lacks the necessary accuracy for applications such as Augmented Reality (AR). Our SPS integrates Geographic Information System (GIS) data, GPS signals, and visual image information to estimate the 6 Degree-of-Freedom (DoF) pose through cross-view semantic matching. This approach has excellent scalability to support GIS context with Levels of Detail (LOD). The map data representation is Digital Elevation Model (DEM), a cost-effective aerial map that allows for fast deployment for large-scale areas. However, the DEM lacks geometric and texture details, making it challenging for traditional visual feature extraction to establish pixel/voxel level cross-view correspondences. To address this, we sample observation pixels from the query ground-view image using predicted semantic labels. We then propose an iterative homography estimation method with semantic correspondences. To improve the efficiency of the overall system, we further employ a heuristic search to speedup the matching process. The proposed method is robust, real-time, and automatic. Quantitative experiments on the challenging Bund dataset show that we achieve a positioning accuracy of 73.24%, surpassing the baseline skyline-based method by 20%. Compared with the state-of-the-art semantic-based approach on the Kitti dataset, we improve the positioning accuracy by an average of 5%. Junxiong Cai, Wensen Feng, Haoxiang Chen 0004, Tai-Jiang Mu |
IEEE Trans. Image Process. | 3 |
| 2023 | Real-Time Globally Consistent 3D Reconstruction With Semantic PriorsabstractMaintaining global consistency continues to be critical for online 3D indoor scene reconstruction. However, it is still challenging to generate satisfactory 3D reconstruction in terms of global consistency for previous approaches using purely geometric analysis, even with bundle adjustment or loop closure techniques. In this article, we propose a novel real-time 3D reconstruction approach which effectively integrates both semantic and geometric cues. The key challenge is how to map this indicative information, i.e., semantic priors, into a metric space as measurable information, thus enabling more accurate semantic fusion leveraging both the geometric and semantic cues. To this end, we introduce a semantic space with a continuous metric function measuring the distance between discrete semantic observations. Within the semantic space, we present an accurate frame-to-model semantic tracker for camera pose estimation, and semantic pose graph equipped with semantic links between submaps for globally consistent 3D scene reconstruction. With extensive evaluation on public synthetic and real-world 3D indoor scene RGB-D datasets, we show that our approach outperforms the previous approaches for 3D scene reconstruction both quantitatively and qualitatively, especially in terms of global consistency. Shi-Sheng Huang, Haoxiang Chen 0004, Hongbo Fu 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | CIRCLE: Convolutional Implicit Reconstruction and Completion for Large-Scale Indoor Scene
Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
ECCV (32) | 1 |
| 2022 | A Neural Galerkin Solver for Accurate Surface ReconstructionabstractTo reconstruct meshes from the widely-available 3D point cloud data, implicit shape representation is among the primary choices as an intermediate form due to its superior representation power and robustness in topological optimizations. Although different parameterizations of the implicit fields have been explored to model the underlying geometry, there is no explicit mechanism to ensure the fitting tightness of the surface to the input. We present in response, NeuralGalerkin, a neural Galerkin-method-based solver designed for reconstructing highly-accurate surfaces from the input point clouds. NeuralGalerkin internally discretizes the target implicit field as a linear combination of a set of spatially-varying basis functions inferred by an adaptive sparse convolution neural network. It then solves differentiably for a variational problem that incorporates both positional and normal constraints from the data in closed form within a single forward pass, highly respecting the raw input points. The reconstructed surface extracted from the implicit interpolants is hence very accurate and incorporates useful inductive biases benefiting from the training data. Extensive evaluations on various datasets demonstrate our method's promising reconstruction performance and scalability. Haoxiang Chen 0004, Shi-Min Hu 0001 |
ACM Trans. Graph. | 2 |