Haihong Xiao

dblp:157/1455 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-3543-9262ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhanced Geometry and Semantics for Camera-Based 3D Semantic Scene Completion
abstract
Giving machines the ability to infer the complete 3D geometry and semantics of complex scenes is crucial for many downstream tasks, such as decision-making and planning. Vision-centric Semantic Scene Completion (SSC) has emerged as a trendy 3D perception paradigm due to its compatibility with task properties, low cost, and rich visual cues. Despite impressive results, current approaches inevitably suffer from problems such as depth errors or depth ambiguities during the 2D-to-3D transformation process. To overcome these limitations, in this paper, we first introduce an Optical Flow-Guided (OFG) DepthNet that leverages the strengths of pretrained depth estimation models, while incorporating optical flow images to improve depth prediction accuracy in regions with significant depth changes. Then, we propose a depth ambiguity-mitigated feature lifting strategy that implements deformable cross-attention in 3D pixel space to avoid depth ambiguities caused by the projection process from 3D to 2D and further enhances the effectiveness of feature updating through the utilization of prior mask indices. Moreover, we customize two subnetworks: a residual voxel network and a sparse UNet, to enhance the network's geometric prediction capabilities and ensure consistent semantic reasoning across varying scales. By doing so, our method achieves performance improvements over state-of-the-art methods on the SemanticKITTI, SSCBench-KITTI-360 and Occ3D-nuScene benchmarks.
Haihong Xiao, Wenxiong Kang, Yulan Guo, Hao Liu 0061, Ying He 0001
IEEE Trans. Image Process.1
2026 Geometry-Aware 3D Gaussian Representation for Real-Time Rendering of Large-Scale Scenes
abstract
Existing NeRF-based methods for reconstructing large-scale scenes face challenges in visual quality and rendering speed due to spectral biases and extensive sampling requirements. Recent 3DGS-based methods for real-time rendering of 3D objects and small scenes outperform NeRF, but several issues persist when extending these techniques to large-scale scenes. These include robust rendering in weak-texture areas, effective densification under memory constraints, finer detail rendering, and achieving natural lighting transitions. To address these, we introduce a geometry-aware 3DGS method for efficient real-time rendering of large scenes. First, we propose a geometry-guided anchor point initialization that reduces noise away from structural surfaces and generates new points in weak texture areas, particularly for large-scale datasets. We also present a structure-aware joint densification strategy combining surface- and curvature-based densification, ensuring Gaussian points are near structural surfaces and increasing density in low-curvature areas. Additionally, we propose a hash grid-assisted, viewpoint-sensitive feature enhancement scheme to improve detail rendering and natural lighting transitions. Our method achieves superior rendering quality compared to state-of-the-art methods while maintaining reasonable memory usage. Extensive experiments across 16 scenes, including 11 from five public datasets (MatrixCity-Aerial, Mill-19, Tanks & Temples, WHU, and UrbanScene3D) and five self-collected scenes from SCUT-CA and plateau regions, demonstrate its generalization capability. Codes are available athttps://github.com/SCUT-BIP-Lab/Geo_gs.
Haihong Xiao, Jianan Zou, Shuai Xing, Wenxiong Kang
IEEE Trans. Multim.1
2025 Learning Multi-View Stereo With Geometry-Aware Prior
abstract
Multi-View Stereo (MVS) reconstructs detailed 3D structures from multi-view images by establishing spatial correspondences. While learning-based methods have significantly advanced the MVS task, challenges such as ambiguous matching caused by textureless surfaces and lighting variations persist. To address these issues, we propose GAP-MVSNet, a framework that leverages surface normals from a monocular normal foundation model as priors to enhance the geometric awareness of reconstruction targets. In this work, surface normal priors are seamlessly integrated into the MVS pipeline to improve depth prediction robustness and accuracy. Specifically, we introduce a structure-aware feature pyramid network that incorporates surface normal information and utilizes uncertainty-aware feature resampling to extract robust image features. Additionally, we present the spatial geometry enhanced regularization that combines sampled depth hypotheses with surface normals to generate a spatial geometric prior, guiding the cost regularization process and enforcing strong spatial coherence, particularly in textureless regions. Furthermore, we design a local consistency depth refinement module that utilizes surface normals to establish depth relationships as a local geometric prior, thereby refining classification-based depth predictions and aligning them with ground truth depth. Extensive experiments on the DTU and Tanks & Temples datasets demonstrate that our method achieves state-of-the-art performance.
Kehua Chen, Zhenlong Yuan, Haihong Xiao, Tianlu Mao
IEEE Trans. Circuits Syst. Video Technol.3
2025 Semantic Scene Completion via Semantic-Aware Guidance and Interactive Refinement Transformer
abstract
Predicting per-voxel occupancy status and corresponding semantic labels in 3D scenes is pivotal to 3D intelligent perception in autonomous driving. In this paper, we propose a novel semantic scene completion framework that can generate complete 3D volumetric semantics from a single image at a low cost. To the best of our knowledge, this is the first endeavor specifically aimed at mitigating the negative impacts of incorrect voxel query proposals caused by erroneous depth estimates and enhancing interactions for positive ones in camera-based semantic scene completion tasks. Specifically, we present a straightforward yet effective Semantic-aware Guided (SAG) module, which seamlessly integrates with task-related semantic priors to facilitate effective interactions between image features and voxel query proposals in a plug-and-play manner. Furthermore, we introduce a set of learnable object queries to better perceive objects within the scene. Building on this, we propose an Interactive Refinement Transformer (IRT) block, which iteratively updates voxel query proposals to enhance the perception of semantics and objects within the scene by leveraging the interaction between object queries and voxel queries through query-to-query cross-attention. Extensive experiments demonstrate that our method outperforms existing state-of-the-art approaches, achieving overall improvements of 0.30 and 2.74 in mIoU metric on the SemanticKITTI and SSCBench-KITTI-360 validation datasets, respectively, while also showing superior performance in the aspect of small object generation.
Haihong Xiao, Wenxiong Kang, Hao Liu 0061, Yuqiong Li, Ying He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 $\mathrm{Tri^{2}plane}$: Advancing Neural Implicit Surface Reconstruction for Indoor Scenes
abstract
Reconstructing 3D indoor scenes presents significant challenges, requiring models capable of inferring both planar surfaces and intricate details. Although recent methods can generate complete surfaces, they often struggle to simultaneously reconstruct low-texture regions and high-frequency details due to non-local effects. In this paper, we introduce a novel triangle-based triplane representation, named (tri$^{2}$plane), specifically designed to account for the diverse spatial feature distribution and information density of indoor environments. Our method begins by projecting point clouds onto three orthogonal planes, followed by 2D Delaunay triangulation. This representation enables adaptive encoding of low-texture and high-frequency regions by employing triangles of variable sizes. Moreover, we develop a dual tri$^{2}$plane framework that incorporates both geometric and semantic information, significantly enhancing the reconstruction quality. We combine these key modules and evaluate our method on benchmark indoor scene datasets. The results unequivocally demonstrate the superiority of our proposed method over the state-of-the-art Occ-SDF. Specifically, our method achieves significant improvements over Occ-SDF, with margins of 1.3, 1.7, and 2.3in F-score on the ScanNet, Tanks & Temples, and Replica datasets, respectively. To facilitate further research, we will make our code publicly available.
Haihong Xiao, Wenxiong Kang
IEEE Trans. Multim.2
2024 EA-MVSNet: Learning Error-Awareness for Enhanced Multi-View Stereo
abstract
Multi-view stereo (MVS) aims to reconstruct the dense 3D geometry of a scene by processing and relating images captured from different viewpoints. Despite impressive successes, most existing techniques simply supervise cost volumes or depth maps through conventional classification or regression methods, thereby inadequately exploring the depth representation’s full potential. Moreover, reconstructing areas with occlusions or weak textures continues to be a long-standing challenge within MVS. Another critical issue, frequently neglected, is the potential inaccuracy of ground truth depths, as evidenced in datasets like DTU. To address these problems, we introduce EA-MVSNet, an innovative error-aware MVS framework designed to enhance depth prediction. The key contributions of this work include three parts: (1) We present a novel error-aware depth representation that enhances depth prediction accuracy through error-aware learning, thereby improving reconstruction quality. (2) We develop a Deformable Feature Pyramid Network (DFPN), meticulously designed to augment reconstruction details in occluded and texture-deficient areas. (3) We introduce a cross-view consistency guidance module into the learning process, effectively mitigating the detrimental effects of ground truth depth inaccuracies and fostering faster convergence. Comprehensive experiments on the DTU dataset and Tanks and Temples dataset validate the superiority of our EA-MVSNet. Compared to the preceding UniMVSNet, EA-MVSNet achieves a notable 7.6% decrease in overall reconstruction error on the DTU dataset, and boosts the mean F-score by 3.0% and 4.1% in the intermediate and advanced groups of the Tanks and Temples dataset, respectively, surpassing most recent state-of-the-art methods.
Wencong Gu, Haihong Xiao, Xueyan Zhao, Wenxiong Kang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Point Cloud Completion via Self-Projected View Augmentation and Implicit Field Constraint
abstract
Recent advances in point cloud completion make it possible to simultaneously recover complete shapes and fine details from partial point clouds captured by professional 3D devices, such as Lidar, or consumer cameras, such as iPhones. Despite significant progress, the potential utilization of self-projected views from partial inputs and the effective reduction of noise in generated point clouds remain under-explored. In this paper, we propose a novel point cloud completion method that leverages self-projected view augmentation and implicit field constraints. Specifically, we introduce a cross-view augmentation (CVA) module and a cross-modal fusion (CMF) module to enhance information interaction and integration at the image and modality levels, respectively. We also propose a bidirection-aware refinement block to improve detail and completeness by considering both complete-to-partial detail perception and partial-to-complete structure perception paths. Additionally, we address the issue of noise reduction from the perspective of implicit field constraints. We evaluate our method on several baseline datasets, including PCN, ShapeNet55/34 and KITTI (car). Extensive experiments demonstrate that our method outperforms state-of-the-art methods, achieving improvements of 0.11 CD-$\ell _{1}$, 0.015 DCD and 0.009 F-score on the standard PCN test set. Furthermore, our approach effectively reduces noise in the generated point clouds, showcasing its promising potential for practical applications.
Haihong Xiao, Ying He 0001, Hao Liu 0061, Wenxiong Kang, Yuqiong Li
IEEE Trans. Circuits Syst. Video Technol.1
2024 Instance-Aware Monocular 3D Semantic Scene Completion
abstract
We study outdoor 3D scene understanding, a challenging task demanding the intelligent system to infer both geometry and semantics from a single-view image – a critical skill for autonomous vehicles to navigate in the real 3D world. Towards this end, we present an instance-aware monocular semantic scene completion framework. To the best of our knowledge, this is the first endeavor specifically targeting the challenge of instance perception in the camera-based semantic scene completion task. Our method consists of two stages. In stage I, we design a region-based VQ-VAE network, providing an effective solution for 3D occupancy prediction. In stage II, we first introduce an instance-aware attention module, explicitly incorporating instance-level cues captured from mask images to enhance the instance features in RGB images. Then we leverage the deformable cross-attention to aggregate image features corresponding to each voxel query and utilize the deformable self-attention to refine query proposals. We combine these key ingredients and evaluate our method on two challenging datasets, namely SemanticKITTI and SSCBench-KITTI-360. The results unequivocally demonstrate the superiority of our proposed method over the state-of-the-art VoxFormer-S. Specifically, our method surpasses VoxFormer-S by 0.22 IoU and 0.72 mIoU on the validation set and achieves an impressive improvement of 3.04 IoU and 1.06 mIoU on the SSCBench-KITTI-360 validation set. Meanwhile, our approach ensures accurate perception of critical instances, thereby exhibiting its exceptional performance and potential for practical deployment.
Haihong Xiao, Wenxiong Kang, Yuqiong Li
IEEE Trans. Intell. Transp. Syst.1
2023 PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering
abstract
Semantic segmentation of point clouds usually requires exhausting efforts of human annotations, hence it attracts wide attention to the challenging topic of learning from unlabeled or weaker forms of annotations. In this paper, we take the first attempt for fully unsupervised semantic segmentation of point clouds, which aims to delineate semantically meaningful objects without any form of annotations. Previous works of unsupervised pipeline on 2D images fails in this task of point clouds, due to: 1) Clustering Ambiguity caused by limited magnitude of data and imbalanced class distribution; 2) Irregularity Ambiguity caused by the irregular sparsity of point cloud. Therefore, we propose a novel framework, PointDC, which is comprised of two steps that handle the aforementioned problems respectively: Cross-Modal Distillation (CMD) and Super-Voxel Clustering (SVC). In the first stage of CMD, multi-view visual features are back-projected to the 3D space and aggregated to a unified point feature to distill the training of the point representation. In the second stage of SVC, the point features are aggregated to super-voxels and then fed to the iterative clustering process for excavating semantic classes. PointDC1yields a significant improvement over the prior state-of-the-art unsupervised methods, on both the ScanNet-v2 (+18.4 mIoU) and S3DIS (+11.5 mIoU) semantic segmentation benchmarks.
Zisheng Chen, Haihong Xiao, Baigui Sun, Xuansong Xie, Wenxiong Kang
ICCV5
2023 Semi-supervised Deep Multi-view Stereo
abstract
Significant progress has been witnessed in learning-based Multi-view Stereo (MVS) under supervised and unsupervised settings. To combine their respective merits in accuracy and completeness, meantime reducing the demand for expensive labeled data, this paper explores the problem of learning-based MVS in a semi-supervised setting that only a tiny part of the MVS data is attached with dense depth ground truth. However, due to huge variation of scenarios and flexible settings in views, it may break the basic assumption in classic semi-supervised learning, that unlabeled data and labeled data share the same label space and data distribution, named as semi-supervised distribution-gap ambiguity in the MVS problem. To handle these issues, we propose a novel semi-supervised distribution-augmented MVS framework, namely SDA-MVS. For the simple case that the basic assumption works in MVS data, consistency regularization encourages the model predictions to be consistent between original sample and randomly augmented sample. For further troublesome case that the basic assumption is conflicted in MVS data, we propose a novel style consistency loss to alleviate the negative effect caused by the distribution gap. The visual style of unlabeled sample is transferred to labeled sample to shrink the gap, and the model prediction of generated sample is further supervised with the label in original labeled sample. The experimental results in semi-supervised settings of multiple MVS datasets show the superior performance of the proposed method. With the same settings in backbone network, our proposed SDA-MVS outperforms its fully-supervised and unsupervised baselines.
Yang Liu 0356, Haihong Xiao, Baigui Sun, Xuansong Xie, Wenxiong Kang
ACM Multimedia5
2023 Distinguishing and Matching-Aware Unsupervised Point Cloud Completion
abstract
Real-scanned point clouds are often incomplete due to occlusion, light reflection and limitations of sensor resolution, which impedes the related progress of downstream tasks, e.g., shape classification and object detection. Although there has been impressive research progress on the point cloud completion topic, they rely on the premise of extensive paired training data. However, collecting complete point clouds in some specified scenarios is labor-intensive and even impractical. To mitigate this problem, we propose DMNet, a distinguishing and matching-aware unsupervised point cloud completion network. Our work belongs to the group of unsupervised completion methods but goes beyond previous studies. Firstly, we propose a distinguishing-aware feature extractor to learn discriminable semantic information for different instances, simultaneously enhancing the robust invariant representation under noise disturbances. Secondly, we design a hierarchy-aware hyperbolic decoder to recover the complete geometry of point clouds, which not only can capture the implicit hierarchical relationships in data but also has an explicit extended nature. Finally, we develop a matching-aware refiner to eliminate noise points via aligning the topology structure of the input and predicted partial point clouds. Extensive experiments on MVP, Completion3D and KITTI datasets prove the effectiveness of our method, which performs favorably over state-of-the-art methods both quantitatively and qualitatively.
Haihong Xiao, Yuqiong Li, Wenxiong Kang, Qiuxia Wu
IEEE Trans. Circuits Syst. Video Technol.1