VLDB 2026 Research / reviewers in the wild / expert
Shi-Sheng Huang
dblp:58/8294
· DBLP profile ↗
28ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised Depth Completion Guided by 3D Perception and Geometry ConsistencyabstractDepth completion which aims at predicting dense depth maps from sparse depth measurements, plays a crucial role in many computer graphics and computer vision applications. Previous supervised learning based approaches have demonstrated overwhelming success in this task, while unsupervised high-precision depth completion without relying on the ground-truth data still remains challenging. One main drawback of most previous unsupervised solutions comes from the ignorance of 3D structural information, which often leads to inaccurate spatial propagation and mixed-depth problems. To alleviate the above challenges, this paper explores the utilization of 3D perceptual features and multi-view geometry consistency to devise a high-precision self-supervised depth completion method. Our key contribution is a 3D perceptual spatial propagation constructed with a point cloud representation and an attention weighting mechanism, to capture more reasonable and favorable neighbors during the depth propagation process. Based on the 3D perceptual spatial propagation, we also introduce multi-view geometric constraints between adjacent views to guide the optimization of the whole depth completion model, which achieves geometry consistent depth completion in a self-supervised manner. Extensive experiments on benchmark datasets of NYU-Depth-v2, VOID and KITTI Depth Completion demonstrate that the proposed model achieves the state-of-the-art depth completion performance compared with other unsupervised methods, and even competitive performance compared with previous supervised methods. Tianyu Shen, Shi-Sheng Huang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | GANG: Geometrically-Aligned Neural Gaussians for Efficient and Realistic RelightingabstractEfficient and realistic relighting of complex scenes with unknown illumination remains a crucial but challenging task. Recent advancements in 3D Gaussian Splatting (3DGS) have shown impressive object-level relighting. However, they still struggle with complex real-world scenes, mainly due to the challenges of accurately decoupling intricate geometry, materials, and lighting using concise 3D Gaussian primitives. In this paper, we propose a new Geometrically-Aligned Neural Gaussian Splatting (GANG) method, which performs efficient physically based rendering (PBR) directly on anchor-based relightable neural Gaussians. Our key idea is to regularize the decoded neural Gaussians geometrically aligned with the latent signed distance field (SDF) surface spawned from anchors using a differentiable implicit indicator function (IIF) solver. It brings effective geometric association to accurate decoupling of materials and lighting for efficient and realistic relighting of complex scenes. Furthermore, we propose a locally consistent geometry regularization to guide more concise neural Gaussian learning with a hybrid lighting model, which combines position-learnable spherical Gaussians (SGs) and an environment map, allowing accurate modeling of both local and global illumination. Experimental results on public datasets demonstrate that GANG consistently outperforms previous PBR methods in material decomposition and relighting quality, while representing complex scenes with concise anchors. To the best of our knowledge, GANG is a new state-of-the-art 3DGS method for realistic relighting, enabling efficient rendering and flexible editing materials and illumination, especially for complex scenes. Deqi Li, Shi-Sheng Huang, Hongbo Fu 0001, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray DiffusionabstractAccurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings, but could easily fail for sparse-view scenarios without sufficient visual overlap. In this paper, we propose a new technique for pose-free surface reconstruction, which follows triplane-based signed distance field (SDF) learning but regularizes the learning by explicit points sampled from ray-based diffusion of camera pose estimation. Our key contribution is a novel Geometric Consistent Ray Diffusion model (GCRayDiffusion), where we represent camera poses as neural bundle rays and regress the distribution of noisy rays via a diffusion model. More importantly, we further condition the denoising process of RGRayDiffusion using the triplane-based SDF of the entire scene, which provides effective 3D consistent regularization to achieve multi-view consistent camera pose estimation. Finally, we incorporate RGRayDiffusion into the triplane-based SDF learning by introducing on-surface geometric regularization from the sampling points of the neural bundle rays, which leads to highly accurate pose-free surface reconstruction results even for sparse-view inputs. Extensive evaluations on public datasets show that our GCRayDiffusion achieves more accurate camera pose estimation than previous approaches, with geometrically more consistent surface reconstruction results, especially given sparse-view inputs. Li-Heng Chen, Zixin Zou, Tianjiao Jing, Yan-Pei Cao 0001, Shi-Sheng Huang, Hongbo Fu 0001, Hua Huang 0001 |
ICCV | 6 |
| 2025 | RTMap: Real-Time Recursive Mapping with Change Detection and LocalizationabstractWhile recent online HD mapping methods relieve burdened offline pipelines and solve map freshness, they remain limited by perceptual inaccuracies, occlusion in dense traffic, and an inability to fuse multi-agent observations. We propose RTMap to enhance these single-traversal methods by persistently crowdsourcing a multi-traversal HD map as a self-evolutional memory. On onboard agents, RTMap simultaneously addresses three core challenges in an end-to-end fashion: (1) Uncertainty-aware positional modeling for HD map elements, (2) probabilistic-aware localization w.r.t. the crowdsourced prior-map, and (3) real-time detection for possible road structural changes. Experiments on several public autonomous driving datasets demonstrate our solid performance on both the prior-aided map quality and the localization accuracy, demonstrating our effectiveness of robustly serving downstream prediction and planning modules while gradually improving the accuracy and freshness of the crowdsourced prior-map asynchronously. Our source-code will be made publicly available at https://github.com/CN-ADLab/RTMap. Yuheng Du, Sheng Yang 0007, Lingxuan Wang, Zhenghua Hou, Chengying Cai, Zhitao Tan, Mingxia Chen, Shi-Sheng Huang |
ICCV | 8 |
| 2025 | Spatio-Temporally Consistent Depth Estimation for Dynamic Scenes using 3D Scene FlowsabstractDynamic depth estimation continues to be crucial but challenging mainly due to the violation of multi-view consistency raised by dynamic areas. Recent approaches have made impressive progress by implicitly fusing the intra-relation features, but is still limited for heterogeneous dynamic scenes. In this paper, we propose a new intra-relation feature fusion, which can significantly improve the fusion quality using an explicit regularization from 3D scene flow cues. We first introduces a Dual Cross-Cue Fusion (D-CCF) module for depth prediction, and further build up an efficient 3D scene flow estimation as explicit 3D spatio-temporal corresponding priors to regularize the depth prediction. Finally, by jointly learning both the depth prediction and 3D scene flow estimation in a unsupervised manner, we achieve more accurate dynamic depth estimation towards spatio-temporal consistency. By extensive evaluation on challenging benchmarks (KITTI and DDAD), our approach can achieve better depth estimation results than state-of-the-art approaches in both static and dynamic areas, which especially maintains the spatio-temporal consistency for dynamic scenes. Tianjiao Jing, Zhengxuan Lian, Shi-Sheng Huang, Hua Huang 0001 |
ICME | 5 |
| 2025 | GS-RoadPatching: Inpainting Gaussians via 3D Searching and Placing for Driving ScenesabstractThis paper presents GS-RoadPatching, an inpainting method for driving scene completion by referring to completely reconstructed regions, which are represented by 3D Gaussian Splatting (3DGS). Unlike existing 3DGS inpainting methods that perform generative completion relying on 2D perspective-view-based diffusion or GAN models to predict limited appearance or depth cues for missing regions, our approach enables substitutional scene inpainting and editing directly through the 3DGS modality, extricating it from requiring spatial-temporal consistency of 2D cross-modals and eliminating the need for time-intensive retraining of Gaussians. Our key insight is that the highly repetitive patterns in driving scenes often share multi-modal similarities within the implicit 3DGS feature space and are particularly suitable for structural matching to enable effective 3DGS-based substitutional inpainting. Practically, we construct feature-embedded 3DGS scenes to incorporate a patch measurement method for abstracting local context at different scales and, subsequently, propose a structural search method to find candidate patches in 3D space effectively. Finally, we propose a simple yet effective substitution-and-fusion optimization for better visual harmony. We conduct extensive experiments on multiple publicly available datasets to demonstrate the effectiveness and efficiency of our proposed method in driving scenes, and the results validate that our method achieves state-of-the-art performance compared to the baseline methods in terms of both quality and interoperability. Additional experiments in general scenes also demonstrate the applicability of the proposed 3D inpainting strategy. The project page and code are available at: https://shanzhaguoo.github.io/GS-RoadPatching/. Jiarun Liu, Sicong Du, Chenming Wu, Deqi Li, Shi-Sheng Huang, Guofeng Zhang 0001, Sheng Yang 0007 |
SIGGRAPH Asia | 6 |
| 2025 | A Survey of Recent Advances in Generative 3D Reconstruction
Shi-Sheng Huang, Shao-Kui Zhang, Sheng Yang 0007, Jian-Wei Guo, Hua Huang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2025 | RGAvatar: Relightable 4D Gaussian Avatar From Monocular VideosabstractRelightable 4D avatar reconstruction which enables high fidelity and real-time rendering continues to be a crucial but challenging problem, especially from monocular videos. Previous NeRF-based 4D avatars enable photo-realistic relighting but are too slow for rendering, while point-based or mesh-based 4D avatars are efficient but have limited rendering quality. The recent success of 3D Gaussian Splatting, i.e., 3DGS, has inspired a series of impressive 4D Gaussian avatars, however, most of which only focus on faithful appearance reconstruction but are not relightable. To address such issues, this article proposes a new Relightable 4D Gaussian Avatar, i.e., RGAvatar, tailored for high fidelity relightable rendering from monocular videos. Our key idea is to introduce a new relightable 4D Gaussian representation, based on which we can directly perform high fidelity Physically Based Rendering, and an effective joint learning mechanism for compact 4D Gaussian reconstruction with SDF regulation and accurate materials and lighting decomposition. By comparing with previous state-of-the-art approaches, RGAvatar can significantly outperform previous approaches in relightable rendering quality and speed. To our best knowledge, RGAvatar contributes a new state-of-the-art 4D Gaussian avatar from monocular videos, which enables high fidelity relightable rendering in a quite efficient manner. Zhe Fan, Shi-Sheng Huang, Dachao Shang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | MPGS: Multi-Plane Gaussian Splatting for Compact Scenes RenderingabstractAccurate reconstruction of heterogeneous scenes for high-fidelity rendering in an efficient manner remains a crucial but challenging task in many Virtual Reality and Augmented Reality applications. The recent 3D Gaussian Splatting (3DGS) has shown impressive quality in scene rendering with real-time performance. However, for heterogeneous scenes with many weak-textured regions, the original 3DGS can easily produce numerously wrong floaters with unbalanced reconstruction using redundant 3D Gaussians, which often leads to unsatisfied scene rendering. This paper proposes a novel multi-plane Gaussian Splatting (MPGS), which aims to achieve high-fidelity rendering with compact reconstruction for heterogeneous scenes. The key insight of our MPGS is the introduction of a novel multi-plane Gaussian optimization strategy, which effectively adjusts the Gaussian distribution for both rich-textured and weak-textured regions in heterogeneous scenes. Moreover, we further propose a multi-scale geometric correction mechanism to effectively mitigate degradation of the 3D Gaussian distribution for compact scene reconstruction. Besides, we regularize the Gaussian distributions using normal information extracted from the compact scene learning. Experimental results on public datasets demonstrate that the proposed MPGS achieves much better rendering quality compared to previous methods, while using less storage and offering more efficient rendering. To our best knowledge, MPGS is a new state-of-the-art 3D Gaussian splatting method for compact reconstruction of heterogeneous scenes, enabling high-fidelity rendering in novel view synthesis, especially improving rendering quality for weak-textured regions. The code will be released at https://github.com/wanglids/MPGS. Deqi Li, Shi-Sheng Huang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | EGAvatar: Efficient GAN Inversion for Generalizable Head Avatar From Few-Shot ImagesabstractControllable head avatar reconstruction via the inversion of few-shot images using 3D generative models has demonstrated significant potential for efficient avatar creation. However, under limited input conditions, existing one-shot inversion methods often fail to produce high-fidelity results, frequently leading to shape distortions, expression deviations, and identity inconsistencies. To address these limitations, we propose EGAvatar, a novel and efficient 3DGAN inversion framework designed to generate high-fidelity, generalizable head avatars from few-shot images. The core principle of EGAvatar is a decoupling-by-inverting strategy, built upon an animatable 3DGAN prior. Specifically, we introduce an effective animatable 3DGAN model that synthesizes high-quality 3D avatars by integrating a coarse 3D triplane representation (derived from a latent 3DGAN) with an offset 3D triplane (learned via a triplane 3DGAN). Leveraging this architecture, we design a 3DGAN-based inversion approach to reconstruct 3D avatars efficiently. Additionally, we incorporate an expression-view disentanglement mechanism to maintain consistent appearance across varying expressions and viewpoints, thereby enhancing the generalizability of avatar reconstruction from limited input images. Extensive experiments conducted on two publicly available benchmarks and a private dataset demonstrate that EGAvatar outperforms existing state-of-the-art methods in both qualitative and quantitative evaluations. Notably, EGAvatar achieves superior performance while requiring significantly fewer input images and offering more efficient training and inference. Hao Pan Ren, Wan Yu Li, Shi-Sheng Huang, Juyong Zhang, Hua Huang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | GP-Recon: Online Monocular Neural 3D Reconstruction With Geometric PriorabstractHigh-fidelity online 3D scene reconstruction from monocular videos continues to be challenging, especially for coherent and fine-grained geometry reconstruction. The previous learning-based online 3D reconstruction approaches with neural implicit representations have shown a promising ability for coherent scene reconstruction, but often fail to consistently reconstruct fine-grained geometric details during online reconstruction. This paper presents a new on-the-fly monocular 3D reconstruction approach, named GP-Recon, to perform high-fidelity online neural 3D reconstruction with fine-grained geometric details. We incorporate geometric prior (GP) into a scene's neural geometry learning to better capture its geometric details and, more importantly, propose an online volume rendering optimization to reconstruct and maintain geometric details during the online reconstruction task. The extensive comparisons with state-of-the-art approaches show that our GP-Recon consistently generates more accurate and complete reconstruction results with much better fine-grained details, both quantitatively and qualitatively. Zixin Zou, Shi-Sheng Huang, Yan-Pei Cao 0001, Tai-Jiang Mu, Ying Shan, Hongbo Fu 0001, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy ViewsabstractThe recent neural surface reconstruction approaches using volume rendering have made much progress by achieving impressive surface reconstruction quality, but are still limited to dense and highly accurate posed views. To overcome such drawbacks, this paper pays special attention on the consistent surface reconstruction from sparse views with noisy camera poses. Unlike previous approaches, the key difference of this paper is to exploit the multi-view constraints directly from the explicit geometry of the neural surface, which can be used as effective regularization to jointly learn the neural surface and refine the camera poses. To build effective multi-view constraints, we introduce a fast differentiable on-surface intersection to generate on-surface points, and propose view-consistent losses on such differentiable points to regularize the neural surface learning. Based on this point, we propose a joint learning strategy, named SC-NeuS, to perform geometry-consistent surface reconstruction in an end-to-end manner. With extensive evaluation on public datasets, our SC-NeuS can achieve consistently better surface reconstruction results with fine-grained details than previous approaches, especially from sparse and noisy camera views. The source code is available at https://github.com/zouzx/sc-neus.git. Shi-Sheng Huang, Zixin Zou, Yan-Pei Cao 0001, Ying Shan |
AAAI | 1 |
| 2024 | Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsabstractReconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distilling pre-trained diffusion priors into 3D representations using score distillation sampling (SDS), these methods often struggle to simultaneously achieve high-quality, consistent, and detailed results for both novel-view synthesis (NVS) and geometry. In this work, we present Sparse3D, a novel 3D reconstruction method tailored for sparse view inputs. Our approach distills robust priors from a multiview-consistent diffusion model to refine a neural radiance field. Specifically, we employ a controller that harnesses epipolar features from input views, guiding a pre-trained diffusion model, such as Stable Diffusion, to produce novel-view images that maintain 3D consistency with the input. By tapping into 2D priors from powerful image diffusion models, our integrated model consistently delivers high-quality results, even when faced with open-world objects. To address the blurriness introduced by conventional SDS, we introduce the category-score distillation sampling (C-SDS) to enhance detail. We conduct experiments on CO3DV2 which is a multi-view dataset of real-world objects. Both quantitative and qualitative evaluations demonstrate that our approach outperforms previous state-of-the-art works on the metrics regarding NVS and geometry reconstruction. Zixin Zou, Weihao Cheng 0002, Yan-Pei Cao 0001, Shi-Sheng Huang, Ying Shan, Song-Hai Zhang |
AAAI | 4 |
| 2024 | NeuralIndicator: Implicit Surface Reconstruction from Neural Indicator PriorsabstractThe neural implicit surface reconstruction from unorganized points is still challenging, especially when the point clouds are incomplete and/or noisy with complex topology structure. Unlike previous approaches performing neural implicit surface learning relying on local shape priors, this paper proposes to utilize global shape priors to regularize the neural implicit function learning for more reliable surface reconstruction. To this end, we first introduce a differentiable module to generate a smooth indicator function, which globally encodes both the indicative prior and local SDFs of the entire input point cloud. Benefit from this, we propose a new framework, called NeuralIndicator, to jointly learn both the smooth indicator function and neural implicit function simultaneously, using the global shape prior encoded by smooth indicator function to effectively regularize the neural implicit function learning, towards reliable and high-fidelity surface reconstruction from unorganized points without any normal information. Extensive evaluations on synthetic and real-scan datasets show that our approach consistently outperforms previous approaches, especially when point clouds are incomplete and/or noisy with complex topology structure. Shi-Sheng Huang, Chen Li Heng, Hua Huang 0001 |
ICML | 1 |
| 2023 | Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse ViewsabstractSignificant progress has been made in realizing novel view synthesis of dynamic scenes from sparse input views. However, achieving spatio-temporal consistency in dynamic view synthesis remains to be challenging for previous approaches, since the spatio-temporal correlation for view synthesis has not been fully explored. In this paper, we propose a spatio-temporal feature warping (STFW) mechanism, which can be embedded into a deep model to produce high-quality and spatio-temporally consistent view synthesis results. The two core components of STFW are: (1) a spatial feature warping (SFW) module, which enables adaptive perception of multi-view context-consistent geometric information with a compact point cloud representation, and (2) a temporal feature warping (TFW) module that implicitly models the dynamic geometry by approaching the pixel shift in image coordinate. In the optimization process of view synthesis, the SFW and TFW are integrated to exploit the spatio-temporal correlation cues across sparse input views and novel views. Leveraging the STFW, we further build an end-to-end dynamic view synthesis model with sparse input views. Qualitative and quantitative evaluation on public multi-view datasets demonstrate that our view synthesis pipeline achieves better performance compared to previous methods in terms of visual quality. Deqi Li, Shi-Sheng Huang, Tianyu Shen, Hua Huang 0001 |
ACM Multimedia | 2 |
| 2023 | VirtualClassroom: A Lecturer-Centered Consumer-Grade Immersive Teaching System in Cyber-Physical-Social SpaceabstractLecturers, as the guidance of the classroom, play a significant role in the teaching process. However, the lecturers’ sense of space immersion has been ignored in current virtual teaching systems. In this article, we explore the cyber–physical–social intelligence for Edu-Metaverse in cyber–physical–social space and specially design a lecturer-centered immersive teaching system, taking the social and lecturers’ factors into consideration. We call this system VirtualClassroom (V-Classroom). Specifically, we first introduce the cyber–physical–social system (CPSS) paradigm of V-Classroom so that the workflow is standardized and significantly simplified, and the systems can be constructed with off-the-shelf hardware. The key component of V-Classroom is a cyber-world representation of a physical-world classroom instrumented with sparse consumer-grade RGBD cameras for capturing the 3-D geometry and texture of the classrooms. We provide each V-Classroom lecturer with a physical device for sending 6DoF view-change messages and showing view-dependent content of the remote classroom. Following the above paradigm, we develop the V-Classroom algorithms, including V-Classroom depth algorithm (V-DA) and V-Classroom view algorithm (V-VA), to achieve the real-time rendering of remote classrooms. V-DA is dedicated to recovering accurate depth information of the classrooms while V-VA is devoted to real-time novel view synthesis. Finally, we illustrate our implemented CPSS-driven V-Classroom prototype, based on real-world classroom scenarios we collected, and discuss the main challenges and future direction. Tianyu Shen, Shi-Sheng Huang, Deqi Li, Fei-Yue Wang 0001, Hua Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Real-Time Globally Consistent 3D Reconstruction With Semantic PriorsabstractMaintaining global consistency continues to be critical for online 3D indoor scene reconstruction. However, it is still challenging to generate satisfactory 3D reconstruction in terms of global consistency for previous approaches using purely geometric analysis, even with bundle adjustment or loop closure techniques. In this article, we propose a novel real-time 3D reconstruction approach which effectively integrates both semantic and geometric cues. The key challenge is how to map this indicative information, i.e., semantic priors, into a metric space as measurable information, thus enabling more accurate semantic fusion leveraging both the geometric and semantic cues. To this end, we introduce a semantic space with a continuous metric function measuring the distance between discrete semantic observations. Within the semantic space, we present an accurate frame-to-model semantic tracker for camera pose estimation, and semantic pose graph equipped with semantic links between submaps for globally consistent 3D scene reconstruction. With extensive evaluation on public synthetic and real-world 3D indoor scene RGB-D datasets, we show that our approach outperforms the previous approaches for 3D scene reconstruction both quantitatively and qualitatively, especially in terms of global consistency. Shi-Sheng Huang, Haoxiang Chen 0004, Hongbo Fu 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | ObjectFusion: Accurate object-level SLAM with neural object priors
Zixin Zou, Shi-Sheng Huang, Tai-Jiang Mu, Yu-Ping Wang 0001 |
Graph. Model. | 2 |
| 2022 | Accurate Dynamic SLAM Using CRF-Based Long-Term ConsistencyabstractAccurate camera pose estimation is essential and challenging for real world dynamic 3D reconstruction and augmented reality applications. In this article, we present a novel RGB-D SLAM approach for accurate camera pose tracking in dynamic environments. Previous methods detect dynamic components only across a short time-span of consecutive frames. Instead, we provide a more accurate dynamic 3D landmark detection method, followed by the use of long-term consistency via conditional random fields, which leverages long-term observations from multiple frames. Specifically, we first introduce an efficient initial camera pose estimation method based on distinguishing dynamic from static points using graph-cut RANSAC. These static/dynamic labels are used as priors for the unary potential in the conditional random fields, which further improves the accuracy of dynamic 3D landmark detection. Evaluation using the TUM and Bonn RGB-D dynamic datasets shows that our approach significantly outperforms state-of-the-art methods, providing much more accurate camera trajectory estimation in a variety of highly dynamic environments. We also show that dynamic 3D reconstruction can benefit from the camera poses estimated by our RGB-D SLAM approach. Zheng-Jun Du, Shi-Sheng Huang, Tai-Jiang Mu, Qunhe Zhao, Ralph R. Martin, Kun Xu 0003 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | DI-Fusion: Online Implicit 3D Reconstruction With Deep PriorsabstractPrevious online 3D dense reconstruction methods struggle to achieve the balance between memory storage and surface quality, largely due to the usage of stagnant underlying geometry representation, such as TSDF (truncated signed distance functions) or surfels, without any knowledge of the scene priors. In this paper, we present DI-Fusion (Deep Implicit Fusion), based on a novel 3D representation, i.e. Probabilistic Local Implicit Voxels (PLIVoxs), for online 3D reconstruction with a commodity RGB-D camera. Our PLIVox encodes scene priors considering both the local geometry and uncertainty parameterized by a deep neural network. With such deep priors, we are able to perform online implicit 3D reconstruction achieving state-of-the-art camera trajectory estimation accuracy and mapping quality, while achieving better storage efficiency compared with previous online 3D reconstruction approaches. Our implementation is available at https://www.github.com/huangjh-pub/di-fusion. Shi-Sheng Huang, Haoxuan Song, Shi-Min Hu 0001 |
CVPR | 2 |
| 2021 | Supervoxel Convolution for Online 3D Semantic SegmentationabstractOnline 3D semantic segmentation, which aims to perform real-time 3D scene reconstruction along with semantic segmentation, is an important but challenging topic. A key challenge is to strike a balance between efficiency and segmentation accuracy. There are very few deep-learning-based solutions to this problem, since the commonly used deep representations based on volumetric-grids or points do not provide efficient 3D representation and organization structure for online segmentation. Observing that on-surface supervoxels, i.e., clusters of on-surface voxels, provide a compact representation of 3D surfaces and brings efficient connectivity structure via supervoxel clustering, we explore a supervoxel-based deep learning solution for this task. To this end, we contribute a novel convolution operation (SVConv) directly on supervoxels. SVConv can efficiently fuse the multi-view 2D features and 3D features projected on supervoxels during the online 3D reconstruction, and leads to an effective supervoxel-based convolutional neural network, termed as Supervoxel-CNN , enabling 2D-3D joint learning for 3D semantic prediction. With the Supervoxel-CNN , we propose a clustering-then-prediction online 3D semantic segmentation approach. The extensive evaluations on the public 3D indoor scene datasets show that our approach significantly outperforms the existing online semantic segmentation systems in terms of efficiency or accuracy. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 1 |
| 2020 | Lidar-Monocular Visual Odometry using Point and Line FeaturesabstractWe introduce a novel lidar-monocular visual odometry approach using point and line features. Compared to previous point-only based lidar-visual odometry, our approach leverages more environment structure information by introducing both point and line features into pose estimation. We provide a robust method for point and line depth extraction, and formulate the extracted depth as prior factors for point-line bundle adjustment. This method greatly reduces the features' 3D ambiguity and thus improves the pose estimation accuracy. Besides, we also provide a purely visual motion tracking method and a novel scale correction scheme, leading to an efficient lidar-monocular visual odometry system with high accuracy. The evaluations on the public KITTI odometry benchmark show that our technique achieves more accurate pose estimation than the state-of-the-art approaches, and is sometimes even better than those leveraging semantic information. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ICRA | 1 |
| 2016 | Structure guided interior scene synthesis via graph matching
Shi-Sheng Huang, Hongbo Fu 0001, Shi-Min Hu 0001 |
Graph. Model. | 1 |
| 2016 | Support Substructures: Support-Induced Part-Level Structural RepresentationabstractIn this work we explore a support-induced structural organization of object parts. We introduce the concept of support substructures, which are special subsets of object parts with support and stability. A bottom-up approach is proposed to identify such substructures in a support relation graph. We apply the derived high-level substructures to part-based shape reshuffling between models, resulting in nontrivial functionally plausible model variations that are difficult to achieve with symmetry-induced substructures by the state-of-the-art methods. We also show how to automatically or interactively turn a single input model to new functionally plausible shapes by structure rearrangement and synthesis, enabled by support substructures. To the best of our knowledge no single existing method has been designed for all these applications. Shi-Sheng Huang, Hongbo Fu 0001, Ling-Yu Wei, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Parametric meta-filter modeling from a single example pair
Shi-Sheng Huang, Guo-Xin Zhang, Yukun Lai, Johannes Kopf 0001, Daniel Cohen-Or, Shi-Min Hu 0001 |
Vis. Comput. | 1 |
| 2013 | Qualitative organization of collections of shapes via quartet analysisabstractWe present a method for organizing a heterogeneous collection of 3D shapes for overview and exploration. Instead of relying on quantitative distances, which may become unreliable between dissimilar shapes, we introduce aqualitativeanalysis which utilizes multiple distance measures but only in cases where the measures can be reliably compared. Our analysis is based on the notion ofquartets, each defined by two pairs of shapes, where the shapes in each pair are close to each other, but far apart from the shapes in the other pair. Combining the information from many quartets computed across a shape collection using several distance measures, we create a hierarchical structure we callcategorization treeof the shape collection. This tree satisfies the topological (qualitative) constraints imposed by the quartets creating an effective organization of the shapes. We present categorization trees computed on various collections of shapes and compare them to ground truth data from human categorization. We further introduce the concept ofdegree of separationchart for every shape in the collection and show the effectiveness of using it for interactive shapes exploration. Shi-Sheng Huang, Ariel Shamir, Chao-Hui Shen, Hao (Richard) Zhang, Alla Sheffer, Shi-Min Hu 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 1 |
| 2011 | Adaptive partitioning of urban facadesabstractAutomatically discovering high-level facade structures in unorganized 3D point clouds of urban scenes is crucial for applications like digitalization of real cities. However, this problem is challenging due to poor-quality input data, contaminated with severe missing areas, noise and outliers. This work introduces the concept of adaptive partitioning to automatically derive a flexible and hierarchical representation of 3D urban facades. Our key observation is that urban facades are largely governed by concatenated and/or interlaced grids. Hence, unlike previous automatic facade analysis works which are typically restricted to globally rectilinear grids, we propose to automatically partition the facade in an adaptive manner, in which the splitting direction, the number and location of splitting planes are all adaptively determined. Such an adaptive partition operation is performed recursively to generate a hierarchical representation of the facade. We show that the concept of adaptive partitioning is also applicable to flexible and robust analysis of image facades. We evaluate our method on a dozen of LiDAR scans of various complexity and styles, and the image facades from the eTRIMS database and the Ecole Centrale Paris database. A series of applications that benefit from our approach are also demonstrated. Chao-Hui Shen, Shi-Sheng Huang, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 2 |
| 2010 | Popup: automatic paper architectures from 3D modelsabstractPaper architectures are 3D paper buildings created by folding and cutting. The creation process of paper architecture is often labor-intensive and highly skill-demanding, even with the aid of existing computer-aided design tools. We propose an automatic algorithm for generating paper architectures given a user-specified 3D model. The algorithm is grounded on geometric formulation of planar layout for paper architectures that can be popped-up in a rigid and stable manner, and sufficient conditions for a 3D surface to be popped-up from such a planar layout. Based on these conditions, our algorithm computes a class of paper architectures containing two sets of parallel patches that approximate the input geometry while guaranteed to be physically realizable. The method is demonstrated on a number of architectural examples, and physically engineered results are presented. Xian-Ying Li, Chao-Hui Shen, Shi-Sheng Huang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |