EDBT 2026 Demo / reviewers in the wild / expert
Baorui Ma
dblp:275/3742
· DBLP profile ↗
24ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-7229-2386ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEE4D: Pose-Free 4D Generation via Auto-Regressive Video InpaintingabstractAbstract Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video‐to‐4D methods typically rely on manually annotated camera poses, which are labor‐intensive and brittle for in‐the‐wild footage. Recent warp‐then‐inpaint approaches mitigate the need for pose labels by warping input frames along a novel camera trajectory and using an inpainting model to fill missing regions, thereby depicting the 4D scene from diverse viewpoints. However, this trajectory‐to‐trajectory formulation often entangles camera motion with scene dynamics and complicates both modeling and inference. We introduce S ee 4D , a pose‐free, trajectory‐to‐camera framework that replaces explicit trajectory prediction with rendering to a bank of fixed virtual cameras, thereby separating camera control from scene modeling. A view‐conditional video inpainting model is trained to learn a robust geometry prior by denoising realistically synthesized warped images and to inpaint occluded or missing regions across virtual viewpoints, eliminating the need for explicit 3D annotations. Building on this inpainting core, we design a spatiotemporal autoregressive inference pipeline that traverses virtual‐camera splines and extends videos with overlapping windows, enabling coherent generation at bounded per‐step complexity. We validate See4D on cross‐view video generation and sparse reconstruction benchmarks. Across quantitative metrics and qualitative assessments, our method achieves superior generalization and improved performance relative to pose‐ or trajectory‐conditioned baselines, advancing practical 4D world modeling from casual videos. Dongyue Lu, Ao Liang, Tianxin Huang, Baorui Ma, Liang Pan, Wei Yin 0006, Lingdong Kong, Wei Tsang Ooi, Ziwei Liu 0002 |
Comput. Graph. Forum | 6 |
| 2026 | UDFStudio: A Unified Framework of Datasets, Benchmarks and Generative Models for Unsigned Distance FunctionsabstractUnsigned distance functions (UDFs) have emerged as powerful representation for modeling and reconstructing geometries with open surfaces. However, the development of 3D generative models for UDFs remains largely unexplored, limiting current methods from generating diverse open-surface 3D content. Moreover, mainstream 3D datasets predominantly consist of watertight meshes, revealing a critical challenge: the absence of standardized datasets and benchmarks specifically tailored for open-surface generation and reconstruction. In this paper, we begin by introducing UDiFF, a novel diffusion-based 3D generative model specifically designed for UDFs. UDiFF supports both conditional and unconditional generation of textured 3D shapes with open surfaces. At its core, UDiFF generates UDFs in the spatial-frequency domain using a learnable wavelet transform. Instead of relying on manually selected wavelet transforms, which are labor-intensive and prone to information loss, we introduce a data-driven approach that learns the optimal wavelet transformation from UDFs datasets. Beyond UDiFF, we present the UWings dataset, comprising 1,509 high-quality 3D open-surface models of winged creatures. Using UWings, we establish comprehensive benchmarks for evaluating both generative and reconstruction methods based on UDFs. Junsheng Zhou, Baorui Ma, Kanle Shi, Yu-Shen Liu, Zhizhong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | You See it, You Got it: Learning 3D Creation on Pose-Free Videos at ScaleabstractRecent 3D generation models typically rely on limited-scale 3D ‘gold-labels’ or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning paradigms. In this work, we present See3D, a visual-conditional multi-view diffusion model trained on large-scale Internet videos for open-world 3D creation. The model aims to Get 3D knowledge by solely Seeing the visual contents from the vast and rapidly growing video data — You See it, You Got it. To achieve this, we first scale up the training data using a proposed data curation pipeline that automatically filters out multi-view inconsistencies and insufficient observations from source videos. This results in a high-quality, richly diverse, large-scale dataset of multi-view images, termed WebVi3D, containing 320M frames from 16M video clips. Nevertheless, learning generic 3D priors from videos without explicit 3D geometry or camera pose annotations is nontrivial, and annotating poses for web-scale videos is prohibitively expensive. To eliminate the need for pose conditions, we introduce an innovative visual-condition - a purely 2D-inductive visual signal generated by adding time-dependent noise to the masked video data. Finally, we introduce a novel visual-conditional 3D generation framework by integrating See3D into a warping-based pipeline for high-fidelity 3D generation. Our numerical and visual comparisons on single and sparse reconstruction benchmarks show that See3D, trained on cost- effective and scalable video data, achieves notable zero-shot and open-world generation capabilities, markedly outperforming models trained on costly and constrained 3D datasets. Additionally, our model naturally supports other image-conditioned 3D creation tasks, such as 3D editing, without further fine-tuning. Please refer to our project page at: https://vision.baai.ac.cn/see3d. Baorui Ma, Huachen Gao, Haoge Deng, Tiejun Huang 0001, Lulu Tang |
CVPR | 1 |
| 2025 | NeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene ReconstructionabstractRecently, it has shown that priors are vital for neural implicit functions to reconstruct high-quality surfaces from multi-view RGB images. However, current priors require large-scale pre-training, and merely provide geometric clues without considering the importance of color. In this paper, we present NeRFPrior, which adopts a neural radiance field as a prior to learn signed distance fields using volume rendering for surface reconstruction. Our NeRF prior can provide both geometric and color clues, and also get trained fast under the same scene without additional data. Based on the NeRF prior, we are enabled to learn a signed distance function (SDF) by explicitly imposing a multi-view consistency constraint on each ray intersection for surface inference. Specifically, at each ray intersection, we use the density in the prior as a coarse geometry estimation, while using the color near the surface as a clue to check its visibility from another view angle. For the textureless areas where the multi-view consistency constraint does not work well, we further introduce a depth consistency loss with confidence weights to infer the SDF. Our experimental results outperform the state-of-the-art methods under the widely used benchmarks. Project page: https://wen-yuan-zhang.github.io/NeRFPrior/. Emily Yue-ting Jia, Junsheng Zhou, Baorui Ma, Kanle Shi, Yu-Shen Liu, Zhizhong Han |
CVPR | 4 |
| 2025 | RelaI2P: Relational Learning for Image-to-Point Cloud RegistrationabstractCross-modality registration between 2D images and 3D point clouds is an important task in autonomous driving and robotics. Existing methods predict the correspondence between images and point clouds by matching patterns of pixel and point features learned by deep neural networks. However, due to the significant differences in their representation and feature processing, the feature spaces are vastly different. The insufficient feature interaction between image and point cloud branches leads to a lack of information necessary for relational reasoning, which is crucial for establishing the correspondence between pixels and points. To address these problems, we propose a Cross-Modality Relation Module (CMRM) that leverages the relations between different levels of features from images and point clouds. This module facilitates rich information exchange between the feature extractor branches corresponding to the two modalities. Additionally, we introduce a Relation-Aware Fusion Module (RAFM) that effectively integrates multimodal features and their relations. The experimental results on KITTI dataset show improvements over the state-of-the-art methods. The code will be publicly available at https://github.com/JLUrob/RelaI2P. Minghui Hou, Gang Wang 0013, Baorui Ma |
ICASSP | 4 |
| 2025 | BLCC: A Benchmark for Multi-LiDAR and Multi-camera Calibration
Minghui Hou, Gang Wang 0013, Tongzhou Zhang 0001, Baorui Ma |
MMM (1) | 5 |
| 2025 | PolarBEVU: Multi-Camera 3D Object Detection in Polar Bird's-Eye View via Unprojectionabstract3D object detection from a Bird’s Eye View (BEV) has emerged as a novel perception paradigm for autonomous driving scenarios. While most current 3D object detection methods still rely on the conventional Cartesian coordinates, they fail to align with the non-aligned coordinate system inherent in image geometry. The Polar coordinates, on the other hand, better fit with the geometric shape corresponding to the perception of cameras. However, transforming between coordinate systems introduces distortions in the perception information, resulting in issues such as “Weak Adaptability to Heatmap Distribution” and “Offset in the Center Point of the Bounding Box.” To address these challenges, this paper proposes a cutting-edge 3D object detection model named PolarBEVU, which leverages the bird’s-eye view under the Polar coordinates along with multi-camera unprojection. The model introduces an innovative “Deformable Uniform Heatmap Distribution” method that adjusts heatmap computations based on box shapes, generating high-quality heatmaps and effectively resolving the issue of “Weak Adaptability to Heatmap Distribution.” Moreover, the model incorporates the concept of “Dynamic High-risk Regression Region” to enhance the accuracy and robustness of the center point regression at the bounding box, thus mitigating the issue of “Offset in the Center Point of the Bounding Box.” In extensive experiments on the nuScenes dataset, PolarBEVU achieves impressive results with 49.9% mAP and 57.4% NDS on the test set, surpassing other comparative approaches and reaching the state-of-the-art (SOTA) performance among methods utilizing Polar coordinates. This clearly demonstrates the efficacy and superiority of PolarBEVU. In addition, the model is successfully deployed on Nvidia Jetson AGX Orin, showcasing real-time inference speeds of 31.42ms. These findings affirm PolarBEVU’s potential for practical applications. Code is available athttps://github.com/JLUrob/PolarBEVU. Minghui Hou, Chuanhao Lyu, Gang Wang 0013, Baorui Ma, Rongtao Xu, Jue Hu, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning Continuous Implicit Field with Local Distance Indicator for Arbitrary-Scale Point Cloud UpsamplingabstractPoint cloud upsampling aims to generate dense and uniformly distributed point sets from a sparse point cloud, which plays a critical role in 3D computer vision. Previous methods typically split a sparse point cloud into several local patches, upsample patch points, and merge all upsampled patches. However, these methods often produce holes, outliers or non-uniformity due to the splitting and merging process which does not maintain consistency among local patches.To address these issues, we propose a novel approach that learns an unsigned distance field guided by local priors for point cloud upsampling. Specifically, we train a local distance indicator (LDI) that predicts the unsigned distance from a query point to a local implicit surface. Utilizing the learned LDI, we learn an unsigned distance field to represent the sparse point cloud with patch consistency. At inference time, we randomly sample queries around the sparse point cloud, and project these query points onto the zero-level set of the learned implicit field to generate a dense point cloud. We justify that the implicit field is naturally continuous, which inherently enables the application of arbitrary-scale upsampling without necessarily retraining for various scales. We conduct comprehensive experiments on both synthetic data and real scans, and report state-of-the-art results under widely used benchmarks. Project page: https://lisj575.github.io/APU-LDI Shujuan Li, Junsheng Zhou, Baorui Ma, Yu-Shen Liu, Zhizhong Han |
AAAI | 3 |
| 2024 | UDiFF: Generating Conditional Unsigned Distance Fields with Optimal Wavelet DiffusionabstractDiffusion models have shown remarkable results for im-age generation, editing and inpainting. Recent works ex-plore diffusion models for 3D shape generation with neural implicit functions, i.e., signed distance function and occu-pancy function. However, they are limited to shapes with closed surfaces, which prevents them from generating di-verse 3D real-world contents containing open surfaces. In this work, we present UDiFF, a 3D diffusion model for unsigned distance fields (UDFs) which is capable to gener-ate textured 3D shapes with open surfaces from text conditions or unconditionally. Our key idea is to generate UDFs in spatial-frequency domain with an optimal wavelet trans-formation, which produces a compact representation space for UDF generation. Specifically, instead of selecting an appropriate wavelet transformation which requires expen-sive manual efforts and still leads to large information loss, we propose a data-driven approach to learn the optimal wavelet transformation for UDFs. We evaluate UDiFF to show our advantages by numerical and visual comparisons with the latest methods on widely used benchmarks. Page: https://weiqi-zhang.github.io/UDiFF. Junsheng Zhou, Baorui Ma, Kanle Shi, Yu-Shen Liu, Zhizhong Han |
CVPR | 3 |
| 2024 | Uni3D: Exploring Unified 3D Representation at ScaleabstractScaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively unexplored. In this work, we present Uni3D, a 3D foundation model to explore the unified 3D representation at scale. Uni3D uses a 2D initialized ViT end-to-end pretrained to align the 3D point cloud features with the image-text aligned features. Via the simple architecture and pretext task, Uni3D can leverage abundant 2D pretrained models as initialization and image-text aligned models as the target, unlocking the great potential of 2D model zoos and scaling-up strategies to the 3D world. We efficiently scale up Uni3D to one billion parameters, and set new records on a broad range of 3D tasks, such as zero-shot classification, few-shot classification, open-world understanding and zero-shot part segmentation. We show that the strong Uni3D representation also enables applications such as 3D painting and retrieval in the wild. We believe that Uni3D provides a new direction for exploring both scaling up and efficiency of the representation in 3D domain. Junsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu, Tiejun Huang 0001 |
ICLR | 3 |
| 2024 | 3D-OAE: Occlusion Auto-Encoders for Self-Supervised Learning on Point CloudsabstractThe manual annotation for large-scale point clouds is still tedious and unavailable for many harsh real-world tasks. Self-supervised learning, which is used on raw and unlabeled data to pre-train deep neural networks, is a promising approach to address this issue. Existing works usually take the common aid from auto-encoders to establish the self-supervision by the self-reconstruction schema. However, the previous auto-encoders merely focus on the global shapes and do not distinguish the local and global geometric features apart. To address this problem, we present a novel and efficient self-supervised point cloud representation learning framework, named 3D Occlusion Auto-Encoder (3D-OAE), to facilitate the detailed supervision inherited in local regions and global shapes. We propose to randomly occlude some local patches of point clouds and establish the supervision via inpainting the occluded patches using the remaining ones. Specifically, we design an asymmetrical encoder-decoder architecture based on standard Transformer, where the encoder operates only on the visible subset of patches to learn local patterns, and a lightweight decoder is designed to leverage these visible patterns to infer the missing geometries via self-attention. We find that occluding a very high proportion of the input point cloud (e.g. 75%) will still yield a nontrivial self-supervisory performance, which enables us to achieve 3-4 times faster during training but also improve accuracy. Experimental results show that our approach outperforms the state-of-the-art on a diverse range of down-stream discriminative and generative tasks. Code is available at https://github.com/junshengzhou/3D-OAE. Junsheng Zhou, Xin Wen 0003, Baorui Ma, Yu-Shen Liu, Yue Gao 0002, Yi Fang 0006, Zhizhong Han |
ICRA | 3 |
| 2024 | Inferring 3D Occupancy Fields through Implicit Reasoning on Silhouette Images
Baorui Ma, Yu-Shen Liu, Matthias Zwicker, Zhizhong Han |
ACM Multimedia | 1 |
| 2024 | Fast Learning of Signed Distance Functions From Noisy Point Clouds via Noise to Noise MappingabstractLearning signed distance functions (SDFs) from point clouds is an important task in 3D computer vision. However, without ground truth signed distances, point normals or clean point clouds, current methods still struggle from learning SDFs from noisy point clouds. To overcome this challenge, we propose to learn SDFs via a noise to noise mapping, which does not require any clean point cloud or ground truth supervision. Our novelty lies in the noise to noise mapping which can infer a highly accurate SDF of a single object or scene from its multiple or even single noisy observations. We achieve this by a novel loss which enables statistical reasoning on point clouds and maintains geometric consistency although point clouds are irregular, unordered and have no point correspondence among noisy observations. To accelerate training, we use multi-resolution hash encodings implemented in CUDA in our framework, which reduces our training time by a factor of ten, achieving convergence within one minute. We further introduce a novel schema to improve multi-view reconstruction by estimating SDFs as a prior. Our evaluations under widely-used benchmarks demonstrate our superiority over the state-of-the-art methods in surface reconstruction from point clouds or multi-view images, point cloud denoising and upsampling. Junsheng Zhou, Baorui Ma, Yu-Shen Liu, Zhizhong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | CAP-UDF: Learning Unsigned Distance Functions Progressively From Raw Point Clouds With Consistency-Aware Field OptimizationabstractSurface reconstruction for point clouds is an important task in 3D computer vision. Most of the latest methods resolve this problem by learning signed distance functions from point clouds, which are limited to reconstructing closed surfaces. Some other methods tried to represent open surfaces using unsigned distance functions (UDF) which are learned from ground truth distances. However, the learned UDF is hard to provide smooth distance fields due to the discontinuous character of point clouds. In this paper, we propose CAP-UDF, a novel method to learn consistency-aware UDF from raw point clouds. We achieve this by learning to move queries onto the surface with a field consistency constraint, where we also enable to progressively estimate a more accurate surface. Specifically, we train a neural network to gradually infer the relationship between queries and the approximated surface by searching for the moving target of queries in a dynamic way. Meanwhile, we introduce a polygonization algorithm to extract surfaces using the gradients of the learned UDF. We conduct comprehensive experiments in surface reconstruction for point clouds, real scans or depth maps, and further explore our performance in unsupervised point normal estimation, which demonstrate non-trivial improvements of CAP-UDF over the state-of-the-art methods. Junsheng Zhou, Baorui Ma, Shujuan Li, Yu-Shen Liu, Yi Fang 0006, Zhizhong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | NeAF: Learning Neural Angle Fields for Point Normal EstimationabstractNormal estimation for unstructured point clouds is an important task in 3D computer vision. Current methods achieve encouraging results by mapping local patches to normal vectors or learning local surface fitting using neural networks. However, these methods are not generalized well to unseen scenarios and are sensitive to parameter settings. To resolve these issues, we propose an implicit function to learn an angle field around the normal of each point in the spherical coordinate system, which is dubbed as Neural Angle Fields (NeAF). Instead of directly predicting the normal of an input point, we predict the angle offset between the ground truth normal and a randomly sampled query normal. This strategy pushes the network to observe more diverse samples, which leads to higher prediction accuracy in a more robust manner. To predict normals from the learned angle fields at inference time, we randomly sample query vectors in a unit spherical space and take the vectors with minimal angle values as the predicted normals. To further leverage the prior learned by NeAF, we propose to refine the predicted normal vectors by minimizing the angle offsets. The experimental results with synthetic data and real scans show significant improvements over the state-of-the-art under widely used benchmarks. Project page: https://lisj575.github.io/NeAF/. Shujuan Li, Junsheng Zhou, Baorui Ma, Yu-Shen Liu, Zhizhong Han |
AAAI | 3 |
| 2023 | Towards Better Gradient Consistency for Neural Signed Distance Functions via Level Set AlignmentabstractNeural signed distance functions (SDFs) have shown remarkable capability in representing geometry with details. However, without signed distance supervision, it is still a challenge to infer SDFs from point clouds or multi-view images using neural networks. In this paper, we claim that gradient consistency in the field, indicated by the parallelism of level sets, is the key factor affecting the inference accuracy. Hence, we propose a level set alignment loss to evaluate the parallelism of level sets, which can be minimized to achieve better gradient consistency. Our novelty lies in that we can align all level sets to the zero level set by constraining gradients at queries and their projections on the zero level set in an adaptive way. Our insight is to propagate the zero level set to everywhere in the field through consistent gradients to eliminate uncertainty in the field that is caused by the discreteness of 3D point clouds or the lack of observations from multi-view images. Our proposed loss is a general term which can be used upon different methods to infer SDFs from 3D point clouds and multi-view images. Our numerical and visual comparisons demonstrate that our loss can significantly improve the accuracy of SDFs inferred from point clouds or multiview images under various benchmarks. Code and data are available at https://github.com/mabaorui/TowardsBetterGradient. Baorui Ma, Junsheng Zhou, Yu-Shen Liu, Zhizhong Han |
CVPR | 1 |
| 2023 | Learning a More Continuous Zero Level Set in Unsigned Distance Fields through Level Set ProjectionabstractLatest methods represent shapes with open surfaces using unsigned distance functions (UDFs). They train neural networks to learn UDFs and reconstruct surfaces with the gradients around the zero level set of the UDF. However, the differential networks struggle from learning the zero level set where the UDF is not differentiable, which leads to large errors on unsigned distances and gradients around the zero level set, resulting in highly fragmented and discontinuous surfaces. To resolve this problem, we propose to learn a more continuous zero level set in UDFs with level set projections. Our insight is to guide the learning of zero level set using the rest non-zero level sets via a projection procedure. Our idea is inspired from the observations that the non-zero level sets are much smoother and more continuous than the zero level set. We pull the non-zero level sets onto the zero level set with gradient constraints which align gradients over different level sets and correct unsigned distance errors on the zero level set, leading to a smoother and more continuous unsigned distance field. We conduct comprehensive experiments in surface reconstruction for point clouds, real scans or depth maps, and further explore the performance in unsupervised point cloud upsampling and unsupervised point normal estimation with the learned UDF, which demonstrate our non-trivial improvements over the state-of-the-art methods. Code is available at https://github.com/junshengzhou/LevelSetUDF. Junsheng Zhou, Baorui Ma, Shujuan Li, Yu-Shen Liu, Zhizhong Han |
ICCV | 2 |
| 2023 | Learning Signed Distance Functions from Noisy 3D Point Clouds via Noise to Noise MappingabstractLearning signed distance functions (SDFs) from 3D point clouds is an important task in 3D computer vision. However, without ground truth signed distances, point normals or clean point clouds, current methods still struggle from learning SDFs from noisy point clouds. To overcome this challenge, we propose to learn SDFs via a noise to noise mapping, which does not require any clean point cloud or ground truth supervision for training. Our novelty lies in the noise to noise mapping which can infer a highly accurate SDF of a single object or scene from its multiple or even single noisy point cloud observations. Our novel learning manner is supported by modern Lidar systems which capture multiple noisy observations per second. We achieve this by a novel loss which enables statistical reasoning on point clouds and maintains geometric consistency although point clouds are irregular, unordered and have no point correspondence among noisy observations. Our evaluation under the widely used benchmarks demonstrates our superiority over the state-of-the-art methods in surface reconstruction, point cloud denoising and upsampling. Our code, data, and pre-trained models are available at https://github.com/mabaorui/Noise2NoiseMapping/ . Baorui Ma, Yu-Shen Liu, Zhizhong Han |
ICML | 1 |
| 2023 | Differentiable Registration of Images and LiDAR Point Clouds with VoxelPoint-to-Pixel MatchingabstractCross-modality registration between 2D images captured by cameras and 3D point clouds from LiDARs is a crucial task in computer vision and robotic. Previous methods estimate 2D-3D correspondences by matching point and pixel patterns learned by neural networks, and use Perspective-n-Points (PnP) to estimate rigid transformation during post-processing. However, these methods struggle to map points and pixels to a shared latent space robustly since points and pixels have very different characteristics with patterns learned in different manners (MLP and CNN), and they also fail to construct supervision directly on the transformation since the PnP is non-differentiable, which leads to unstable registration results. To address these problems, we propose to learn a structured cross-modality latent space to represent pixel features and 3D features via a differentiable probabilistic PnP solver. Specifically, we design a triplet network to learn VoxelPoint-to-Pixel matching, where we represent 3D elements using both voxels and points to learn the cross-modality latent space with pixels. We design both the voxel and pixel branch based on CNNs to operate convolutions on voxels/pixels represented in grids, and integrate an additional point branch to regain the information lost during voxelization. We train our framework end-to-end by imposing supervisions directly on the predicted pose distribution with a probabilistic PnP solver. To explore distinctive patterns of cross-modality features, we design a novel loss with adaptive-weighted optimization for cross-modality feature description. The experimental results on KITTI and nuScenes datasets show significant improvements over the state-of-the-art methods. Junsheng Zhou, Baorui Ma, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han |
NeurIPS | 2 |
| 2022 | Reconstructing Surfaces for Sparse Point Clouds with On-Surface PriorsabstractIt is an important task to reconstruct surfaces from 3D point clouds. Current methods are able to reconstruct surfaces by learning Signed Distance Functions (SDFs) from single point clouds without ground truth signed distances or point normals. However, they require the point clouds to be dense, which dramatically limits their performance in real applications. To resolve this issue, we propose to reconstruct highly accurate surfaces from sparse point clouds with an on-surface prior. We train a neural network to learn SDFs via projecting queries onto the surface represented by the sparse point cloud. Our key idea is to infer signed distances by pushing both the query projections to be on the surface and the projection distance to be the minimum. To achieve this, we train a neural network to capture the on-surface prior to determine whether a point is on a sparse point cloud or not, and then leverage it as a differentiable function to learn SDFs from unseen sparse point cloud. Our method can learn SDFs from a single s parse point cloud without ground truth signed distances or point normals. Our numerical evaluation under widely used benchmarks demonstrates that our method achieves state-of-the-art reconstruction accuracy, especially for sparse point clouds. Code and data are available at https://github.com/mabaorui/OnSurfacePrior. Baorui Ma, Yu-Shen Liu, Zhizhong Han |
CVPR | 1 |
| 2022 | Surface Reconstruction from Point Clouds by Learning Predictive Context PriorsabstractSurface reconstruction from point clouds is vital for 3D computer vision. State-of-the-art methods leverage large datasets to first learn local context priors that are represented as neural network-based signed distance functions (SDFs) with some parameters encoding the local contexts. To reconstruct a surface at a specific query location at inference time, these methods then match the local reconstruction target by searching for the best match in the local prior space (by optimizing the parameters encoding the local context) at the given query location. However, this requires the local context prior to generalize to a wide variety of unseen target regions, which is hard to achieve. To resolve this issue, we introduce Predictive Context Priors by learning Predictive Queries for each specific point cloud at inference time. Specifically, we first train a local context prior using a large point cloud dataset similar to previous techniques. For surface reconstruction at inference time, however, we specialize the local context prior into our Predictive Context Prior by learning Predictive Queries, which predict adjusted spatial query locations as displacements of the original locations. This leads to a global SDF that fits the specific point cloud the best. Intuitively, the query prediction enables us to flexibly search the learned local context prior over the entire prior space, rather than being restricted to the fixed query locations, and this improves the generalizability. Our method does not require ground truth signed distances, normals, or any additional procedure of signed distance fusion across overlapping regions. Our experimental results in surface reconstruction for single shapes or complex scenes show significant improvements over the state-of-the-art under widely used benchmark-s. Code and data are available at https://github.com/mabaorui/PredictableContextPrior. Baorui Ma, Yu-Shen Liu, Matthias Zwicker, Zhizhong Han |
CVPR | 1 |
| 2022 | Learning Consistency-Aware Unsigned Distance Functions Progressively from Raw Point CloudsabstractSurface reconstruction for point clouds is an important task in 3D computer vision. Most of the latest methods resolve this problem by learning signed distance functions (SDF) from point clouds, which are limited to reconstructing shapes or scenes with closed surfaces. Some other methods tried to represent shapes or scenes with open surfaces using unsigned distance functions (UDF) which are learned from large scale ground truth unsigned distances. However, the learned UDF is hard to provide smooth distance fields near the surface due to the noncontinuous character of point clouds. In this paper, we propose a novel method to learn consistency-aware unsigned distance functions directly from raw point clouds. We achieve this by learning to move 3D queries to reach the surface with a field consistency constraint, where we also enable to progressively estimate a more accurate surface. Specifically, we train a neural network to gradually infer the relationship between 3D queries and the approximated surface by searching for the moving target of queries in a dynamic way, which results in a consistent field around the surface. Meanwhile, we introduce a polygonization algorithm to extract surfaces directly from the gradient field of the learned UDF. The experimental results in surface reconstruction for synthetic and real scan data show significant improvements over the state-of-the-art under the widely used benchmarks. Junsheng Zhou, Baorui Ma, Yu-Shen Liu, Yi Fang 0006, Zhizhong Han |
NeurIPS | 2 |
| 2021 | Neural-Pull: Learning Signed Distance Function from Point clouds by Learning to Pull Space onto SurfaceabstractReconstructing continuous surfaces from 3D point clouds is a fundamental operation in 3D geometry processing. Several recent state-of-the-art methods address this problem using neural networks to learn signed distance functions (SDFs). In this paper, we introduce Neural-Pull, a new approach that is simple and leads to high quality SDFs. Specifically, we train a neural network to pull query 3D locations to their closest points on the surface using the predicted signed distance values and the gradient at the query locations, both of which are computed by the network itself. The pulling operation moves each query location with a stride given by the distance predicted by the network. Based on the sign of the distance, this may move the query location along or against the direction of the gradient of the SDF. This is a differentiable operation that allows us to update the signed distance value and the gradient simultaneously during training. Our outperforming results under widely used benchmarks demonstrate that we can learn SDFs more accurately and flexibly for surface reconstruction and single image reconstruction than the state-of-the-art methods. Our code and data are available at https://github.com/mabaorui/NeuralPull. Baorui Ma, Zhizhong Han, Yu-Shen Liu, Matthias Zwicker |
ICML | 1 |
| 2020 | Reconstructing 3D Shapes From Multiple Sketches Using Direct Shape Optimizationabstract3D shape reconstruction from multiple hand-drawn sketches is an intriguing way to 3D shape modeling. Currently, state-of-the-art methods employ neural networks to learn a mapping from multiple sketches from arbitrary view angles to a 3D voxel grid. Because of the cubic complexity of 3D voxel grids, however, neural networks are hard to train and limited to low resolution reconstructions, which leads to a lack of geometric detail and low accuracy. To resolve this issue, we propose to reconstruct 3D shapes from multiple sketches using direct shape optimization (DSO), which does not involve deep learning models for direct voxel-based 3D shape generation. Specifically, we first leverage a conditional generative adversarial network (CGAN) to translate each sketch into an attenuance image that captures the predicted geometry from a given viewpoint. Then, DSO minimizes a project-and-compare loss to reconstruct the 3D shape such that it matches the predicted attenuance images from the view angles of all input sketches. Based on this, we further propose a progressive update approach to handle inconsistencies among a few hand-drawn sketches for the same 3D shape. Our experimental results show that our method significantly outperforms the state-of-the-art methods under widely used benchmarks and produces intuitive results in an interactive application. Zhizhong Han, Baorui Ma, Yu-Shen Liu, Matthias Zwicker |
IEEE Trans. Image Process. | 2 |