Honghua Chen

dblp:53/5080 · DBLP profile ↗
← Back
47ranked-venue papers
7as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 26 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion
abstract
Existing diffusion-based 3D shape completion methods typically use a conditional paradigm, injecting incomplete shape information into the denoising network via deep feature interactions (e.g., concatenation, cross-attention) to guide sampling toward complete shapes, often represented by voxel-based distance functions. However, these approaches fail to explicitly model the optimal global transport path, leading to suboptimal completions. Moreover, performing diffusion directly in voxel space imposes resolution constraints, limiting the generation of fine-grained geometric details. To address these challenges, we propose BridgeShape, a novel framework for 3D shape completion via latent diffusion Schrödinger bridge. The key innovations lie in two aspects: (i) BridgeShape formulates shape completion as an optimal transport problem, explicitly modeling the transition between incomplete and complete shapes to ensure a globally coherent transformation. (ii) We introduce a Depth-Enhanced Vector Quantized Variational Autoencoder (VQ-VAE) to encode 3D shapes into a compact latent space, leveraging self-projected multi-view depth information enriched with strong DINOv2 features to enhance geometric structural perception. By operating in a compact yet structurally informative latent space, BridgeShape effectively mitigates resolution constraints and enables more efficient and high-fidelity 3D shape completion. BridgeShape achieves state-of-the-art performance on 3D shape completion benchmarks, demonstrating superior fidelity at higher resolutions and for unseen object classes.
Dequan Kong, Honghua Chen, Zhe Zhu, Mingqiang Wei
AAAI2
2026 Fusing shape descriptors and geometric details for robust category-level object pose estimation
Yun Liu 0002, Weiming Wang 0002, Fu Lee Wang, Haoran Xie 0001, Honghua Chen, Xue Xue, Mingqiang Wei, Harry Qin
Multim. Tools Appl.5
2026 CrossTracker: Robust Multi-Modal 3D Multi-Object Tracking via Cross Correction
abstract
Inaccurate detections remain a critical bottleneck in 3D multi-object tracking (MOT). Recent detection fusion-based methods incorporate camera detections as supplementary to reduce false detections and compensate for missing ones in LiDAR. However, their unidirectional camera-LiDAR correction lacks a feedback mechanism, precluding iterative mutual refinement between modalities for more robust LiDAR-based tracking. Inspired by the coarse-to-fine strategy in two-stage object detection, we introduceCrossTracker, a novel two-stage framework for online multi-modal 3D MOT. CrossTracker first constructs coarse camera and LiDAR trajectories independently, then performs trajectory fusion using both current and historical frames, without requiring future data. This ensures more robust mutual refinement between modalities. Specifically, CrossTracker comprises three core modules: i) the multi-modal modeling (M3) module, which fuses data from images, point clouds, and even planar geometry derived from images to establish a robust tracking constraint; ii) the coarse trajectory generation (C-TG) module, which independently generates coarse trajectories for both modalities using the M3constraint; and iii) the trajectory fusion (TF) module, which applies mutual refinement between coarse LiDAR and camera trajectories through cross correction to ensure robust LiDAR trajectories. Extensive experiments show that CrossTracker outperforms 19 state-of-the-art methods, highlighting its effectiveness in leveraging the synergistic strengths of camera and LiDAR sensors for robust multi-modal 3D MOT. The code is available at https://github.com/lipeng-gu/CrossTracker.
Lipeng Gu, Xuefeng Yan 0001, Weiming Wang 0002, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei
IEEE Trans. Circuits Syst. Video Technol.4
2026 CoreEditor: Correspondence-Constrained Diffusion for Consistent 3D Editing
abstract
Text-driven 3D editing is an emerging task that focuses on modifying scenes based on text prompts. Current methods often adapt pre-trained 2D image editors to multi-view observations, using specific strategies to combine information across views. However, these approaches still struggle with ensuring consistency across views, as they lack precise control over the sharing of information, resulting in edits with insufficient visual changes and blurry details. In this paper, we propose CoreEditor, a novel framework for consistent text-to-3D editing. At the core of our approach is a novel correspondence-constrained attention mechanism, which enforces structured interactions between corresponding pixels that are expected to remain visually consistent during the diffusion denoising process. Unlike conventional wisdom that relies solely on scene geometry, we enhance the correspondence by incorporating semantic similarity derived from the diffusion denoising process. This combined support from both geometry and semantics ensures a robust multi-view editing process. Additionally, we introduce a selective editing pipeline that enables users to choose their preferred edits from multiple candidates, creating a more flexible and user-centered 3D editing process. Extensive experiments demonstrate the effectiveness of CoreEditor, showing its ability to generate high-quality 3D edits, significantly outperforming existing methods.
Zhe Zhu, Honghua Chen, Peng Li 0064, Mingqiang Wei
IEEE Trans. Vis. Comput. Graph.2
2025 CosCAD: Cross-Modal CAD Model Retrieval and Pose Alignment from a Single Image
Zhikun Wen, Honghua Chen, Zhe Zhu, Zeyong Wei, Liangliang Nan, Mingqiang Wei
CVM (2)2
2025 STAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point Clouds
abstract
Extracting geometric edges from unstructured point clouds remains a significant challenge, particularly in thin-walled structures that are commonly found in everyday objects. Traditional geometric methods and recent learning-based approaches frequently struggle with these structures, as both rely heavily on sufficient contextual information from local point neighborhoods. However, 3D measurement data of thin-walled structures often lack the accurate, dense, and regular neighborhood sampling required for reliable edge extraction, resulting in degraded performance.In this work, we introduce STAR-Edge, a novel approach designed for detecting and refining edge points in thin-walled structures. Our method leverages a unique representation—the local spherical curve—to create structure-aware neighborhoods that emphasize co-planar points while reducing interference from close-by, non-co-planar surfaces. This representation is transformed into a rotation-invariant descriptor, which, combined with a lightweight multi-layer perceptron, enables robust edge point classification even in the presence of noise and sparse or irregular sampling. Besides, we also use the local spherical curve representation to estimate more precise normals and introduce an optimization function to project initially identified edge points exactly on the true edges. Experiments conducted on the ABC dataset and thin-walled structure-specific datasets demonstrate that STAR-Edge outperforms existing edge detection methods, showcasing better robustness under various challenging conditions. The source code is available at https://github.com/miraclelzk/star-edge.
Zikuan Li, Honghua Chen, Yuecheng Wang, Sibo Wu, Mingqiang Wei, Jun Wang 0039
CVPR2
2025 Textured 3D Regenerative Morphing with 3D Diffusion Prior
abstract
Textured 3D morphing creates smooth and plausible interpolation sequences between two 3D objects, focusing on transitions in both shape and texture. This is important for creative applications like visual effects in filmmaking. Previous methods rely on establishing point-to-point correspondences and determining smooth deformation trajectories, which inherently restrict them to shape-only morphing on untextured, topologically aligned datasets. This restriction leads to labor-intensive preprocessing and poor generalization. To overcome these challenges, we propose a method for 3D regenerative morphing using a 3D diffusion prior. Unlike previous methods that depend on explicit correspondences and deformations, our method eliminates the additional need for obtaining correspondence and uses the 3D diffusion prior to generate morphing. Specifically, we introduce a 3D diffusion model and interpolate the source and target information at three levels: initial noise, model parameters, and condition features. We then explore an Attention Fusion strategy to generate more smooth morphing sequences. To further improve the plausibility of semantic interpolation and the generated 3D surfaces, we propose two strategies: (a) Token Reordering, where we match approximate tokens based on semantic analysis to guide implicit correspondences in the denoising process of the diffusion model, and (b) Low-Frequency Enhancement, where we enhance low-frequency signals in the tokens to improve the quality of generated surfaces. Experimental results show that our method achieves superior smoothness and plausibility in 3D morphing across diverse cross-category object pairs, offering a novel regenerative method for 3D morphing with textured representations.
Yushi Lan, Honghua Chen, Xingang Pan
ICCV3
2025 RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot Learning
Honghua Chen, Mingqiang Wei
ICCV3
2025 ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
abstract
We propose ArtiLatent, a generative framework that synthesizes human-made 3D objects with fine-grained geometry, accurate articulation, and realistic appearance. Our approach jointly models part geometry and articulation dynamics by embedding sparse voxel representations and associated articulation properties—including joint type, axis, origin, range, and part category—into a unified latent space via a variational autoencoder. A latent diffusion model is then trained over this space to enable diverse yet physically plausible sampling. To reconstruct photorealistic 3D shapes, we introduce an articulation-aware Gaussian decoder that accounts for articulation-dependent visibility changes (e.g., revealing the interior of a drawer when opened). By conditioning appearance decoding on articulation state, our method assigns plausible texture features to regions that are typically occluded in static poses, significantly improving visual realism across articulation configurations. Extensive experiments on furniture-like objects from PartNet-Mobility and ACD datasets demonstrate that ArtiLatent outperforms existing approaches in geometric consistency and appearance fidelity. Our framework provides a scalable solution for articulated 3D object synthesis and manipulation.
Honghua Chen, Yushi Lan, Yongwei Chen, Xingang Pan
SIGGRAPH Asia1
2025 PointSea: Point Cloud Completion via Self-structure Augmentation
Zhe Zhu, Honghua Chen, Mingqiang Wei
Int. J. Comput. Vis.2
2025 Revisiting Tradition and Beyond: A Customized Bilateral Filtering Framework for Point Cloud Denoising
abstract
Deep learning-based methods have become the dominant solution for point cloud denoising, offering strong generalization capabilities through data-driven training. However, traditional methods, despite their drawbacks of heavy parameter tuning and weak generalization, retain unique advantages in interpretability and theoretical robustness. This complementarity motivates us to explore a hybrid solution that leverages data-driven paradigms to overcome the performance constraints of traditional methods. In this paper, we revisit the classic bilateral filter (BF) as a case study and identify three key limitations hindering its performance: excessive parameter tuning, suboptimal neighborhood quality, and fixed parameters across the entire model. To address them, we propose CustomBF, a novel framework for customizing BF components at a per-point level. CustomBF employs multigraph encoders and a mutual guidance strategy to analyze local patches, enabling the customization of BF components including center point normal, neighborhood point coordinates, Gaussian function parameters, and neighborhood radius for each point. Experimental results demonstrate that this component-customized bilateral filter outperforms state-of-the-art methods and achieves robust denoising even in complex scenarios. It highlights the potential of hybrid methods to extend the applicability and effectiveness of traditional techniques.
Peng Li 0064, Zeyong Wei, Honghua Chen, Xuefeng Yan 0001, Mingqiang Wei
ACM Trans. Graph.3
2025 PointCG: Self-Supervised Point Cloud Learning via Joint Completion and Generation
abstract
The core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance.
Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei
IEEE Trans. Vis. Comput. Graph.6
2024 Shape Descriptor Guided Learning for Category-Level Object Pose Estimation
Yun Liu 0002, Weiming Wang 0002, Fu Lee Wang, Haoran Xie 0001, Honghua Chen, Mingqiang Wei, Harry Qin
CGI (3)5
2024 MVIP-NeRF: Multi-View 3D Inpainting on NeRF Scenes via Diffusion Prior
abstract
Despite the emergence of successful NeRF inpainting methods built upon explicit RGB and depth 2D inpainting supervisions, these methods are inherently constrained by the capabilities of their underlying 2D inpainters. This is due to two key reasons: (i) independently inpainting constituent images results in view-inconsistent imagery, and (ii) 2D inpainters struggle to ensure high-quality geometry completion and alignment with inpainted RGB images. To overcome these limitations, we propose a novel approach called MVIP-NeRF that harnesses the potential of diffusion priors for NeRF inpainting, addressing both appearance and geometry aspects. MVIP-NeRF performs joint inpainting across multiple views to reach a consistent solution, which is achieved via an iterative optimization process based on Score Distillation Sampling (SDS). Apart from recovering the rendered RGB images, we also extract normal maps as a geometric representation and define a normal SDS loss that motivates accurate geometry inpaint- ing and alignment with the appearance. Additionally, we formulate a multi-view SDS score function to distill generative priors simultaneously from different view images, ensuring consistent visual completion when dealing with large view variations. Our experimental results show better appearance and geometry recovery than previous NeRF in painting methods.
Honghua Chen, Chen Change Loy, Xingang Pan
CVPR1
2024 PointeNet: A lightweight framework for effective and efficient point cloud analysis
Lipeng Gu, Xuefeng Yan 0001, Liangliang Nan, Dingkun Zhu, Honghua Chen, Weiming Wang 0002, Mingqiang Wei
Comput. Aided Geom. Des.5
2024 PathNet: Path-Selective Point Cloud Denoising
abstract
Current point cloud denoising (PCD) models optimize single networks, trying to make their parameters adaptive to each point in a large pool of point clouds. Such a denoising network paradigm neglects that different points are often corrupted by different levels of noise and they may convey different geometric structures. Thus, the intricacy of both noise and geometry poses side effects including remnant noise, wrongly-smoothed edges, and distorted shape after denoising. We propose PathNet, a path-selective PCD paradigm based on reinforcement learning (RL). Unlike existing efforts, PathNet enables dynamic selection of the most appropriate denoising path for each point, best moving it onto its underlying surface. We have two more contributions besides the proposed framework of path-selective PCD for the first time. First, to leverage geometry expertise and benefit from training data, we propose a noise- and geometry-aware reward function to train the routing agent in RL. Second, the routing agent and the denoising network are trained jointly to avoid under- and over-smoothing. Extensive experiments show promising improvements of PathNet over its competitors, in terms of the effectiveness for removing different levels of noise and preserving multi-scale surface geometries. Furthermore, PathNet generalizes itself more smoothly to real scans than cutting-edge models.
Zeyong Wei, Honghua Chen, Liangliang Nan, Jun Wang 0039, Harry Qin, Mingqiang Wei
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 GeoDC: Geometry-Constrained Depth Completion With Depth Distribution Modeling
abstract
Depth completion is a fundamental, yet not well-solved problem in 3-D vision. Current wisdom attempts to employ implicit geometric spatial cues from point clouds to assist in depth completion. However, these methods encounter challenges in extracting rich geometric features due to the absence of explicit constraints. In this article, we propose GeoDC, a geometry-constrained depth completion network with depth distribution modeling. GeoDC employs point cloud upsampling as an auxiliary task to guide the network in learning more robust and effective geometric features. Simultaneously, a novel image and point cloud fusion module, denoted as IP-Interaction, is implemented to holistically integrate features from images and point clouds. Besides, recognizing the presence of uncertainty and ambiguity in the ground-truth (GT) data, we construct a prior network and a posterior network to model depth feature distributions and leverage the distributions to guide depth map inference. GeoDC can solve both the problems of geometric constraint inadequacies in feature extraction and data uncertainty within depth maps well. Extensive experiments underscore the efficacy of our method, demonstrating comparable or superior performance when compared to existing state-of-the-art methods.
Peng Li 0064, Xuefeng Yan 0001, Honghua Chen, Mingqiang Wei
IEEE Trans. Geosci. Remote. Sens.4
2024 Point Transformer-Based Salient Object Detection Network for 3-D Measurement Point Clouds
abstract
While salient object detection (SOD) on 2D images has been extensively studied, there is very little SOD work on 3D measurement surfaces. We propose an effective point transformer-based SOD network for 3D measurement point clouds, termed PSOD-Net. PSOD-Net is an encoder-decoder network that takes full advantage of transformers to model the contextual information in both multi-scale point- and scene-wise manners. In the encoder, we develop a Point Context Transformer (PCT) module to capture region contextual features at the point level; PCT contains two different transformers to excavate the relationship among points. In the decoder, we develop a Scene Context Transformer (SCT) module to learn context representations at the scene level; SCT contains both Upsampling-and-Transformer blocks and Multi-context Aggregation units to integrate the global semantic and multi-level features from the encoder into the global scene context. Experiments show clear improvements of PSOD-Net over its competitors and validate that PSOD-Net is more robust to challenging cases such as small objects, multiple objects, and objects with complex structures. Code is available at: https://github.com/ZeyongWei/PSOD-Net.
Zeyong Wei, Baian Chen, Weiming Wang 0002, Honghua Chen, Mingqiang Wei, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 PN-Internet: Point-and-Normal Interactive Network for Noisy Point Clouds
abstract
Point cloud denoising and normal estimation are two fundamental yet dependent problems in digital geometry processing. However, both are often independently researched, leading to inconsistent geometry on 3D surfaces. To address it, we propose PN-Internet, an end-to-end Point-and-Normal Interactive Network for joint point cloud denoising and normal estimation. PN-Internet leverages the geometric dependency between point positions and normals to design two interactive graph convolution networks (GCNs): a point-to-normal network and a normal-to-point network. It adopts a coarse-to-fine learning paradigm, where two GCNs are exploited to respectively perform point cloud denoising and normal estimation. The point-to-normal network improves the quality of the normals using an MLP module, while the normal-to-point network refines the point positions using a parameter-free projection module based on the constraints from the normals. In addition, we introduce a feature-aware loss function to preserve the quality of 3D shape features. Unlike most existing methods, PN-Internet takes advantage of the geometric dependency between points and normals and benefits from training data. Our experimental results demonstrate that PN-Internet achieves geometric consistency between point cloud denoising and normal estimation. Furthermore, we show significant improvements over state-of-the-art methods.
Zeyong Wei, Jingbo Qiu, Honghua Chen, Jun Wang 0039, Mingqiang Wei
IEEE Trans. Geosci. Remote. Sens.4
2024 RegiFormer: Unsupervised Point Cloud Registration via Geometric Local-to-Global Transformer and Self-Augmentation
abstract
Representation learning for two partially overlapping point clouds remains an open challenge in unsupervised point cloud registration (U-PCR). In this article, we introduce RegiFormer, a geometric local-to-global transformer (GLGT)-based unsupervised framework equipped with a self-augmentation (SA) strategy, for point cloud registration. The GLGT not only aggregates features from local neighborhoods but also extracts global intrarelationships within the entire point cloud using a transformation-invariant geometry embedding. In addition, it enhances the interrelationships between paired point clouds. To overcome the limited ability of U-PCR methods to learn alignment knowledge, we design an SA strategy that can be flexibly integrated into advanced models, significantly boosting their registration performance. Extensive experiments, conducted on five popular synthetic and real-scanned benchmarks, demonstrate the superior performance of RegiFormer compared to state-of-the-art methods, both qualitatively and quantitatively.
Mengjiao Ma, Zhilei Chen, Honghua Chen, Weiming Wang 0002, Mingqiang Wei
IEEE Trans. Geosci. Remote. Sens.4
2024 Geometric and Learning-Based Mesh Denoising: A Comprehensive Survey
abstract
Mesh denoising is a fundamental problem in digital geometry processing. It seeks to remove surface noise while preserving surface intrinsic signals as accurately as possible. While traditional wisdom has been built upon specialized priors to smooth surfaces, learning-based approaches are making their debut with great success in generalization and automation. In this work, we provide a comprehensive review of the advances in mesh denoising, containing both traditional geometric approaches and recent learning-based methods. First, to familiarize readers with the denoising tasks, we summarize four common issues in mesh denoising. We then provide two categorizations of the existing denoising methods. Furthermore, three important categories, including optimization-, filter-, and data-driven-based techniques, are introduced and analyzed in detail, respectively. Both qualitative and quantitative comparisons are illustrated, to demonstrate the effectiveness of the state-of-the-art denoising methods. Finally, potential directions of future work are pointed out to solve the common problems of these approaches. A mesh denoising benchmark is also built in this work, and future researchers will easily and conveniently evaluate their methods with state-of-the-art approaches. To aid reproducibility, we release our datasets and used results at https://github.com/chenhonghua/Mesh-Denoiser .
Honghua Chen, Zhiqi Li 0002, Mingqiang Wei, Jun Wang 0039
ACM Trans. Multim. Comput. Commun. Appl.1
2024 CSDN: Cross-Modal Shape-Transfer Dual-Refinement Network for Point Cloud Completion
abstract
How will you repair a physical object with some missings? You may imagine its original shape from previously captured images, recover its overall (global) but coarse shape first, and then refine its local details. We are motivated to imitate the physical repair procedure to address point cloud completion. To this end, we propose a cross-modal shape-transfer dual-refinement network (termed CSDN), a coarse-to-fine paradigm with images of full-cycle participation, for quality point cloud completion. CSDN mainly consists of "shape fusion" and "dual-refinement" modules to tackle the cross-modal challenge. The first module transfers the intrinsic shape characteristics from single images to guide the geometry generation of the missing regions of point clouds, in which we propose IPAdaIN to embed the global features of both the image and the partial point cloud into completion. The second module refines the coarse output by adjusting the positions of the generated points, where the local refinement unit exploits the geometric relation between the novel and the input points by graph convolution, and the global constraint unit utilizes the input image to fine-tune the generated offset. Different from most existing approaches, CSDN not only explores the complementary information from images but also effectively exploits cross-modal data in the whole coarse-to-fine completion procedure. Experimental results indicate that CSDN performs favorably against twelve competitors on the cross-modal benchmark.
Zhe Zhu, Liangliang Nan, Haoran Xie 0001, Honghua Chen, Jun Wang 0039, Mingqiang Wei, Harry Qin
IEEE Trans. Vis. Comput. Graph.4
2024 GeoSegNet: point cloud semantic segmentation via geometric encoder-decoder modeling
Chen Chen 0161, Yisen Wang 0003, Honghua Chen, Xuefeng Yan 0001, Dayong Ren, Yanwen Guo 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei
Vis. Comput.3
2023 Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising
abstract
We propose GeoGCN, a novel geometric dual-domain graph convolution network for point cloud denoising (PCD). Beyond the traditional wisdom of PCD, to fully exploit the geometric information of point clouds, we define two kinds of surface normals, one is called Real Normal (RN), and the other is Virtual Normal (VN). RN preserves the local details of noisy point clouds while VN avoids the global shape shrinkage during denoising. GeoGCN is a new PCD paradigm that, 1) first regresses point positions by spatial-based GCN with the help of VNs, 2) then estimates initial RNs by performing Principal Component Analysis on the regressed points, and 3) finally regresses fine RNs by normal-based GCN. Unlike existing PCD methods, GeoGCN not only exploits two kinds of geometry expertise (i.e., RN and VN) but also benefits from training data. Experiments validate that GeoGCN outperforms SOTAs in terms of both noise-robustness and local-and-global feature preservation.
Zhaowei Chen, Peng Li 0064, Zeyong Wei, Honghua Chen, Haoran Xie 0001, Mingqiang Wei, Fu Lee Wang
ICASSP4
2023 SVDFormer: Complementing Point Cloud via Self-view Augmentation and Self-structure Dual-generator
abstract
In this paper, we propose a novel network, SVDFormer, to tackle two specific challenges in point cloud completion: understanding faithful global shapes from incomplete point clouds and generating high-accuracy local structures. Current methods either perceive shape patterns using only 3D coordinates or import extra images with well-calibrated intrinsic parameters to guide the geometry estimation of the missing parts. However, these approaches do not always fully leverage the cross-modal self-structures available for accurate and high-quality point cloud completion. To this end, we first design a Self-view Fusion Network that leverages multiple-view depth image information to observe incomplete self-shape and generate a compact global shape. To reveal highly detailed structures, we then introduce a refinement module, called Self-structure Dual-generator, in which we incorporate learned shape priors and geometric self-similarities for producing new points. By perceiving the incompleteness of each point, the dual-path design disentangles refinement strategies conditioned on the structural type of each point. SVDFormer absorbs the wisdom of self-structures, avoiding any additional paired information such as color images with precisely calibrated camera intrinsic parameters. Comprehensive experiments indicate that our method achieves state-of-the-art performance on widely-used benchmarks. Code is available at https://github.com/czvvd/SVDFormer.
Zhe Zhu, Honghua Chen, Weiming Wang 0002, Harry Qin, Mingqiang Wei
ICCV2
2023 Deep Learning for Scene Flow Estimation on Point Clouds: A Survey and Prospective Trends
abstract
Abstract Aiming at obtaining structural information and 3D motion of dynamic scenes, scene flow estimation has been an interest of research in computer vision and computer graphics for a long time. It is also a fundamental task for various applications such as autonomous driving. Compared to previous methods that utilize image representations, many recent researches build upon the power of deep analysis and focus on point clouds representation to conduct 3D flow estimation. This paper comprehensively reviews the pioneering literature in scene flow estimation based on point clouds. Meanwhile, it delves into detail in learning paradigms and presents insightful comparisons between the state‐of‐the‐art methods using deep learning for scene flow estimation. Furthermore, this paper investigates various higher‐level scene understanding tasks, including object tracking, motion segmentation, etc. and concludes with an overview of foreseeable research trends for scene flow estimation.
Zhiqi Li 0002, Honghua Chen, Jian J. Zhang 0001, Xiaosong Yang
Comput. Graph. Forum3
2023 Refine-Net: Normal Refinement Neural Network for Noisy Point Clouds
abstract
Point normal, as an intrinsic geometric property of 3D objects, not only serves conventional geometric tasks such as surface consolidation and reconstruction, but also facilitates cutting-edge learning-based techniques for shape analysis and generation. In this paper, we propose a normal refinement network, called Refine-Net, to predict accurate normals for noisy point clouds. Traditional normal estimation wisdom heavily depends on priors such as surface shapes or noise distributions, while learning-based solutions settle for single types of hand-crafted features. Differently, our network is designed to refine the initial normal of each point by extracting additional information from multiple feature representations. To this end, several feature modules are developed and incorporated into Refine-Net by a novel connection module. Besides the overall network architecture of Refine-Net, we propose a new multi-scale fitting patch selection scheme for the initial normal estimation, by absorbing geometry domain knowledge. Also, Refine-Net is a generic normal estimation framework: 1) point normals obtained from other methods can be further refined, and 2) any feature module related to the surface geometric structures can be potentially integrated into the framework. Qualitative and quantitative evaluations demonstrate the clear superiority of Refine-Net over the state-of-the-arts on both synthetic and real-scanned datasets.
Honghua Chen, Yingkui Zhang, Mingqiang Wei, Haoran Xie 0001, Jun Wang 0039, Tong Lu 0002, Harry Qin, Xiao-Ping Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 GeoDualCNN: Geometry-Supporting Dual Convolutional Neural Network for Noisy Point Clouds
abstract
We propose a geometry-supporting dual convolutional neural network (GeoDualCNN) for both point cloud normal estimation and denoising. GeoDualCNN fuses the geometry domain knowledge that the underlying surface of a noisy point cloud is piecewisely smooth with the fact that a point normal is properly defined only when local surface smoothness is guaranteed. Centered around this insight, we define the homogeneous neighborhood (HoNe) which stays clear of surface discontinuities, and associate each HoNe with a point whose geometry and normal orientation is mostly consistent with that of HoNe. Thus, we not only obtain initial estimates of the point normals by performing PCA on HoNes, but also for the first time optimize these initial point normals by learning the mapping from two proposed geometric descriptors to the ground-truth point normals. GeoDualCNN consists of two parallel branches that remove noise using the first geometric descriptor (a homogeneous height map, which encodes the point-position information), while preserving surface features using the second geometric descriptor (a homogeneous normal map, which encodes the point-normal information). Such geometry-supporting network architectures enable our model to leverage previous geometry expertise and to benefit from training data. Experiments with noisy point clouds show that GeoDualCNN outperforms the state-of-the-art methods in terms of both noise-robustness and feature preservation.
Mingqiang Wei, Honghua Chen, Yingkui Zhang, Haoran Xie 0001, Yanwen Guo 0001, Jun Wang 0039
IEEE Trans. Vis. Comput. Graph.2
2022 UTOPIC: Uncertainty-aware Overlap Prediction Network for Partial Point Cloud Registration
abstract
Abstract High‐confidence overlap prediction and accurate correspondences are critical for cutting‐edge models to align paired point clouds in a partial‐to‐partial manner. However, there inherently exists uncertainty between the overlapping and non‐overlapping regions, which has always been neglected and significantly affects the registration performance. Beyond the current wisdom, we propose a novel uncertainty‐aware overlap prediction network, dubbed UTOPIC, to tackle the ambiguous overlap prediction problem; to our knowledge, this is the first to explicitly introduce overlap uncertainty to point cloud registration. Moreover, we induce the feature extractor to implicitly perceive the shape knowledge through a completion decoder, and present a geometric relation embedding for Transformer to obtain transformation‐invariant geometry‐aware feature representations. With the merits of more reliable overlap scores and more precise dense correspondences, UTOPIC can achieve stable and accurate registration results, even for the inputs with limited overlapping areas. Extensive quantitative and qualitative experiments on synthetic and real benchmarks demonstrate the superiority of our approach over state‐of‐the‐art methods.
Zhilei Chen, Honghua Chen, Lina Gong, Xuefeng Yan 0001, Jun Wang 0039, Yanwen Guo 0001, Harry Qin, Mingqiang Wei
Comput. Graph. Forum2
2022 SPCNet: Stepwise Point Cloud Completion Network
abstract
Abstract How will you repair a physical object with large missings? You may first recover its global yet coarse shape and stepwise increase its local details. We are motivated to imitate the above physical repair procedure to address the point cloud completion task. We propose a novel stepwise point cloud completion network (SPCNet) for various 3D models with large missings. SPCNet has a hierarchical bottom‐to‐up network architecture. It fulfills shape completion in an iterative manner, which 1) first infers the global feature of the coarse result; 2) then infers the local feature with the aid of global feature; and 3) finally infers the detailed result with the help of local feature and coarse result. Beyond the wisdom of simulating the physical repair, we newly design a cycle loss to enhance the generalization and robustness of SPCNet. Extensive experiments clearly show the superiority of our SPCNet over the state‐of‐the‐art methods on 3D point clouds with large missings. Code is available at https://github.com/1127368546/SPCNet .
Honghua Chen, Xuequan Lu, Zhe Zhu, Jun Wang 0039, Weiming Wang 0002, Fu Lee Wang, Mingqiang Wei
Comput. Graph. Forum2
2022 Semi-MoreGAN: Semi-supervised Generative Adversarial Network for Mixture of Rain Removal
abstract
Abstract Real‐world rain is a mixture of rain streaks and rainy haze. However, current efforts formulate image rain streaks removal and rainy haze removal as separated models, worsening the loss of image details. This paper attempts to solve the mixture of rain removal problem in a single model by estimating the scene depths of images. To this end, we propose a novel SEMI‐ supervised M ixture O f rain RE moval G enerative A dversarial N etwork (Semi‐MoreGAN). Unlike most of existing methods, Semi‐MoreGAN is a joint learning paradigm of mixture of rain removal and depth estimation; and it effectively integrates the image features with the depth information for better rain removal. Furthermore, it leverages unpaired real‐world rainy and clean images to bridge the gap between synthetic and real‐world rain. Extensive experiments show clear improvements of our approach over twenty representative state‐of‐the‐arts on both synthetic and real‐world rainy images. Source code is available at https://github.com/syy-whu/Semi-MoreGAN .
Yiyang Shen, Yongzhen Wang 0001, Mingqiang Wei, Honghua Chen, Haoran Xie 0001, Gary Cheng 0001, Fu Lee Wang
Comput. Graph. Forum4
2022 RePCD-Net: Feature-Aware Recurrent Point Cloud Denoising Network
Honghua Chen, Zeyong Wei, Xianzhi Li 0001, Yabin Xu, Mingqiang Wei, Jun Wang 0039
Int. J. Comput. Vis.1
2022 Multiscale Feature Line Extraction From Raw Point Clouds Based on Local Surface Variation and Anisotropic Contraction
abstract
Recent 3-D scanning techniques can produce various kinds of digitized 3-D data. Most of these scanned data are in a format of unstructured point clouds. Such low-level representation of 3-D data usually contains only geometric properties (point positions), while lacking higher level structure cues, for example, feature lines. Feature lines can be defined as a visually prominent characteristic of the shape, including edges, ridges, and valley lines in multiple scales, which can support a lot of downstream applications, such as shape reconstruction and analysis. We present a two-phase algorithm for extracting line-type features on point clouds. To extract both large-scale and shallow feature lines, we first define a statistical metric to detect all potential feature points while immune to the noise to some extent. Then, for correctly reconstructing the feature lines from these identified coarse feature points, we introduce an anisotropic contracting scheme to force feature points lying on the underlying real feature lines. To illustrate the reliability of our method, various experiments have been conducted on both synthetic and raw data. Both visual and quantitative comparisons show that our method is robust to noise and can correctly extract multiscale feature lines. In addition, our method is generally applicable to robotic picking.Note to Practitioners—This article was motivated by the problem of the feature line extraction for real scanned point clouds. Feature lines, as one kind of the most important structure information, depict the basic shape of the real object in our life. Extracting this kind of shape features from the unstructured point clouds can facilitate a variety of downstream practical applications, such as product design, workpiece manufacturing, and robotic grasping. Existing approaches to detect features either heavily rely on differential quantities, which are sensitive to the noise, or need an elaborately designed local descriptor but fail to recognize small-scale features. These challenges motivate us to design a new approach aiming at extracting multiscale feature lines while keeping robustness to heavy noise. The technique developed in this work can produce high-quality feature points and feature lines, which would serve as higher level structural information and facilitate many applications. Additional applications in 6-degree-of-freedom (6-DoF) pose estimation demonstrate the potential of our method for robotic picking.
Honghua Chen, Yaoran Huang, Qian Xie 0001, Mingqiang Wei, Jun Wang 0039
IEEE Trans Autom. Sci. Eng.1
2021 Part-in-whole point cloud registration for aircraft partial scan automated localization
Qian Xie 0001, Xuanming Cao, Yabin Xu, Dening Lu, Honghua Chen, Jun Wang 0039
Comput. Aided Des.6
2021 Selective Guidance Normal Filter for Geometric Texture Removal
abstract
There is typically a trade-off between removing the detailed appearance (i.e., geometric textures) and preserving the intrinsic properties (i.e., geometric structures) of 3D surfaces. The conventional use of mesh vertex/facet-centered patches in many filters leads to side-effects including remnant textures, improperly filtered structures, and distorted shapes. We propose a selective guidance normal filter (SGNF) which adapts the Relative Total Variation (RTV) to a maximal/minimal scheme (mmRTV). The mmRTV measures the geometric flatness of surface patches, which helps in finding adaptive patches whose boundaries are aligned with the facet being processed. The adaptive patches provide selective guidance normals, which are subsequently used for normal filtering. The filtering smooths out the geometric textures by using guidance normals estimated from patches with maximal RTV (the least flatness), and preserves the geometric structures by using normals estimated from patches with minimal RTV (the most flatness). This simple yet effective modification of the RTV makes our SGNF specialized rather than trade off between texture removal and structure preservation, which is distinct from existing mesh filters. Experiments show that our approach is visually and numerically comparable to the state-of-the-art mesh filters, in most cases. In addition, the mmRTV is generally applicable to bas-relief modeling and image texture removal.
Mingqiang Wei, Yidan Feng, Honghua Chen
IEEE Trans. Vis. Comput. Graph.3
2020 Geometry and Learning Co-Supported Normal Estimation for Unstructured Point Cloud
abstract
In this paper, we propose a normal estimation method for unstructured point cloud. We observe that geometric estimators commonly focus more on feature preservation but are hard to tune parameters and sensitive to noise, while learning-based approaches pursue an overall normal estimation accuracy but cannot well handle challenging regions such as surface edges. This paper presents a novel normal estimation method, under the co-support of geometric estimator and deep learning. To lowering the learning difficulty, we first propose to compute a suboptimal initial normal at each point by searching for a best fitting patch. Based on the computed normal field, we design a normal-based height map network (NH-Net) to fine-tune the suboptimal normals. Qualitative and quantitative evaluations demonstrate the clear improvements of our results over both traditional methods and learning-based methods, in terms of estimation accuracy and feature recovery.
Honghua Chen, Yidan Feng, Qiong Wang 0001, Harry Qin, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei, Jun Wang 0039
CVPR2
2020 Aircraft Skin Rivet Detection Based on 3D Point Cloud via Multiple Structures Fitting
Qian Xie 0001, Dening Lu, Kunpeng Du, Jinxuan Xu, Jiajia Dai, Honghua Chen, Jun Wang 0039
Comput. Aided Des.6
2020 Multi-Patch Collaborative Point Cloud Denoising via Low-Rank Recovery with Graph Constraint
abstract
Point cloud is the primary source from 3D scanners and depth cameras. It usually contains more raw geometric features, as well as higher levels of noise than the reconstructed mesh. Although many mesh denoising methods have proven to be effective in noise removal, they hardly work well on noisy point clouds. We propose a new multi-patch collaborative method for point cloud denoising, which is solved as a low-rank matrix recovery problem. Unlike the traditional single-patch based denoising approaches, our approach is inspired by the geometric statistics which indicate that a number of surface patches sharing approximate geometric properties always exist within a 3D model. Based on this observation, we define a rotation-invariant height-map patch (HMP) for each point by robust Bi-PCA encoding bilaterally filtered normal information, and group its non-local similar patches together. Within each group, all patches are geometrically similar, while suffering from noise. We pack the height maps of each group into an HMP matrix, whose initial rank is high, but can be significantly reduced. We design an improved low-rank recovery model, by imposing a graph constraint to filter noise. Experiments on synthetic and raw datasets demonstrate that our method outperforms state-of-the-art methods in both noise removal and feature preservation.
Honghua Chen, Mingqiang Wei, Yangxing Sun, Xingyu Xie, Jun Wang 0039
IEEE Trans. Vis. Comput. Graph.1
2019 Structure-guided shape-preserving mesh texture smoothing via joint low-rank matrix recovery
Honghua Chen, Oussama Remil, Haoran Xie 0001, Harry Qin, Yanwen Guo 0001, Mingqiang Wei, Jun Wang 0039
Comput. Aided Des.1
2019 3D Shape Synthesis via Content-Style Revealing Priors
Oussama Remil, Qian Xie 0001, Honghua Chen, Jun Wang 0039
Comput. Aided Des.3
2019 Reliable Rolling-guided Point Normal Filtering for Surface Texture Removal
abstract
Abstract Semantic surface decomposition (SSD) facilitates various geometry processing and product re‐design tasks. Filter‐based techniques are meaningful and widely used to achieve the SSD, which however often leads to surface either under‐fitting or over‐fitting. In this paper, we propose a reliable rolling‐guided point normal filtering method to decompose textures from a captured point cloud surface. Our method is built on the geometry assumption that 3D surfaces are comprised of an underlying shape (US) and a variety of bump ups and downs (BUDs) on the US. We have three core contributions. First, by considering the BUDs as surface textures, we present a RANSAC‐based sub‐neighborhood detection scheme to distinguish the US and the textures. Second, to better preserve the US (especially the prominent structures), we introduce a patch shift scheme to estimate the guidance normal for feeding the rolling‐guided filter. Third, we formulate a new position updating scheme to alleviate the common uneven distribution of points. Both visual and numerical experiments demonstrate that our method is comparable to state‐of‐the‐art methods in terms of the robustness of texture removal and the effectiveness of the underlying shape preservation.
Yangxing Sun, Honghua Chen, Harry Qin, Mingqiang Wei, Hua Zong
Comput. Graph. Forum2
2018 Unsupervised Articulated Skeleton Extraction From Point Set Sequences Captured by a Single Depth Camera
abstract
How to robustly and accurately extract articulated skeletons from point set sequences captured by a single consumer-grade depth camera still remains to be an unresolved challenge to date. To address this issue, we propose a novel, unsupervised approach consisting of three contributions (steps): (i) a non-rigid point set registration algorithm to first build one-to-one point correspondences among the frames of a sequence; (ii) a skeletal structure extraction algorithm to generate a skeleton with reasonable numbers of joints and bones; (iii) a skeleton joints estimation algorithm to achieve accurate joints. At the end, our method can produce a quality articulated skeleton from a single 3D point sequence corrupted with noise and outliers. The experimental results show that our approach soundly outperforms state of the art techniques, in terms of both visual quality and accuracy.
Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Zhigang Deng 0001, Wenzhi Chen
AAAI2
2018 To Ask or Not to Ask: The Roles of Interpersonal Trust in Knowledge Seeking
abstract
This article looks to investigate the roles of interpersonal trust in knowledge seeking. Specifically, the article examines and tests the effects of two distinct types of interpersonal trust (affect-based trust and cognition-based trust) on willingness to seek two different types of knowledge (explicit and tacit). Using data from a survey of 143 employees from Chinese firms, the article found that both types of interpersonal trust positively related to explicit knowledge seeking, as well as tacit knowledge seeking. The article also found that cognition-based trust had a stronger relationship with seeking of both explicit and tacit knowledge than affect-based trust. Implications for future research and practice are discussed.
Michael Jijin Zhang, Honghua Chen
Int. J. Knowl. Manag.2
2018 GPF: GMM-Inspired Feature-Preserving Point Set Filtering
abstract
Point set filtering, which aims at reconstructing noise-free point sets from their corresponding noisy inputs, is a fundamental problem in 3D geometry processing. The main challenge of point set filtering is to preserve geometric features of the underlying geometry while at the same time removing the noise. State-of-the-art point set filtering methods still struggle with this issue: some are not designed to recover sharp features, and others cannot well preserve geometric features, especially fine-scale features. In this paper, we propose a novel approach for robust feature-preserving point set filtering, inspired by the Gaussian Mixture Model (GMM). Taking a noisy point set and its filtered normals as input, our method can robustly reconstruct a high-quality point set which is both noise-free and feature-preserving. Various experiments show that our approach can soundly outperform the selected state-of-the-art methods, in terms of both filtering quality and reconstruction accuracy.
Xuequan Lu, Honghua Chen, Sai-Kit Yeung, Wenzhi Chen, Matthias Zwicker
IEEE Trans. Vis. Comput. Graph.3
2007 A Method for Building Concept Lattice Based on Matrix Operation
Yajun Du, Dan Xiang, Honghua Chen, Zhenwen Liao
ICIC (2)4
2007 Representation of Rough Sets Based on Intuitionistic Fuzzy Special Sets
Zheng Pei 0001, Honghua Chen
IFSA (1)3
2007 Extracting Fuzzy Linguistic Summaries Based on Including Degree Theory and FCA
Zheng Pei 0001, Honghua Chen
IFSA (1)3