Jiacheng Deng 0002

dblp:320/4938-2 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0003-2838-0378ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration
abstract
Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect correspondences. Moreover, without dedicated designs, it remains difficult to effectively select informative and correlated representations across modalities, thereby limiting the robustness and accuracy of registration. To address these challenges, we propose a novel cross-modal registration framework composed of two key modules: the Iterative Agents Selection (IAS) module and the Reliable Agents Interaction (RAI) module. IAS enhances structural feature awareness with phase maps and employs reinforcement learning principles to efficiently select reliable agents. RAI then leverages these selected agents to guide cross-modal interactions, effectively reducing mismatches and improving overall robustness. Extensive experiments on the RGB-D Scenes v2 and 7-Scenes benchmarks demonstrate that our method consistently achieves state-of-the-art performance.
Zhixin Cheng, Xiaotian Yin, Jiacheng Deng 0002, Bohao Liao, Baoqun Yin, Tianzhu Zhang 0001
AAAI3
2026 GLASS: Geometry-Aware Local Alignment and Structure Synchronization Network for 2D-3D Registration
abstract
Image-to-point cloud registration methods typically follow a coarse-to-fine pipeline, extracting patch-level correspondences and refining them into dense pixel-to-point matches. However, in scenes with repetitive patterns, images often lack sufficient 3D structural cues and alignment with point clouds, leading to incorrect matches. Moreover, prior methods usually overlook structural consistency, limiting the full exploitation of correspondences. To address these issues, we propose two novel modules: the Local Geometry Enhancement (LGE) module and the Graph Distribution Consistency (GDC) module. LGE enhances both image and point cloud features with normal vectors, injecting geometric structure into image features to reduce mismatches. GDC constructs a graph from matched points to update features and explicitly constrain similarity distributions. Extensive experiments and ablations on two benchmarks, RGB-D Scenes v2 and 7-Scenes, demonstrate that our approach achieves state-of-the-art performance in image-to-point cloud registration.
Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Bohao Liao, Li Liu 0067, Xiaotian Yin, Baoqun Yin, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 ScaleDepth: Decomposing Metric Depth Estimation Into Semantic-Aware Scale Prediction and Adaptive Relative Depth Estimation
abstract
Estimating the depth map of an image in the wild is a challenging visual task. Compared to relative depth estimation, metric depth estimation attracts more attention due to its practical physical significance and critical applications in real-life scenarios. However, existing depth estimation methods typically focus only on generalization of relative depth, neglecting the importance of metric depth generalization. To address this challenge, we propose a novel monocular depth estimation method called ScaleDepth. It decomposes metric depth into scene scale and relative depth and predicts them through a semantic-aware scale prediction (SASP) module and an adaptive relative depth estimation (ARDE) module, respectively. Our proposed approach has several merits. First, the SASP module can implicitly combine structural and semantic features of the images to predict precise scene scales. Second, the ARDE module can adaptively estimate the relative depth distribution of each image within a normalized depth space. Third, our method achieves metric depth estimation for both indoor and outdoor scenes in a unified framework, without the need to set the depth range or fine-tune the model. Extensive experiments demonstrate that our method achieves competitive performance in indoor, outdoor, unconfined, and unseen scenes. Project page: https://ruijiezhu94.github.io/ScaleDepth.
Ruijie Zhu 0002, Chuxin Wang, Ziyang Song 0001, Li Liu 0067, Jiacheng Deng 0002, Tianzhu Zhang 0001, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Adaptive Siamese Masked Autoencoder with Global Optimization for Unsupervised Point Cloud Shape Correspondence
abstract
Unsupervised point cloud shape correspondence aims to establish point-wise correspondences between point clouds without annotated data. Ensuring efficiency and accuracy is crucial for practically implementing point cloud shape correspondence. Although the current methods have achieved desirable performance, the nature of encoding at dense points limits their application in actual scenarios. Moreover, independently computing per-point correspondences results in numerous multiple-to-one erroneous correspondences. To address these issues, we present an Adaptive siamese Masked autoencoder with Global Optimization (AMIGO), comprising a siamese masked autoencoder and a global optimization module. In the siamese masked autoencoder, we downsample the input point cloud and employ adaptive siamese mask operations to boost the coding capabilities of the encoder, thereby mitigating the information loss caused by downsampling. In the global optimization module, optimal transport is only utilized to generate pseudo-labels during the training phase, facilitating the efficient global planning of the correspondence results. Extensive experiments on four standard human and animal benchmarks demonstrate that AMIGO surpasses existing methods with remarkable margins, achieving new state-of-the-art results.
Jiacheng Deng 0002, Jiahao Lu 0001
AAAI1
2025 DiffCorr: Conditional Diffusion Model with Reliable Pseudo-Label Guidance for Unsupervised Point Cloud Shape Correspondence
abstract
Unsupervised point cloud shape correspondence aims to establish dense correspondences between source and target point clouds. Existing methods universally follow a one-step paradigm to obtain shape correspondence directly, but it often fails in large-scale motions of humans and animals. To address this challenge, we propose a conditional Diffusion model with reliable pseudo-label guidance for unsupervised point cloud shape Correspondence (DiffCorr), including a transformer-based conditional diffusion model and a reliable pseudo-label generator. The proposed DiffCorr enjoys several merits. Firstly, the transformer-based conditional diffusion model implements a coarse-to-fine optimization for coarse correspondences. Secondly, we design a reliable pseudo-label generator to provide high-quality pseudo-labels for training. Extensive experiments on four human and animal datasets demonstrate that DiffCorr surpasses state-of-the-art methods and exhibits favorable generalization capabilities.
Jiacheng Deng 0002, Jiahao Lu 0001, Zhixin Cheng, Wenfei Yang
AAAI1
2025 Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain Alignment
abstract
The method for image-to-point cloud registration typically determines the rigid transformation using a coarse-to-fine pipeline. However, directly and uniformly matching image patches with point cloud patches may lead to focusing on incorrect noise patches during matching while ignoring key ones. Moreover, due to the significant differences between image and point cloud modalities, it may be challenging to bridge the domain gap without specific improvements in design. To address the above issues, we innovatively propose the Uncertainty-aware Hierarchical Matching Module (UHMM) and the Adversarial Modal Alignment Module (AMAM). Within the UHMM, we model the uncertainty of critical information in image patches and facilitate multi-level fusion interactions between image and point cloud features. In the AMAM, we design an adversarial approach to reduce the domain gap between image and point cloud. Extensive experiments and ablation studies on RGB-D Scene V2 and 7-Scenes benchmarks demonstrate the superiority of our method, making it a state-of-the-art approach for image-to-point cloud registration tasks.
Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Baoqun Yin, Tianzhu Zhang 0001
AAAI2
2025 Pamba: Enhancing Global Interaction in Point Clouds via State Space Model
abstract
Transformers have demonstrated impressive results for 3D point cloud semantic segmentation. However, the quadratic complexity of transformer makes computation costs high, limiting the number of points that can be processed simultaneously and impeding the modeling of long-range dependencies between objects in a single scene. Drawing inspiration from the great potential of recent state space models (SSM) for long sequence modeling, we introduce Mamba, an SSM-based architecture, to the point cloud domain and propose Pamba, a novel architecture with strong global modeling capability under linear complexity. Specifically, to make the disorderness of point clouds fit in with the causal nature of Mamba, we propose a multi-path serialization strategy applicable to point clouds. Besides, we propose the ConvMamba block to compensate for the shortcomings of Mamba in modeling local geometries and in unidirectional modeling. Pamba obtains state-of-the-art results on several 3D point cloud segmentation tasks, including ScanNet v2, ScanNet200, S3DIS and nuScenes, while its effectiveness is validated by extensive experiments.
Yubo Ai, Jiahao Lu 0001, Chuxin Wang, Jiacheng Deng 0002, Hanzhi Chang, Yanzhe Liang, Wenfei Yang, Tianzhu Zhang 0001
AAAI5
2025 Implicit Correspondence Learning for Image-to-Point Cloud Registration
abstract
Image-to-point cloud registration aims to estimate the camera pose of a given image within a 3D scene point cloud. In this area, matching-based methods have achieved leading performance by first detecting the overlapping region, then matching point and pixel features learned by neural networks and finally using the PnP-RANSAC algorithm to estimate camera pose. However, achieving accurate image-to-point cloud registration remains challenging because the overlapping region detection is unreliable merely relying on point-wise classification, direct alignment of cross-modal data is difficult and indirect optimization objective leads to unstable registration results. To address these challenges, we propose a novel implicit correspondence learning method, including a Geometric Prior-guided overlapping region Detection Module (GPDM), an Implicit Correspondence Learning Module (ICLM), and a Pose Regression Module (PRM). The proposed method enjoys several merits. First, the proposed GPDM can precisely detect the overlapping region. Second, the ICLM can generate robust cross-modality correspondences. Third, the PRM can enable end-to-end optimization. Extensive experimental results on KITTI and nuScenes datasets demonstrate that the proposed model sets a new state-of-the-art performance in registration accuracy.
Xinjun Li, Wenfei Yang, Jiacheng Deng 0002, Zhixin Cheng, Tianzhu Zhang 0001
CVPR3
2025 Relation3D : Enhancing Relation Modeling for Point Cloud Instance Segmentation
abstract
3D instance segmentation aims to predict a set of object instances in a scene, representing them as binary foreground masks with corresponding semantic labels. Currently, transformer-based methods are gaining increasing attention due to their elegant pipelines and superior predictions. However, these methods primarily focus on modeling the external relationships between scene features and query features through mask attention. They lack effective modeling of the internal relationships among scene features as well as between query features. In light of these disadvantages, we propose Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation. Specifically, we introduce an adaptive superpoint aggregation module and a contrastive learning-guided superpoint refinement module to better represent superpoint features (scene features) and leverage contrastive learning to guide the updates of these features. Furthermore, our relation-aware self-attention mechanism enhances the capabilities of modeling relationships between queries by incorporating positional and geometric relationships into the self-attention mechanism. Extensive experiments on the ScanNetV2, ScanNet++, ScanNet200 and S3DIS datasets demonstrate the superior performance of Relation3D. Code is available at this website.
Jiahao Lu 0001, Jiacheng Deng 0002
CVPR2
2025 CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection
abstract
Detection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching results, ultimately impairing registration accuracy. Furthermore, similar structures in the scene could lead to redundant correspondences in cross-modal matching. To address these issues, we propose Channel Adaptive Adjustment Module (CAA) and Global Optimal Selection Module (GOS). CAA enhances intra-modal features and suppresses cross-modal sensitivity, while GOS replaces local selection with global optimization. Experiments on RGB-D Scenes V2 and 7-Scenes demonstrate the superiority of our method, achieving state-of-the-art performance in image-to-point cloud registration.
Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Xiaotian Yin, Bohao Liao, Baoqun Yin, Wenfei Yang, Tianzhu Zhang 0001
ICCV2
2025 SAS: Segment Any 3D Scene with Integrated 2D Priors
abstract
The open vocabulary capability of 3D models is increasingly valued, as traditional methods with models trained with fixed categories fail to recognize unseen objects in complex dynamic 3D scenes. In this paper, we propose a simple yet effective approach, SAS, to integrate the open vocabulary capability of multiple 2D models and migrate it to 3D domain. Specifically, we first propose Model Alignment via Text to map different 2D models into the same embedding space using text as a bridge. Then, we propose Annotation-Free Model Capability Construction to explicitly quantify the 2D model's capability of recognizing different categories using diffusion models. Following this, point cloud features from different 2D models are fused with the guide of constructed model capabilities. Finally, the integrated 2D open vocabulary capability is transferred to 3D domain through feature distillation. SAS outperforms previous methods by a large margin across multiple datasets, including ScanNet v2, Matterport3D, and nuScenes, while its generalizability is further validated on downstream tasks, e.g., gaussian segmentation and instance segmentation.
Jiahao Lu 0001, Jiacheng Deng 0002, Hanzhi Chang, Yanzhe Liang, Tianzhu Zhang 0001
ICCV3
2025 TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
abstract
Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic motion and cannot densely track newly emerging dynamic subjects in the videos. To address these two limitations, we propose TrackingWorld, a novel pipeline for dense 3D tracking of almost all pixels within a world-centric 3D coordinate system. First, we introduce a tracking upsampler that efficiently lifts the arbitrary sparse 2D tracks into dense 2D tracks. Then, to generalize the current tracking methods to newly emerging objects, we apply the upsampler to all frames and reduce the redundancy of 2D tracks by eliminating the tracks in overlapped regions. Finally, we present an efficient optimization-based framework to back-project dense 2D tracks into world-centric 3D trajectories by estimating the camera poses and the 3D coordinates of these 2D tracks. Extensive evaluations on both synthetic and real-world datasets demonstrate that our system achieves accurate and dense 3D tracking in a world-centric coordinate frame.
Jiahao Lu 0001, Weitao Xiong, Jiacheng Deng 0002, Zhiyang Dou, Cheng Lin 0001, Sai-Kit Yeung, Yuan Liu 0025
NeurIPS3
2025 Quantity-Quality Enhanced Self-Training Network for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Point cloud semantic segmentation is essential for understanding 3D scenes. Contemporary techniques often require extensive annotated training data, yet obtaining point-wise annotations for point clouds is time-consuming and laborious. Recent developments in weakly supervised methods seek to mitigate this problem by generating pseudo-labels using limited annotations. However, these pseudo-labels frequently suffer from either insufficient quantity or inferior quality. To overcome these hurdles, we introduce a Quantity-Quality Enhanced Self-training Network for Weakly Supervised Point Cloud Semantic Segmentation (Q2E). Specifically, an image-assisted pseudo-label generator is proposed to exploit 2D images to extend pseudo-labels for point clouds. Additionally, a hierarchical pseudo-label optimizer is developed to refine the quality of the pseudo-labels by hierarchically grouping them into broader categories. Extensive experiments on the ScanNet-v2, S3DIS, Semantic3D, and SemanticKITTI datasets demonstrate that Q2E outperforms state-of-the-art weakly supervised methods and rivals fully supervised approaches for point cloud semantic segmentation. Remarkably, as of the initial submission on February 2, 2024, our method ranked the first place in various settings of the ScanNet-v2 benchmark.
Jiacheng Deng 0002, Jiahao Lu 0001, Tianzhu Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Plane2Depth: Hierarchical Adaptive Plane Guidance for Monocular Depth Estimation
abstract
Monocular depth estimation aims to infer a dense depth map from a single image, which is a fundamental and prevalent task in computer vision. Many previous works have shown impressive depth estimation results through carefully designed network structures, but they usually ignore the planar information and therefore perform poorly in low-texture areas of indoor scenes. In this paper, we propose Plane2Depth, which adaptively utilizes plane information to improve depth prediction within a hierarchical framework. Specifically, in the proposed plane guided depth generator (PGDG), we design a set of plane queries as prototypes to softly model planes in the scene and predict per-pixel plane coefficients. Then the predicted plane coefficients can be converted into metric depth values with the pinhole camera model. In the proposed adaptive plane query aggregation (APGA) module, we introduce a novel feature interaction approach to improve the aggregation of multi-scale plane features in a top-down manner. Extensive experiments show that our method can achieve outstanding performance, especially in low-texture or repetitive areas. Furthermore, under the same backbone network, our method outperforms the state-of-the-art methods on the NYU-Depth-v2 dataset, achieves competitive results with state-of-the-art methods KITTI dataset and can be generalized to unseen scenes effectively.
Li Liu 0067, Ruijie Zhu 0002, Jiacheng Deng 0002, Ziyang Song 0001, Wenfei Yang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 ER-Depth: Enhancing the Robustness of Self-Supervised Monocular Depth Estimation in Challenging Scenes
abstract
Self-supervised monocular depth estimation holds significant importance in the fields of autonomous driving and robotics. However, existing methods are typically trained and evaluated on clear, sunny datasets, overlooking the impact of various adverse conditions commonly encountered in real-world applications, such as rainy weather, low visibility, and motion blur. As a result, they often struggle in challenging scenarios and produce artifacts. To address this issue, we propose ER-Depth, a novel two-stage self-supervised framework designed for robust depth estimation. In the first stage, we propose perturbation-invariant depth consistency regularization to propagate reliable supervision from standard to challenging scenes. In the second stage, we adopt the Mean Teacher paradigm for self-distillation and present a novel consistency-based pseudo-label filtering strategy to improve the quality of pseudo-labels. Extensive experiments demonstrate that our method exhibits exceptional robustness in challenging scenarios while maintaining high performance in standard scenes, significantly outperforming existing state-of-the-art methods on challenging KITTI-C, DrivingStereo, and NuScenes-Night benchmarks. Project page: https://ruijiezhu94.github.io/ERDepth_page .
Ziyang Song 0001, Ruijie Zhu 0002, Chuxin Wang, Jiacheng Deng 0002, Wenfei Yang, Tianzhu Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Unsupervised Template-assisted Point Cloud Shape Correspondence Network
abstract
Unsupervised point cloud shape correspondence aims to establish point-wise correspondences between source and target point clouds. Existing methods obtain correspon-dences directly by computing point-wise feature similar-ity between point clouds. However, non-rigid objects pos-sess strong deformability and unusual shapes, making it a longstanding challenge to directly establish correspon-dences between point clouds with unconventional shapes. To address this challenge, we propose an unsupervised Template-Assisted point cloud shape correspondence Net-work, termed TANet, including a template generation mod-ule and a template assistance module. The proposed TANet enjoys several merits. Firstly, the template generation mod-ule establishes a set of learnable templates with explicit structures. Secondly, we introduce a template assistance module that extensively leverages the generated templates to establish more accurate shape correspondences from multiple perspectives. Extensive experiments on four hu-man and animal datasets demonstrate that TANet achieves favorable performance against state-of-the-art methods.
Jiacheng Deng 0002, Jiahao Lu 0001, Tianzhu Zhang 0001
CVPR1
2024 BSNet: Box-Supervised Simulation-Assisted Mean Teacher for 3D Instance Segmentation
abstract
3D instance segmentation (3DIS) is a crucial task, but point-level annotations are tedious in fully supervised settings. Thus, using bounding boxes (bboxes) as annotations has shown great potential. The current mainstream approach is a two-step process, involving the generation of pseudo-labels from box annotations and the training of a 3DIS network with the pseudo-labels. However, due to the presence of intersections among bboxes, not every point has a determined instance label, especially in over-lapping areas. To generate higher quality pseudo-labels and achieve more precise weakly supervised 3DIS results, we propose the Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation (BSNet), which devises a novel pseudo-labeler called Simulation-assisted Transformer. The labeler consists of two main components. The first is Simulation-assisted Mean Teacher, which introduces Mean Teacher for the first time in this task and constructs simulated samples to assist the labeler in acquiring prior knowledge about overlapping areas. To better model local-global structure, we also propose Local-Global Aware Attention as the decoder for teacher and student labelers. Extensive experiments conducted on the ScanNetV2 and S3DIS datasets verify the superiority of our designs. Code is available at https://github.com/peopleu/BSNet.
Jiahao Lu 0001, Jiacheng Deng 0002, Tianzhu Zhang 0001
CVPR2
2024 Diff3DETR: Agent-Based Diffusion Model for Semi-supervised 3D Object Detection
Jiacheng Deng 0002, Jiahao Lu 0001, Tianzhu Zhang 0001
ECCV (34)1
2024 MotionGS: Exploring Explicit Motion Guidance for Deformable 3D Gaussian Splatting
abstract
Dynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to dynamic scenes, they often lack explicit constraints on object motion, leading to optimization difficulties and performance degradation. To address the above issues, we propose a novel deformable 3D Gaussian splatting framework called MotionGS, which explores explicit motion priors to guide the deformation of 3D Gaussians. Specifically, we first introduce an optical flow decoupling module that decouples optical flow into camera flow and motion flow, corresponding to camera movement and object motion respectively. Then the motion flow can effectively constrain the deformation of 3D Gaussians, thus simulating the motion of dynamic objects. Additionally, a camera pose refinement module is proposed to alternately optimize 3D Gaussians and camera poses, mitigating the impact of inaccurate camera poses. Extensive experiments in the monocular dynamic scenes validate that MotionGS surpasses state-of-the-art methods and exhibits significant superiority in both qualitative and quantitative results. Project page: https://ruijiezhu94.github.io/MotionGS_page.
Ruijie Zhu 0002, Yanzhe Liang, Hanzhi Chang, Jiacheng Deng 0002, Jiahao Lu 0001, Wenfei Yang, Tianzhu Zhang 0001, Yongdong Zhang 0001
NeurIPS4
2024 DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering
abstract
Dynamic scenes rendering is an intriguing yet challenging problem. Although current methods based on NeRF have achieved satisfactory performance, they still can not reach real-time levels. Recently, 3D Gaussian Splatting (3DGS) has garnered researchers' attention due to their outstanding rendering quality and real-time speed. Therefore, a new paradigm has been proposed: defining a canonical 3D gaussians and deforming it to individual frames in deformable fields. However, since the coordinates of canonical 3D gaussians are filled with noise, which can transfer noise into the deformable fields, and there is currently no method that adequately considers the aggregation of 4D information. Therefore, we propose Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering (DN-4DGS). Specifically, a Noise Suppression Strategy is introduced to change the distribution of the coordinates of the canonical 3D gaussians and suppress noise. Additionally, a Decoupled Temporal-Spatial Aggregation Module is designed to aggregate information from adjacent points and frames. Extensive experiments on various real-world datasets demonstrate that our method achieves state-of-the-art rendering quality under a real-time level. Code is available at https://github.com/peoplelu/DN-4DGS.
Jiahao Lu 0001, Jiacheng Deng 0002, Ruijie Zhu 0002, Yanzhe Liang, Wenfei Yang, Tianzhu Zhang 0001
NeurIPS2
2023 SE-ORNet: Self-Ensembling Orientation-Aware Network for Unsupervised Point Cloud Shape Correspondence
abstract
Unsupervised point cloud shape correspondence aims to obtain dense point-to-point correspondences between point clouds without manually annotated pairs. However, humans and some animals have bilateral symmetry and various orientations, which lead to severe mispredictions of symmetrical parts. Besides, point cloud noise disrupts consistent representations for point cloud and thus degrades the shape correspondence accuracy. To address the above issues, we propose a Self-Ensembling ORientation-aware Network termed SE-ORNet. The key of our approach is to exploit an orientation estimation module with a domain adaptive discriminator to align the orientations of point cloud pairs, which significantly alleviates the mispredictions of symmetrical parts. Additionally, we design a self-ensembling framework for unsupervised point cloud shape correspondence. In this framework, the disturbances of point cloud noise are overcome by perturbing the inputs of the student and teacher networks with different data augmentations and constraining the consistency of predictions. Extensive experiments on both human and animal datasets show that our SE-ORNet can surpass state-of-the-art unsupervised point cloud shape correspondence methods.
Jiacheng Deng 0002, Chuxin Wang, Jiahao Lu 0001, Tianzhu Zhang 0001, Jiyang Yu
CVPR1
2023 Query Refinement Transformer for 3D Instance Segmentation
abstract
3D instance segmentation aims to predict a set of object instances in a scene and represent them as binary foreground masks with corresponding semantic labels. However, object instances are diverse in shape and category, and point clouds are usually sparse, unordered, and irregular, which leads to a query sampling dilemma. Besides, noise background queries interfere with proper scene perception and accurate instance segmentation. To address the above issues, we propose the Query Refinement Transformer termed QueryFormer. The key to our approach is to exploit a query initialization module to optimize the initialization process for the query distribution with a high coverage and low repetition rate. Additionally, we design an affiliated transformer decoder that suppresses the interference of noise background queries and helps the foreground queries focus on instance discriminative parts to predict final segmentation results. Extensive experiments on ScanNetV2 and S3DIS datasets show that our QueryFormer can surpass state-of-the-art 3D instance segmentation methods.
Jiahao Lu 0001, Jiacheng Deng 0002, Chuxin Wang, Tianzhu Zhang 0001
ICCV2
2023 Long-Short Range Adaptive Transformer With Dynamic Sampling for 3D Object Detection
abstract
3D object detection in point cloud aims at simultaneously localizing and recognizing 3D objects from a 3D point set. However, since point clouds are usually sparse, unordered, and irregular, it is challenging to learn robust point representations and sample high-quality object queries. To deal with the above issues, we propose a Long-short rangE Adaptive transformer with Dynamic sampling (LeadNet), including a point representation encoder, a dynamic object query sampling decoder, and an object detection decoder in a unified architecture for 3D object detection. Specifically, in the point representation encoder, we combine an attention layer and a channel attentive kernel convolution layer to consider the local structure and the long-range context simultaneously. In the dynamic object query sampling decoder, we utilize multiple dynamic prototypes to adapt to various point clouds. In the object detection decoder, we incorporate a dynamic Gaussian weight map into the cross-attention mechanism to help the detection decoder focus on the proper visual regions near the object, further accelerating the training process. Extensive experimental results on two standard benchmarks show that our LeadNet outperforms the 3DETR baseline by$11.6\%~\text {mAP}_{50}$on the ScanNet v2 dataset and achieves the new state-of-the-art results on ScanNet v2 and SUN RGB-D benchmarks for the geometric-only approaches.
Chuxin Wang, Jiacheng Deng 0002, Tianzhu Zhang 0001, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Hierarchical Shape-Consistent Transformer for Unsupervised Point Cloud Shape Correspondence
abstract
Point cloud shape correspondence aims at accurately mapping one point cloud to another point cloud with various 3D shapes. Since point clouds are usually sparse, disordered, irregular, and with diverse shapes, it is challenging to learn consistent point cloud representations and achieve the accurate matching of different point cloud shapes. To address the above issues, we propose a Hierarchical Shape-consistent TRansformer for unsupervised point cloud shape correspondence (HSTR), including a multi-receptive-field point representation encoder and a shape-consistent constrained module in a unified architecture. The proposed HSTR enjoys several merits. In the multi-receptive-field point representation encoder, we set progressively larger receptive fields in different blocks to simultaneously consider the local structure and the long-range context. In the shape-consistent constrained module, we design two novel shape selective whitening losses, which can complement each other to achieve suppression of features sensitive to shape change. Extensive experimental results on four standard benchmarks demonstrate the superiority and generalization ability of our approach to existing methods at the similar model scale, and our method achieves the new state-of-the-art results.
Jiacheng Deng 0002, Tianzhu Zhang 0001, Yongdong Zhang 0001
IEEE Trans. Image Process.2