Ronghan Chen

dblp:300/4552 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-6307-2923ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 81% Robot manipulation · 11% Motion planning and robot control · 6%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
0.912025
DetailRecon: Focusing on Detailed Regions for Online Monocular 3D Reconstruction · IEEE Trans. Multim. 2025
Computer vision › 3D vision
3d scene understanding
0.912025
DetailRecon: Focusing on Detailed Regions for Online Monocular 3D Reconstruction · IEEE Trans. Multim. 2025
Computer vision › 3D vision › 3d reconstruction
online reconstruction
0.912025
DetailRefine: Towards Fine-Grained and Efficient Online Monocular 3D Reconstruction · ICRA 2025
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.912025
DetailRefine: Towards Fine-Grained and Efficient Online Monocular 3D Reconstruction · ICRA 2025
Computer vision › 3D vision › feature matching › 3d correspondence
2d-3d correspondence
0.812024
Marrying NeRF with Feature Matching for One-step Pose Estimation · ICRA 2024
Computer vision › 3D vision › pose estimation › learning-based pose estimation
NeRF-based pose estimation
0.812024
Marrying NeRF with Feature Matching for One-step Pose Estimation · ICRA 2024
Computer vision › 3D vision
object pose estimation
0.812024
Marrying NeRF with Feature Matching for One-step Pose Estimation · ICRA 2024
Computer vision › 3D vision
pose estimation
0.812024
Marrying NeRF with Feature Matching for One-step Pose Estimation · ICRA 2024
Robotics › Robot manipulation › object manipulation
category-level manipulation
0.712023
Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration · CVPR 2023
Robotics › Robot manipulation
deformable object manipulation
0.712023
Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration · CVPR 2023
Robotics › Motion planning and robot control
robot learning
0.712023
Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration · CVPR 2023
Computer vision › 3D vision › geometric deep learning
3d deep learning
0.612022
The Devil is in the Pose: Ambiguity-free 3D Rotation-invariant Learning via Pose-aware Convolution · CVPR 2022
Computer vision › 3D vision › geometric deep learning
rotation-invariant convolution
0.612022
The Devil is in the Pose: Ambiguity-free 3D Rotation-invariant Learning via Pose-aware Convolution · CVPR 2022
Geometric modeling and processing
deformation modeling
0.512021
Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence · ICCV 2021
Geometric modeling and processing › shape matching
non-rigid shape matching
0.512021
Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence · ICCV 2021
Geometric modeling and processing
shape correspondence
0.512021
Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence · ICCV 2021
Computer vision › Segmentation and scene understanding
scene understanding
0.312025
DetailRefine: Towards Fine-Grained and Efficient Online Monocular 3D Reconstruction · ICRA 2025
Computer vision › 3D vision
neural radiance field
0.212024
Marrying NeRF with Feature Matching for One-step Pose Estimation · ICRA 2024
Geometric modeling and processing › point cloud processing
point cloud representation
0.212022
The Devil is in the Pose: Ambiguity-free 3D Rotation-invariant Learning via Pose-aware Convolution · CVPR 2022

Methods — techniques the papers use, named apart from their topics

hierarchical hybrid sparsification · 0.9hierarchical consistency loss · 0.9dynamic detail refinement · 0.9connectivity-aware sparsification · 0.9coarse-to-fine pipeline · 0.9adaptive hybrid fusion · 0.9sampling strategy · 0.8point mining · 0.8feature matching · 0.8neural spatial encoding · 0.7factorized dynamic kernel · 0.6augmented point pair feature · 0.6maximum mean discrepancy · 0.5deformation graph · 0.5autoencoder · 0.5
YearPublicationVenuePosition
2026 Generalizable Multistage Assembly via One-Shot Category-Level Demonstration
abstract
Imitation learning offers a flexible approach for robot skill acquisition, enabling robots to learn complex tasks directly from demonstrations. However, most existing methods require a large number of demonstrations, whereas humans typically only need one or a few demonstrations. This discrepancy results in significant time consumption for data collection. Furthermore, these methods often assume that test scenarios will always be identical to the demonstration, which can lead to substantial performance degradation when facing novel scenarios, such as manipulating objects from the same category but with different shapes and sizes, or encountering object collisions during manipulation. To address these challenges, we propose a generalized multistage manipulation network for category-level robot assembly tasks. This network allows a robot to learn a multistage screw-nut assembly task from a single demonstration and generalize to new object instances with varying shapes and sizes. Specifically, the network uses category-level pose estimation to extract manipulation trajectories from the demonstration and applies manipulation-pose generalization to transfer these trajectories to novel instances. In addition, real-time action correction adjusts the trajectory based on real-time force feedback, enabling the robot to adapt to unexpected collisions during execution. We validate our method through experiments in both simulation and real-world environments, verifying its effectiveness and flexibility.
Yang Cong, Ronghan Chen, Wei Cong, Gan Sun
IEEE Trans. Neural Networks Learn. Syst.3
2025 DetailRefine: Towards Fine-Grained and Efficient Online Monocular 3D Reconstruction
abstract
Online monocular 3D reconstruction has attracted widespread attention as it promotes the application of robots in interactive scenarios. Most existing methods focus on 1) real-time reconstruction, 2) accurate voxel featuring learning, and 3) effective voxel sparsification algorithm. To this end, 1) they adopt a coarse-to-fine pipeline, where all non-empty voxels are sent to the next level for refinement. However, this results in over-refinement of flat regions, leading to unnecessary computational overhead. Furthermore, 2) advanced methods focus on exploring view visibility but overlook the discriminability among visible views, which limits the representation of learned voxel features. Moreover, 3) existing sparsification algorithms struggle to distinguish detailed and empty voxels, resulting in either the loss of detailed voxels or the retention of empty voxels. To tackle these challenges, 1) we present Dynamic Detail Refinement (DDR) to allocate more voxels to detailed regions for refinement, which could alleviate the computational burden. Furthermore, 2) we propose Discriminability-Aware Fusion (DAF) to focus on discriminative views, which helps to capture accurate voxel features. In addition, 3) we propose Hierarchical Hybrid Sparsification (HHS) to balance global completeness and local refinement, which helps to preserve detailed voxels at hierarchical levels effectively. Extensive experiments conducted on the representative ScanNet (V2) and 7-Scenes datasets demonstrate the superiority of the proposed method.
Fupeng Chu, Yang Cong, Ronghan Chen
ICRA3
2025 Learning Generalizable 3D Manipulation With 10 Demonstrations
abstract
Learning robust and generalizable manipulation skills from few demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. Although recent imitation learning methods have achieved impressive results, they often require a large amount of demonstration data and struggle to generalize across different spatial variants. In this work, we propose a framework that learns 3D manipulation policies from only 10 demonstrations while achieving robust generalization to unseen spatial configurations through semantic-guided perception and spatial-equivariant policy learning. Our framework consists of two key modules: a Semantic Guided Perception module that extracts task-aware 3D representations from RGB-D inputs using semantic priors and a Spatial Generalized Decision module implementing a diffusion-based policy that preserves spatial equivariance through denoising. Central to our framework is a spatially equivariant training strategy, which adapts 2D data augmentation principles to 3D manipulation by maintaining gripper-object spatial relationships during trajectory augmentation. We validate our framework through extensive experiments on both simulation benchmarks and real-world robotic systems. Our method demonstrates a significant improvement in success rates over state-of-the-art approaches on a series of challenging tasks, particularly under significant object pose variations. This work shows significant potential to advance efficient and generalizable manipulation skill learning in real-world applications.
Yang Cong, Bohao Huang, Jiahao Long, Ronghan Chen, Huijie Fan
IROS5
2025 DetailRecon: Focusing on Detailed Regions for Online Monocular 3D Reconstruction
abstract
Learning-based online monocular 3D reconstruction has emerged with great potential recently. Most state-of-the-art methods focus on two key questions, namely 1) how to exploit accurate voxel features and 2) how to preserve detailed voxels in the sparsification process. However, 1) most methods adopt the same receptive field to extract features for both informative and uninformative regions, which struggle to capture geometric details. Furthermore, 2) they mainly utilize a fixed threshold or a straightforward ray-based algorithm to discard voxels in the sparsification process. However, some detailed regions (especially thin regions) may be discarded incorrectly. To tackle these challenges, we present a novel method named DetailRecon to focus on detailed regions that contain more geometric information. Specifically, we first propose an Adaptive Hybrid Fusion (AHF) module and a Connectivity-Aware Sparsification (CAS) module for voxel feature learning and voxel sparsification, respectively. 1) The AHF receives multiple feature maps with different receptive fields as input, and adaptively adopts a smaller receptive field for regions with fine structures to exploit accurate geometric details. 2) The CAS updates the occupancy value of voxels based on the connected voxels within its neighbor space, which could expand the radiation range of reliable voxels in detailed regions and eventually reduce their probability of being discarded. Moreover, 3) we introduce a lightweight yet effective pipeline named Focus On Fine (FOF) to accelerate our DetailRecon. In addition, 4) we propose a Hierarchical Consistency Loss (HCL) to align multi-level volume features, which assists in exploring accurate volume features for recovering more details. Extensive experiments conducted on the ScanNet (V2) and 7-Scenes datasets demonstrate the superiority of our DetailRecon.
Fupeng Chu, Yang Cong, Ronghan Chen
IEEE Trans. Multim.4
2024 Marrying NeRF with Feature Matching for One-step Pose Estimation
abstract
Given the image collection of an object, we aim at building a real-time image-based pose estimation method, which requires neither its CAD model nor hours of object-specific training. Recent NeRF-based methods provide a promising solution by directly optimizing the pose from pixel loss between rendered and target images. However, during inference, they require long converging time, and suffer from local minima, making them impractical for real-time robot applications. We aim at solving this problem by marrying image matching with NeRF. With 2D matches and depth rendered by NeRF, we directly solve the pose in one step by building 2D-3D correspondences between target and initial view, thus allowing for real-time prediction. Moreover, to improve the accuracy of 2D-3D correspondences, we propose a 3D consistent point mining strategy, which effectively discards unfaithful points reconstruted by NeRF. Moreover, current NeRF-based methods naively optimizing pixel loss fail at occluded images. Thus, we further propose a 2D matches based sampling strategy to preclude the occluded area. Experimental results on representative datasets prove that our method outperforms state-of-the-art methods, and improves inference efficiency by 90×, achieving real-time prediction at 6 FPS.
Ronghan Chen, Yang Cong
ICRA1
2024 OPEN: Occlusion-Invariant Perception Network for Single Image-Based 3D Shape Retrieval
abstract
Single image-based 3D shape retrieval (IBSR) has attracted appealing academic interests recently, which aims to find the corresponding 3D shape from a shape repository for a given single 2D image. However, state-of-the-art methods neglect the discrepancy in the image domain due to unavoidable occlusion. The occluded image representations acting as noise, may perturb the alignment of the normal 2D representations with the 3D representations, resulting in occlusion-sensitive image-shape retrieval. To tackle this crucial challenge, in this paper, we propose a novel Occlusion-invariant PErception Network (OPEN) to learn occlusion-invariant image representations and image-shape correspondence. Specifically, we propose a hard occlusion example mining strategy to sample a hard image pair. Hereafter, to enforce the consistency between normal and occluded 2D images, we propose an Occlusion-invariant Image Consistency (OIC) based on hard image pairs, which gathers 2D image representations of the same instance while pushing away other 2D image representations. In addition, to prevent the 3D representations from perturbation by the occluded 2D representations, we design an Occlusion-invariant Correspondence Consistency (OCC) based on hard image pairs, which pulls the image-specific 3D shape embedding derived by attention mechanism close to the other 2D image representation of the same instance. The combination of OIC and OCC leads to accurate 2D-3D shape matching in challenging occluded scenarios. Our OPEN outperforms state-of-the-art methods by 6%~11% in terms of Top-1 retrieval accuracy on several representative benchmark datasets.
Fupeng Chu, Yang Cong, Ronghan Chen
IEEE Trans. Circuits Syst. Video Technol.3
2023 Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration
abstract
In comparison with most methods focusing on$3D$rigid object recognition and manipulation, deformable objects are more common in our real life but attract less attention. Generally, most existing methods for deformable object manipulation suffer two issues, 1) Massive demonstration: repeating thousands of robot-object demonstrations for model training of one specific instance; 2) Poor generalization: inevitably re-training for transferring the learned skill to a similar/new instance from the same category. Therefore, we propose a category-level deformable$3D$object manipulation framework, which could manipulate deformable$3D$objects with only one demonstration and generalize the learned skills to new similar instances without re-training. Specifically, our proposed framework consists of two modules. The Nocs State Transform$(NST)$module transfers the observed point clouds of the target to a pre-defined unified pose state (i.e.,Nocs state), which is the foundation for the category-level manipulation learning; the Neural Spatial Encoding$(NSE)$module generalizes the learned skill to novel instances by encoding the category-level spatial information to pursue the expected grasping point without re-training. The relative motion path is then planned to achieve autonomous manipulation. Both the simulated results via our$\text{Cap}_{40}$dataset and real robotic experiments justify the effectiveness of our framework.
Ronghan Chen, Yang Cong
CVPR2
2023 A Comprehensive Study of 3-D Vision-Based Robot Manipulation
abstract
Robot manipulation, for example, pick-and-place manipulation, is broadly used for intelligent manufacturing with industrial robots, ocean engineering with underwater robots, service robots, or even healthcare with medical robots. Most traditional robot manipulations adopt 2-D vision systems with plane hypotheses and can only generate 3-DOF (degrees of freedom) pose accordingly. To mimic human intelligence and endow the robot with more flexible working capabilities, 3-D vision-based robot manipulation has been studied. However, this task is still challenging in the open world especially for general object recognition and pose estimation with occlusion in cluttered backgrounds and human-like flexible manipulation. In this article, we propose a comprehensive analysis of recent progress about the 3-D vision for robot manipulation, including 3-D data acquisition and representation, robot-vision calibration, 3-D object detection/recognition, 6-DOF pose estimation, grasping estimation, and motion planning. We then present some public datasets, evaluation criteria, comparisons, and challenges. Finally, the related application domains of robot manipulation are given, and some future directions and open problems are studied as well.
Yang Cong, Ronghan Chen, Bingtao Ma, Hongsen Liu, Dongdong Hou, Chenguang Yang 0001
IEEE Trans. Cybern.2
2022 The Devil is in the Pose: Ambiguity-free 3D Rotation-invariant Learning via Pose-aware Convolution
abstract
Recent progress in introducing rotation invariance (RI) to 3D deep learning methods is mainly made by designing RI features to replace 3D coordinates as input. The key to this strategy lies in how to restore the global information that is lost by the input RI features. Most state-of-the-arts achieve this by incurring additional blocks or complex global representations, which is time-consuming and ineffective. In this paper, we real that the global information loss stems from an unexplored pose information loss problem, i.e., common convolution layers cannot capture the relative poses between RI features, thus hindering the global information to be hierarchically aggregated in the deep networks. To address this problem, we develop a Poseaware Rotation Invariant Convolution (i.e., PaRI-Conv), which dynamically adapts its kernels based on the relative poses. Specifically, in each PaRI-Conv layer, a lightweight Augmented Point Pair Feature (APPF) is designed to fully encode the RI relative pose information. Then, we propose to synthesize a factorized dynamic kernel, which reduces the computational cost and memory burden by decomposing it into a shared basis matrix and a pose-aware diagonal matrix that can be learned from the APPF. Extensive experiments on shape classification and part segmentation tasks show that our PaRI-Conv surpasses the state-of-the-art RI methods while being more compact and efficient.
Ronghan Chen, Yang Cong
CVPR1
2021 Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence
abstract
Shape correspondence from 3D deformation learning has attracted appealing academy interests recently. Nevertheless, current deep learning based methods require the supervision of dense annotations to learn per-point translations, which severely over-parameterize the deformation process. Moreover, they fail to capture local geometric details of original shape via global feature embedding. To address these challenges, we develop a new Unsupervised Dense Deformation Embedding Network (i.e., UD2E-Net), which learns to predict deformations between non-rigid shapes from dense local features. Since it is non-trivial to match deformation-variant local features for deformation prediction, we develop an Extrinsic-Intrinsic Autoencoder to first encode extrinsic geometric features from source into intrinsic coordinates in a shared canonical shape, with which the decoder then synthesizes corresponding target features. Moreover, a bounded maximum mean discrepancy loss is developed to mitigate the distribution divergence between the synthesized and original features. To learn natural deformation without dense supervision, we introduce a coarse parameterized deformation graph, for which a novel trace and propagation algorithm is proposed to improve both the quality and efficiency of the deformation. Our UD2E-Net outperforms state-of-the-art unsupervised methods by 24% on Faust Inter challenge and even supervised methods by 13% on Faust Intra challenge.
Ronghan Chen, Yang Cong, Jiahua Dong 0001
ICCV1