Rui Chen 0019

dblp:02/1003-19 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-4041-4131ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021
YearPublicationVenuePosition
2025 AllTact Fin Ray: A Compliant Robot Gripper With Omni-Directional Tactile Sensing
abstract
Tactile sensing plays a crucial role in robot grasping and manipulation by providing essential contact information between the robot and the environment. In this paper, we present AllTact Fin Ray, a novel compliant gripper design with omni-directional and local tactile sensing capabilities. The finger body is unibody-casted using transparent elastic silicone, and a camera positioned at the base of the finger captures the deformation of the whole body and the contact face. Due to the global deformation of the adaptive structure, existing vision-based tactile sensing approaches that assume constant illumination are no longer applicable. To address this, we propose a novel sensing method where the global deformation is first reconstructed from the image using edge features and spatial constraints. Then, detailed contact geometry is computed from the brightness difference against a dynamically retrieved reference image. Extensive experiments validate the effectiveness of our proposed gripper design and sensing method in contact detection, force estimation, object grasping, and precise manipulation.
Siwei Liang, Jing Xu 0011, Hongyu Qian, Xiangjun Zhang, Dan Wu 0008, Wenbo Ding 0001, Rui Chen 0019
IEEE Trans Autom. Sci. Eng.8
2025 GlassMolder: Transparent Object Reconstruction With Silhouette-Guided Object-Centric Diffusion
abstract
Depth reconstruction for transparent objects is a challenging problem, where surface feature matching methods are hindered by complex refraction and reflection. Existing learning-based reconstruction methods by regressing or completing depth maps for entire scenes are data-costly and lack generalization in different environments. To solve this problem, we propose a novel transparent object reconstruction pipeline with a guided object-centric 3D diffusion model. Specifically, we train an unconditional 3D diffusion model with only 3D point cloud data. To control the output of the diffusion model, we design a silhouette-based guidance function and a completion framework with outline points for each step of diffusion process. Specifically, for each step, we design a re-projection pipeline to estimate a silhouette with uncertainty and constrain the partially-noised point cloud to align with it. We further apply stereo matching to compute the outline points in the stereo silhouettes and use a completion framework to fuse them with the partially-denoised point cloud. Finally, we transform the transparent objects to the world frame by applying the transformation from pose estimation. Experiment results show that our method can achieve state-of-the-art performance for transparent object depth reconstruction compared to existing depth regression and completion methods.
Changping Hu, Jing Xu 0011, Chifai Pun, Fei Chen 0007, Rui Chen 0019
IEEE Trans. Circuits Syst. Video Technol.5
2025 DexSim2Real$^{\mathbf{2}}$: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation
abstract
Articulated objects are ubiquitous in daily life. In this paper, we present DexSim2Real$^{\mathbf{2}}$, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of unseen articulated objects through active interactions, which enables sampling-based model predictive control to plan trajectories achieving different goals without requiring demonstrations or RL. It first predicts an interaction using an affordance network trained on self-supervised interaction data or videos of human manipulation. After executing the interactions on the real robot to move the object parts, we propose a novel modeling pipeline based on 3D AIGC to build a digital twin of the object in simulation from multiple frames of observations. For dexterous hands, we utilize eigengrasp to reduce the action dimension, enabling more efficient trajectory searching. Experiments validate the framework's effectiveness for precise manipulation using a suction gripper, a two-finger gripper and two dexterous hands. The generalizability of the explicit world model also enables advanced manipulation strategies like manipulating with tools.
Taoran Jiang, Liqian Ma, Jing Xu 0011, Jiaojiao Meng, Weihang Chen, Zecui Zeng, Lusong Li, Dan Wu 0008, Rui Chen 0019
IEEE Trans. Robotics10
2025 ThinTact: Thin Vision-Based Tactile Sensor by Lensless Imaging
abstract
Vision-based tactile sensors have drawn increasing interest in the robotics community. However, traditional lens-based designs impose minimum thickness constraints on these sensors, limiting their applicability in space-restricted settings. In this article, we propose ThinTact, a novel lensless vision-based tactile sensor with a sensing field of over 200 mm${}^{2}$and a thickness of less than 10 mm. ThinTact utilizes the mask-based lensless imaging technique to map the contact information to CMOS signals. To ensure real-time tactile sensing, we propose a real-time lensless reconstruction algorithm that leverages a frequency-spatial-domain joint filter based on discrete cosine transform. This algorithm achieves computation significantly faster than existing optimization-based methods. In addition, to improve the sensing quality, we develop a mask optimization method based on the generic algorithm and the corresponding system matrix calibration algorithm. We evaluate the performance of our proposed lensless reconstruction and tactile sensing through qualitative and quantitative experiments. Furthermore, we demonstrate ThinTact's practical applicability in diverse applications, including texture recognition and contact-rich object manipulation.
Jing Xu 0011, Weihang Chen, Hongyu Qian, Dan Wu 0008, Rui Chen 0019
IEEE Trans. Robotics5
2024 GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
abstract
This paper presents GenH2R, a framework for learning generalizable vision-based human-to-robot (H2R) handover skills. The goal is to equip robots with the ability to reliably receive objects with unseen geometry handed over by humans in various complex trajectories. We acquire such generalizability by learning H2R handover at scale with a comprehensive solution including procedural simulation assets creation, automated demonstration generation, and effective imitation learning. We leverage large-scale 3D model repositories, dexterous grasp generation methods, and curve-based 3D animation to create an H2R handover simulation environment named GenH2R-Sim, surpassing the number of scenes in existing simulators by three orders of magnitude. We further introduce a distillation-friendly demonstration generation method that automati-cally generates a million high-quality demonstrations suitable for learning. Finally, we present a 4D imitation learning method augmented by a future forecasting objective to distill demonstrations into a visuo-motor handover policy. Experimental evaluations in both simulators and the real world demonstrate significant improvements (at least +10% success rate) over baselines in all cases.
Junyu Chen 0003, Ziqing Chen, Pengwei Xie, Rui Chen 0019, Li Yi 0001
CVPR5
2024 General-Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors
abstract
Visuotactile sensors can provide rich contact information, having great potential in contact-rich manipulation tasks with reinforcement learning (RL) policies. Sim2Real technique tackles the challenge of RL's reliance on a large amount of interaction data. However, most Sim2Real methods for manipulation tasks with visuotactile sensors rely on rigid-body physics simulation, which fails to simulate the real elastic deformation precisely. Moreover, these methods do not exploit the characteristic of tactile signals for designing the network architecture. In this paper, we build a general-purpose Sim2Real protocol for manipulation policy learning with marker-based visuotactile sensors. To improve the simulation fidelity, we employ an FEM-based physics simulator that can simulate the sensor deformation accurately and stably for arbitrary geometries. We further propose a novel tactile feature extraction network that directly processes the set of pixel coordinates of tactile sensor markers and a self-supervised pre-training strategy to improve the efficiency and generalizability of RL policies. We conduct extensive Sim2Real experiments on the peg-in-hole task to validate the effectiveness of our method. And we further show its generalizability on additional tasks including plug adjustment and lock opening. The protocol, including the simulator and the policy learning framework, will be open-sourced for community usage.
Weihang Chen, Jing Xu 0011, Fanbo Xiang, Xiaodi Yuan, Hao Su 0001, Rui Chen 0019
IEEE Trans. Robotics6
2023 ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Jiayuan Gu, Fanbo Xiang, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen 0019, Hao Su 0001
ICLR14
2023 Sim2Real2: Actively Building Explicit Physics Model for Precise Articulated Object Manipulation
abstract
Accurately manipulating articulated objects is a challenging yet important task for real robot applications. In this paper, we present a novel framework called Sim2Real2to enable the robot to manipulate an unseen articulated object to the desired state precisely in the real world with no human demonstrations. We leverage recent advances in physics simulation and learning-based perception to build the interactive explicit physics model of the object and use it to plan a long-horizon manipulation trajectory to accomplish the task. However, the interactive model cannot be correctly estimated from a static observation. Therefore, we learn to predict the object affordance from a single-frame point cloud, control the robot to actively interact with the object with a one-step action, and capture another point cloud. Further, the physics model is constructed from the two point clouds. Experimental results show that our framework achieves about 70% manipulations with < 30% relative error for common articulated objects, and 30% manipulations for difficult objects. Our proposed framework also enables advanced manipulation strategies, such as manipulating with different tools. Code and videos are available on our project webpage: https://ttimelord.github.io/Sim2Real2-site/
Liqian Ma, Jiaojiao Meng, Shuntao Liu, Weihang Chen, Jing Xu 0011, Rui Chen 0019
ICRA6
2023 TransTouch: Learning Transparent Objects Depth Sensing Through Sparse Touches
abstract
Transparent objects are common in daily life. However, depth sensing for transparent objects remains a challenging problem. While learning-based methods can leverage shape priors to improve the sensing quality, the labor-intensive data collection in real world and the sim-to-real domain gap restrict these methods' scalability. In this paper, we propose a method to finetune a stereo network with sparse depth labels automatically collected using a probing system with tactile feedback. We present a novel utility function to evaluate the benefit of touches. By approximating and optimizing the utility function, we can optimize the probing locations given a fixed touching budget to better improve the network's performance on real objects. We further combine tactile depth supervision with a confidence-based regularization to prevent over-fitting during finetuning. To evaluate the effectiveness of our method, we construct a real-world dataset including both diffuse and transparent objects. Experimental results on this dataset show that our method can significantly improve real-world depth sensing accuracy, especially for transparent objects.
Liuyu Bian, Pengyang Shi, Weihang Chen, Jing Xu 0011, Li Yi 0001, Rui Chen 0019
IROS6
2023 ActiveZero++: Mixed Domain Learning Stereo and Confidence-Based Depth Completion With Zero Annotation
abstract
Learning-based stereo methods usually require a large scale dataset with depth, however obtaining accurate depth in the real domain is difficult, but groundtruth depth is readily available in the simulation domain. In this article we propose a new framework, ActiveZero++, which is a mixed domain learning solution for active stereovision systems that requires no real world depth annotation. In the simulation domain, we use a combination of supervised disparity loss and self-supervised loss on a shape primitives dataset. By contrast, in the real domain, we only use self-supervised loss on a dataset that is out-of-distribution from either training simulation data or test real data. To improve the robustness and accuracy of our reprojection loss in hard-to-perceive regions, our method introduces a novel self-supervised loss called temporal IR reprojection. Further, we propose the confidence-based depth completion module, which uses the confidence from the stereo network to identify and improve erroneous areas in depth prediction through depth-normal consistency. Extensive qualitative and quantitative evaluations on real-world data demonstrate state-of-the-art results that can even outperform a commercial depth sensor. Furthermore, our method can significantly narrow the Sim2Real domain gap of depth maps for state-of-the-art learning based 6D pose estimation algorithms.
Rui Chen 0019, Isabella Liu, Edward Yang, Jianyu Tao, Xiaoshuai Zhang, Qing Ran, Jing Xu 0011, Hao Su 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 A Hierarchical Compliance-Based Contextual Policy Search for Robotic Manipulation Tasks With Multiple Objectives
abstract
Contextual policy search methods have demonstrated the potential to acquire robotic skill generalization on trajectory-shaping-based tasks. However, it is still challenging for robotic contact-rich manipulation tasks because contact force regulation, reference trajectory adaptation, and task generalization must be fulfilled simultaneously. To this end, a hierarchical compliance-based contextual policy search (HC-CPS) approach is proposed to learn the robotic compliant skills for force, motion, and task adaptation. Specifically, the parameterized impedance-conditioned action space is proposed for reinforcement learning lower-level policy to obtain the compliance for reference motion regulation and contact force control, while a linear Gaussian contextual policy is formulated as the higher-level policy to optimize the context-conditioned impedance parameters for task generalization; therefore, a family of contact-rich manipulation tasks with multiple objectives is achieved. Moreover, data efficiency is further improved by two aspects: first, a variation encoder-decoder model is proposed to estimate the underlying constraints of impedance parameters over the actions, leading to the mitigated extrapolation error for lower-level policy off-policy learning; second, a composite forward model is proposed to generate artificial trajectories and reduce the reward bias for higher-level contextual policy learning. The HC-CPS approach is validated by three simulated manipulation tasks and the real-world dual peg-in-hole assembly tasks with two kinds of objectives, and the results demonstrate the effectiveness of HC-CPS.
Zhimin Hou, Rui Chen 0019, Pingfa Feng, Jing Xu 0011
IEEE Trans. Ind. Informatics3
2023 Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation
abstract
In this article, we focus on the simulation of active stereovision depth sensors, which are popular in both academic and industry communities. Inspired by the underlying mechanism of the sensors, we designed a fully physics-grounded simulation pipeline that includes material acquisition, ray-tracing-based infrared (IR) image rendering, IR noise simulation, and depth estimation. The pipeline is able to generate depth maps with material-dependent error patterns similar to a real depth sensor in real time. We conduct real experiments to show that perception algorithms and reinforcement learning policies trained in our simulation platform could transfer well to the real-world test cases without any fine-tuning. Furthermore, due to the high degree of realism of this simulation, our depth sensor simulator can be used as a convenient testbed to evaluate the algorithm performance in the real world, which will largely reduce the human effort in developing robotic algorithms. The entire pipeline has been integrated into the SAPIEN simulator and is open-sourced to promote the research of vision and robotics communities.
Xiaoshuai Zhang, Rui Chen 0019, Ang Li 0010, Fanbo Xiang, Yuzhe Qin, Jiayuan Gu, Zhan Ling, Minghua Liu, Peiyu Zeng, Songfang Han, Zhiao Huang, Tongzhou Mu, Jing Xu 0011, Hao Su 0001
IEEE Trans. Robotics2
2022 ActiveZero: Mixed Domain Learning for Active Stereovision with Zero Annotation
abstract
Traditional depth sensors generate accurate real world depth estimates that surpass even the most advanced learning approaches trained only on simulation domains. Since ground truth depth is readily available in the simulation domain but quite difficult to obtain in the real domain, we propose a method that leverages the best of both worlds. In this paper we present a new framework, ActiveZero, which is a mixed domain learning solution for active stereovision systems that requires no real world depth annotation. First, we demonstrate the transferability of our method to out-of-distribution real data by using a mixed domain learning strategy. In the simulation domain, we use a combination of supervised disparity loss and self-supervised losses on a shape primitives dataset. By contrast, in the real domain, we only use self-supervised losses on a dataset that is out-of-distribution from either training simulation data or test real data. Second, our method introduces a novel self-supervised loss called temporal IR reprojection to increase the robustness and accuracy of our reprojections in hard-to-perceive regions. Finally, we show how the method can be trained end-to-end and that each module is important for attaining the end result. Extensive qualitative and quantitative evaluations on real data demonstrate state of the art results that can even beat a commercial depth sensor. The codes of ActiveZero are available at: httis://github.com/haosulab/active_zero.
Isabella Liu, Edward Yang, Jianyu Tao, Rui Chen 0019, Xiaoshuai Zhang, Qing Ran, Hao Su 0001
CVPR4
2022 Ellipse detection using the edges extracted by deep learning
Chicheng Liu, Rui Chen 0019, Ken Chen 0002, Jing Xu 0011
Mach. Vis. Appl.2
2021 Visibility-Aware Point-Based Multi-View Stereo Network
abstract
We introduce VA-Point-MVSNet, a novel visibility-aware point-based deep framework for multi-view stereo (MVS). Distinct from existing cost volume approaches, our method directly processes the target scene as point clouds. More specifically, our method predicts the depth in a coarse-to-fine manner. We first generate a coarse depth map, convert it into a point cloud and refine the point cloud iteratively by estimating the residual between the depth of the current iteration and that of the ground truth. Our network leverages 3D geometry priors and 2D texture information jointly and effectively by fusing them into a feature-augmented point cloud, and processes the point cloud to estimate the 3D flow for each point. This point-based architecture allows higher accuracy, more computational efficiency and more flexibility than cost-volume-based counterparts. Furthermore, our visibility-aware multi-view feature aggregation allows the network to aggregate multi-view appearance cues while taking into account visibility. Experimental results show that our approach achieves a significant improvement in reconstruction quality compared with state-of-the-art methods on the DTU and the Tanks and Temples dataset. The code of VA-Point-MVSNet proposed in this work will be released at https://github.com/callmeray/PointMVSNet.
Rui Chen 0019, Songfang Han, Jing Xu 0011, Hao Su 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Normal Assisted Stereo Depth Estimation
abstract
Accurate stereo depth estimation plays a critical role in various 3D tasks in both indoor and outdoor environments. Recently, learning-based multi-view stereo methods have demonstrated competitive performance with limited number of views. However, in challenging scenarios, especially when building cross-view correspondences is hard, these methods still cannot produce satisfying results. In this paper, we study how to enforce the consistency between surface normal and depth at training time to improve the performance. We couple the learning of a multi-view normal estimation module and a multi-view depth estimation module. In addition, we propose a novel consistency loss to train an independent consistency module that refines the depths from depth/normal pairs. We find that the joint learning can improve both the prediction of normal and depth, and the accuracy and smoothness can be further improved by enforcing the consistency. Experiments on MVS, SUN3D, RGBD and Scenes11 demonstrate the effectiveness of our method and state-of-the-art performance.
Uday Kusupati, Rui Chen 0019, Hao Su 0001
CVPR3
2019 Point-Based Multi-View Stereo Network
abstract
We introduce Point-MVSNet, a novel point-based deep framework for multi-view stereo (MVS). Distinct from existing cost volume approaches, our method directly processes the target scene as point clouds. More specifically, our method predicts the depth in a coarse-to-fine manner. We first generate a coarse depth map, convert it into a point cloud and refine the point cloud iteratively by estimating the residual between the depth of the current iteration and that of the ground truth. Our network leverages 3D geometry priors and 2D texture information jointly and effectively by fusing them into a feature-augmented point cloud, and processes the point cloud to estimate the 3D flow for each point. This point-based architecture allows higher accuracy, more computational efficiency and more flexibility than cost-volume-based counterparts. Experimental results show that our approach achieves a significant improvement in reconstruction quality compared with state-of-the-art methods on the DTU and the Tanks and Temples dataset. Our source code and trained models are available at https://github.com/callmeray/PointMVSNet.
Rui Chen 0019, Songfang Han, Jing Xu 0011, Hao Su 0001
ICCV1
2017 Set space visual servoing of a 6-DOF manipulator
abstract
This article develops a set space visual servoing method that is quiet different from state-of-the-art approaches. Our approach does not require complex image processing techniques for the extraction, matching and tracking of image features. Instead, it only requires a simple matching algorithm and builds visual errors in set space. Each error is mainly related to one degree of freedom of the camera; therefore, we can design a decoupled control law. This control law is robust and does not require calibrated inner parameters of the camera. Our approach has been validated in 4-degree-of-freedom (DOF) visual servoing simulations with common image patterns and 6-DOF visual servoing experiments with specific image patterns. These visual servoing tasks are properly achieved even when partial occlusions occur.
Chicheng Liu, Rui Chen 0019, Jing Xu 0011, Heping Chen, Ning Xi 0001, Ken Chen 0002
ICRA2