Shaowei Cui

dblp:221/1502 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-4750-3011ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HydroPalm: Dual-Mode Visual-Tactile Sensing for Underwater Humanoid Robot Hands
abstract
Underwater humanoid robots hold great potential for complex marine tasks thanks to their dexterous and versatile hands. However, their perception capabilities are severely hindered in turbid and low-light environments, where vision-only sensing becomes unreliable. To address this challenge, we present HydroPalm, the first bionic dual-modal visual-tactile sensor designed for hands of underwater humanoid robots. HydroPalm integrates a wide-field binocular vision module with a high-resolution soft tactile interface. Specifically, an iterative concentric angular topology sorting (ICATS) algorithm is proposed to resolve marker-matching ambiguity caused by background distortions. A contact-based refractive stereo ray tracing (CRSRT) method is introduced to perform accurate 3D reconstruction in water with variable refractive indices. Experiments across 0-1285.2 NTU demonstrate that HydroPalm improves reconstruction quality by over 230% compared to vision-only baselines, while maintaining a mean absolute error below 5% in waters with varying refractive indices. When deployed on a robotic hand, HydroPalm further enables reliable grasping inside a fully dark and highly turbid underwater cavity. The results suggest a new dual-modal sensing paradigm tailored for underwater humanoid robots, with promising applications in seafood harvesting, delicate ecological sampling, and archaeological excavation.
Shaowei Cui, Hongfei Chu, Min Tan 0001, Shuo Wang 0001, Yu Wang 0062
IEEE Trans Autom. Sci. Eng.2
2026 DexTac: Learning Contact-Aware Visuotactile Policies via Hand-by-Hand Teaching
abstract
For contact-intensive tasks, the ability to generate policies that produce comprehensive tactile-aware motions is essential. However, existing data collection and skill learning systems for dexterous manipulation often suffer from low-dimensional tactile information. To address this limitation, we propose DexTac, a visuo-tactile manipulation learning framework based on kinesthetic teaching. DexTac captures multi-dimensional tactile data—including contact force distributions and spatial contact regions—directly from human demonstrations. By integrating these rich tactile modalities into a policy network, the resulting contact-aware agent enables a dexterous hand to autonomously select and maintain optimal contact regions during complex interactions. We evaluate our framework on a challenging unimanual injection task. Experimental results demonstrate that DexTac achieves a 91.67% success rate. Notably, in high-precision scenarios involving small-scale syringes, our approach outperforms force-only baselines by 31.67%. These results underscore that learning multi-dimensional tactile priors from human demonstrations is critical for achieving robust, human-like dexterous manipulation in contact-rich environments.
Chaofan Zhang, Boyue Zhang 0002, Zhinan Peng, Shaowei Cui, Shuo Wang 0001
IEEE Trans Autom. Sci. Eng.5
2025 Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous Manipulations
abstract
Long-horizon dexterous tasks have been a long-standing problem in robotic manipulation. Previous studies have developed task and motion planning, imitation learning, and reinforcement learning methods for long-horizon manipulations. However, these methods are hard to achieve efficient planning for new tasks. Empowered with the Vision Language Model (VLM), recent studies significantly improve the generalization of robot systems. However, these works are only verified in simple pick-and-place tasks due to limited skills. To this end, we propose the Learn-Gen-Plan (LGP), which combines the VLM and learning-based primitives to endow robots with the ability to efficiently plan and complete various long-horizon dexterous tasks. LGP contains two key phases: skill generation and task planning. In skill generation, the Skill Generator is proposed to utilize the learned key primitives and hand-crafted trivial primitives to generate adaptive robot skills. In task planning, the Multimodal Planner generates the robot plan based on image observation, generated skills, and text prompts. We set up a series of dexterous tasks (e.g., cable routing, peg-in-hole assembly) in a real-world lighting circuit wiring scenario to evaluate LGP. The experimental results show that LGP efficiently generates robot plans with learned skills, controlling the robot to complete various multi-step cable wiring tasks.
Peng Hao 0003, Shaowei Cui, Junhang Wei, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
IEEE Trans Autom. Sci. Eng.2
2025 Dexterity-Guided Dimensional Synthesis and Multi-Task Control for Fingertip Manipulation
abstract
The geometric parameters severely impact the performance of dexterous manipulation, but manually adjusting them is time-consuming. In this paper, we employ the dexterity-guided dimensional synthesis to explore the geometric design space of a robotic hand, and propose a multi-objective optimization framework to improve its manipulation dexterity. Specifically, we first establish a screw-based mathematical model to describe the hand-object system’s kinetostatic properties. Three objectives are specified to achieve Pareto optimality: maximizing the system’s feasible position space, feasible orientation space, and manipulation stability. Furthermore, to validate the optimal design parameters in practice, we develop a multi-task object motion controller and apply it to a letter handwriting task. Finally, the optimized results are quantitatively analyzed by simulation, yielding a Pareto front and 112 optimal solutions. The applicability of the optimal parameters is assessed through a letter “O” handwriting experiment. Results show that the optimized hand can write the “O” with a maximum radius of 30.0 mm, which is 27.12% larger than that of the non-optimized hand. This controller is also used to schedule subtasks with varying priorities to avoid constrained regions. In contrast to the neural network-based PID controller, the designed controller prioritizes the task’s critical parts, ensuring the legibility of the letters.
Congjia Su, Rui Wang 0031, Shaowei Cui, Shuo Wang 0001
IEEE Trans Autom. Sci. Eng.3
2025 GelStereo Tip: A Spherical Fingertip Visuotactile Sensor for Multi-Finger Screwing Manipulation
abstract
Dexterous hands are the key element for robots to achieve human-like manipulation capabilities. An outstanding challenge is to provide fingertips of dexterous hands with precise tactile deformation sensing capabilities. In this paper, we present the GelStereo Tip, a spherical and easy-to-integrate GelStereo-type visuotactile sensor capable of sensing high-resolution 3D elastomer deformation. Previous calibration method does not take into account the impact of imaging errors caused by the sensor’s compact and high-curvature structural characteristics on the accuracy of tactile sensing. Therefore, we propose a novel self-calibration method based on the Refractive Stereo Ray Tracing model, named GTSC, and demonstrate the accuracy of less than 0.3 mm for deformation sensing. Furthermore, we also propose a Contact Retention Tactile Controller to address the issue of fingertips being unable to overcome obstructive torque during the multi-finger bottle cap screwing. After integrating GelStereo Tip into fingertips of Allegro Hand, the controller adjusts the joint positions of the given trajectory using proportional control based on the difference between the sensor’s actual deformation and the reference state for contact retention. We believe that the GelStereo Tip sensor combined with robotic dexterous hands has great application potential in the field of multi-finger fingertip manipulation. Note to Practitioners—The motivation of this paper is to design a fingertip visuotactile sensor with high-precision 3D tactile deformation sensing capabilities for multi-finger robotic hands and to validate its sensing performance. Additionally, it aims to address the issue of overcoming resistance in multi-finger screwing manipulations. Currently, most sensors do not consider the refraction effect or ignore the impact of planar imaging errors in refractive calibration. This paper proposes a visuotactile sensor along with a corresponding self-calibration method to ensure its sensing accuracy. Experiments show that our sensor possesses high-precision and robust 3D deformation sensing capabilities. On the other hand, multi-finger hands often struggle to complete screwing tasks along the given trajectory due to disturbances from torque resistance. This paper proposes a tactile controller that evaluates the contact state through aforementioned tactile sensing to improve subsequent trajectory and achieve continuous screwing. Comparative experiments highlight the necessity of this controller and the reliability of tactile sensing. We hope that the design of our sensor, the self-calibration method, and the tactile controller applied to multi-finger screwing can provide new insights for other practitioners.
Boyue Zhang 0002, Shaowei Cui, Chaofan Zhang, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001
IEEE Trans Autom. Sci. Eng.2
2025 TacFlex: Multimode Tactile Imprints Simulation for Visuotactile Sensors With Coating Patterns
abstract
Visuotactile sensors have been shown to provide rich contact information for robots. However, how to build a high-fidelity visuotactile simulator that supports multi-mode tactile imprints and various sensor configurations (such as coating patterns) remains a challenging problem. In this paper, we present TacFlex, an efficient and flexible simulator for visuotactile sensors, which physically simulates the elastomer deformation using Finite Element Methods (FEM), and focuses on linking the deformed elastomer mesh to diverse tactile imprints, including tactile images with arbitrary coating patterns and tactile 3D point clouds. We further propose a ray tracing-based rectification method to deal with multi-medium refraction effects to make the simulated tactile images more realistic. Extensive qualitative and quantitative experiments are conducted to demonstrate the effectiveness of TacFlex on several visuotactile sensors. Furthermore, we explore the Sim2Real performance of different tactile imprints provided by TacFlex in tactile perception and manipulation tasks, such as cylindrical object pose estimation and peg-in-hole. The perception/policy models trained in simulation are successfully deployed in the real world. Finally, we present the outlook on the potential of TacFlex in visuotactile manipulation learning. The TacFlex simulator is open-sourced to the community. See supplementary video, code, and results athttps://sites.google.com/view/tacflex/.
Chaofan Zhang, Shaowei Cui, Jingyi Hu, Tiandong Zhang, Rui Wang 0031, Shuo Wang 0001
IEEE Trans. Robotics2
2025 FlowSight: Vision-Based Artificial Lateral Line Sensor for Water Flow Perception
abstract
This paper presents a novel vision-based artificial lateral line (ALL) sensor, FlowSight, enhancing the perception capabilities of underwater robots. Through an autonomous vision system, FlowSight allows for simultaneous sensing the speed and direction of local water flow without relying on external auxiliary equipment. Inspired by the lateral line neuromast of fish, a flexible bionic tentacle is designed to sense water flow. Deformation and motion characteristics of the tentacle are modeled and analyzed using bidirectional fluid-structure interaction (FSI) simulation. Upon contact with water flow, the tentacle converts water flow information into elastic deformation information, which is captured and processed into an image sequence by the autonomous vision system. Subsequently, a water flow perception method based on deep neural networks is proposed to estimate the flow speed and direction from the captured image sequence. The perception network is trained and tested using data collected from practical experiments conducted in a controllable swim tunnel. Finally, the FlowSight sensor is integrated into the bionic underwater robot RoboDact, and a closed-loop motion control experiment based on water flow perception is conducted. Experiments conducted in the swim tunnel and water pool demonstrate the feasibility and effectiveness of FlowSight sensor and the water flow perception method.
Tiandong Zhang, Rui Wang 0031, Qiyuan Cao, Shaowei Cui, Gang Zheng 0002, Shuo Wang 0001
IEEE Trans. Robotics4
2024 Learning-Based Slip Detection for Dexterous Manipulation Using GelStereo Sensing
abstract
Endowing the robot with tactile perception can effectively improve manipulation dexterity, along with various benefits of human-like touch. Using GelStereo (GS) tactile sensing, which gives high-resolution contact geometry information, including 2-D displacement field, and 3-D point cloud of the contact surface, we present a learning-based slip detection system in this study. The results reveal that the well-trained network achieves 95.79% accuracy on the never-seen testing dataset, which surpasses the current model-based and learning-based methods using visuotactile sensing. We also propose a general framework for slip feedback adaptive control for dexterous robot manipulation tasks. The experimental results show the effectiveness and efficiency of the proposed control framework using GS tactile feedback when deployed on real-world grasping and screwing manipulation tasks on various robot setups.
Shaowei Cui, Shuo Wang 0001, Rui Wang 0031, Chaofan Zhang
IEEE Trans. Neural Networks Learn. Syst.1
2023 GelStereo Palm: A Novel Curved Visuotactile Sensor for 3-D Geometry Sensing
abstract
Recently, visuotactile sensors have shown promising potential in robotics due to their high-resolution sensing ability. Unfortunately, the majority of available visuotactile sensors are limited to flat shapes, which severely limits their application possibilities. In this article, we propose a novel curved visuotactile sensor, the GelStereo Palm, which senses the 3-D contact geometry on a curved surface using a binocular vision system. Meanwhile, to solve the light refraction problem in the binocular stereo vision system under a curved medium, a refractive stereo ray tracing model for GelStereo Palm is presented. Moreover, a 3-D tactile point cloud sensing pipeline is introduced to reconstruct the 3-D contact geometry in real-time. Finally, extensive experiments are conducted to verify the accuracy and robustness of the 3-D contact geometry sensing of our GelStereo Palm sensor.
Jingyi Hu, Shaowei Cui, Shuo Wang 0001, Chaofan Zhang, Rui Wang 0031, Lipeng Chen
IEEE Trans. Ind. Informatics2
2022 Learning-based Six-axis Force/Torque Estimation Using GelStereo Fingertip Visuotactile Sensing
abstract
Visuotactile sensors have recently attracted much attention in robot communities due to the benefit of high spatial resolution sensing. However, force/torque estimation by visuotactile sensors remains a challenging problem. In this paper, we propose a learning-based six-axis force/torque estimation network using GelStereo visuotactile sensor, which can provide two-dimensional (2D) and three-dimensional (3D) displacements of markers embedded in the sensor surface. The convolutional neural networks are employed to extract multi-modal tactile deformation features; and a novel contact positional encoding method is proposed to eliminate the influence of translation invariance in convolutional operators. The well-trained model achieves the best RMSE of 0.290 N in force and 0.0084 Nm in torque. Furthermore, the proposed force/torque estimation network is integrated with a force-feedback policy for adaptive grasping tasks. The experimental results demonstrate the effectiveness of the proposed method and its potential application in robotic grasping and manipulation tasks.
Chaofan Zhang, Shaowei Cui, Yinghao Cai, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001
IROS2
2022 Multimodal Unknown Surface Material Classification and Its Application to Physical Reasoning
abstract
Unknown surface material classification (SMC) can inform a robot about material properties, enabling it to interact with environments appropriately. Recent research has leveraged multimodal data using deep learning to improve the performance of SMC. In this article, we present a deep learning model, multimodal temporal convolutional neural network (MTCNN), which integrates energy spectrum, dilated convolutions, and sequence poolings into a unified network architecture. The proposed model can learn material representations from auditory and multitactile (i.e., acceleration, normal force, and friction force) data generated by dragging a tool along surfaces, and distinguish unknown object surface materials into categories. For surface material data collection, a tool is also designed to detect different object surfaces. The performance of MTCNN is evaluated on a public dataset and the highest classification accuracy is 87.55%. A robotic curling example is provided to illustrate how the presented model helps the robot in manipulation.
Junhang Wei, Shaowei Cui, Jingyi Hu, Peng Hao 0003, Shuo Wang 0001, Zheng Lou
IEEE Trans. Ind. Informatics2
2020 Grasp State Assessment of Deformable Objects Using Visual-Tactile Fusion Perception
abstract
Humans can quickly determine the force required to grasp a deformable object to prevent its sliding or excessive deformation through vision and touch, which is still a challenging task for robots. To address this issue, we propose a novel 3D convolution-based visual-tactile fusion deep neural network (C3D-VTFN) to evaluate the grasp state of various deformable objects in this paper. Specifically, we divide the grasp states of deformable objects into three categories of sliding, appropriate and excessive. Also, a dataset for training and testing the proposed network is built by extensive grasping and lifting experiments with different widths and forces on 16 various deformable objects with a robotic arm equipped with a wrist camera and a tactile sensor. As a result, a classification accuracy as high as 99.97% is achieved. Furthermore, some delicate grasp experiments based on the proposed network are implemented in this paper. The experimental results demonstrate that the C3D-VTFN is accurate and efficient enough for grasp state assessment, which can be widely applied to automatic force control, adaptive grasping, and other visual-tactile spatiotemporal sequence learning problems.
Shaowei Cui, Rui Wang 0031, Junhang Wei, Fanrong Li, Shuo Wang 0001
ICRA1