VLDB 2026 Research / reviewers in the wild / expert
Shuo Wang 0001
dblp:63/1591-1
· DBLP profile ↗
61ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0002-1390-9219ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 18 since 2021Systems, architecture and hardware · 17 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HydroPalm: Dual-Mode Visual-Tactile Sensing for Underwater Humanoid Robot HandsabstractUnderwater humanoid robots hold great potential for complex marine tasks thanks to their dexterous and versatile hands. However, their perception capabilities are severely hindered in turbid and low-light environments, where vision-only sensing becomes unreliable. To address this challenge, we present HydroPalm, the first bionic dual-modal visual-tactile sensor designed for hands of underwater humanoid robots. HydroPalm integrates a wide-field binocular vision module with a high-resolution soft tactile interface. Specifically, an iterative concentric angular topology sorting (ICATS) algorithm is proposed to resolve marker-matching ambiguity caused by background distortions. A contact-based refractive stereo ray tracing (CRSRT) method is introduced to perform accurate 3D reconstruction in water with variable refractive indices. Experiments across 0-1285.2 NTU demonstrate that HydroPalm improves reconstruction quality by over 230% compared to vision-only baselines, while maintaining a mean absolute error below 5% in waters with varying refractive indices. When deployed on a robotic hand, HydroPalm further enables reliable grasping inside a fully dark and highly turbid underwater cavity. The results suggest a new dual-modal sensing paradigm tailored for underwater humanoid robots, with promising applications in seafood harvesting, delicate ecological sampling, and archaeological excavation. Shaowei Cui, Hongfei Chu, Min Tan 0001, Shuo Wang 0001, Yu Wang 0062 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | DexTac: Learning Contact-Aware Visuotactile Policies via Hand-by-Hand TeachingabstractFor contact-intensive tasks, the ability to generate policies that produce comprehensive tactile-aware motions is essential. However, existing data collection and skill learning systems for dexterous manipulation often suffer from low-dimensional tactile information. To address this limitation, we propose DexTac, a visuo-tactile manipulation learning framework based on kinesthetic teaching. DexTac captures multi-dimensional tactile data—including contact force distributions and spatial contact regions—directly from human demonstrations. By integrating these rich tactile modalities into a policy network, the resulting contact-aware agent enables a dexterous hand to autonomously select and maintain optimal contact regions during complex interactions. We evaluate our framework on a challenging unimanual injection task. Experimental results demonstrate that DexTac achieves a 91.67% success rate. Notably, in high-precision scenarios involving small-scale syringes, our approach outperforms force-only baselines by 31.67%. These results underscore that learning multi-dimensional tactile priors from human demonstrations is critical for achieving robust, human-like dexterous manipulation in contact-rich environments. Chaofan Zhang, Boyue Zhang 0002, Zhinan Peng, Shaowei Cui, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | HOCOpt: Hand-Object Contact Optimization to Improve Pose Estimation in Physical InteractionsabstractReconstructing hand–object physical interaction through visual sensing is crucial to understanding human intentions and guaranteeing the safety of human–robot collaboration in industrial applications. Due to heavy occlusion and cluttered backgrounds, existing methods generate inaccurate pose estimation, leading to unrealistic physical interactions between hands and objects. In this article, we present a novel contact-driven pose optimization framework called hand–object contact optimization (HOCOpt) to achieve accurate hand–object pose estimation. The HOCOpt includes two parts: contact estimation and pose optimization. In the contact estimation, we propose a contact region estimation network (CREN) to predict the potential contact across the hand–object meshes with inaccurate poses. A novel contact entropy weight and an auxiliary network are introduced to the training process of CREN to accelerate the model learning and improve the prediction accuracy. For pose optimization, a two-stage hand–object pose optimization method is utilized to refine inaccurate poses by considering both contact distribution and contact stability. During optimization, an orientation-aware differentiable contact model is introduced to account for hand deformation and contact forces to achieve accurate contact modeling. Extensive experiments on ContactPose, HO3D, and DexYCB datasets show that our approach outperforms the existing baselines. Besides, experiments on physical interaction tasks for human–robot collaboration are conducted to demonstrate the practical significance of HOCOpt in industrial scenarios. Xiaoge Cao, Tao Lu 0006, Wenhao Yu 0011, Yinghao Cai, Shuo Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | SAFT: Real-Time Tracking and Mapping With Self-Supervised Robust Stereo Matching for Underwater VehiclesabstractRobust and efficient tracking and mapping are critical for underwater vehicles, but remain challenging due to degraded visual quality, ambiguous features, and limited computational resources. Although recent deep learning-based stereo matching methods have significantly improved geometric perception for robots, most existing approaches struggle to simultaneously achieve high speed and strong generalization. To address these challenges, we propose SAFT, a tracking and mapping framework based on self-supervised, robust, and real-time stereo matching. SAFT introduces three key innovations: 1) SAFT-Stereo, a novel stereo matching network that integrates cost aggregation with iterative optimization to enable efficient disparity estimation in feature-sparse regions; 2) a spatiotemporal self-supervised loss that leverages both spatial and temporal constraints to provide stable training signals in textureless regions; and 3) SAFT-DSOL, a real-time tracking and mapping algorithm that integrates the self-supervised models to achieve robust localization and dense reconstruction. Extensive experiments on both public and custom underwater datasets demonstrate that SAFT-Stereo achieves the best generalization performance among all real-time methods, while requiring only 1/6 of the inference time of RT-IGEV++. Moreover, the proposed SAFT-DSOL enables stable and efficient tracking and achieves real-time dense reconstruction in indoor shipwreck scenarios. The code is available at github.com/c237814486/SAFT-Stereo. Yaozhong Cao, Xiaolong Hui, Xuejian Bai, Yu Wang 0062, Shuo Wang 0001, Min Tan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | EA-Vit: Efficient Adaptation for Elastic Vision TransformerabstractVision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT. Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036 |
ICCV | 6 |
| 2025 | NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp DetectionabstractRobotic grasping in scenes with transparent and specular objects presents great challenges for methods relying on accurate depth information. In this paper, we introduce NeuGrasp, a neural surface reconstruction method that leverages background priors for material-agnostic grasp detection. NeuGrasp integrates transformers and global prior volumes to aggregate multi-view features with spatial encoding, enabling robust surface reconstruction in narrow and sparse viewing conditions. By focusing on foreground objects through residual feature enhancement and refining spatial perception with an occupancy-prior volume, NeuGrasp excels in handling objects with transparent and specular surfaces. Extensive experiments in both simulated and real-world scenarios show that NeuGrasp outperforms state-of-the-art methods in grasping while maintaining comparable reconstruction quality. More details are available at https://neugrasp.github.io/. Qingyu Fan, Yinghao Cai, Wenzhe He, Tao Lu 0006, Shuo Wang 0001 |
ICRA | 8 |
| 2025 | MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric GraspingabstractRobotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interaction between high-level and low-level features through the Insight Transformer, while the Empower Transformer selectively attends to the highest-level features, which synergistically strikes a balance between focusing on fine geometric details and overall geometric structures. Furthermore, MISCGrasp utilizes multi-scale contrastive learning to exploit similarities among positive grasp samples, ensuring consistency across multi-scale features. Extensive experiments in both simulated and real-world environments demonstrate that MISCGrasp outperforms baseline and variant methods in tabletop decluttering tasks. More details are available at https://miscgrasp.github.io/. Qingyu Fan, Yinghao Cai, Chunting Jiao, Tao Lu 0006, Shuo Wang 0001 |
IROS | 8 |
| 2025 | SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement LearningabstractPreference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel efficient query selection and preference-guided exploration method, called SENIOR, which could select the meaningful and easy-to-comparison behavior segment pairs to improve human feedback-efficiency and accelerate policy learning with the designed preference-guided intrinsic rewards. Our key idea is twofold: (1) We designed a Motion-Distinction-based Selection scheme (MDS). It selects segment pairs with apparent motion and different directions through kernel density estimation of states, which is more task-related and easy for human preference labeling; (2) We proposed a novel preference-guided exploration method (PGE). It encourages the exploration towards the states with high preference and low visits and continuously guides the agent achieving the valuable samples. The synergy between the two mechanisms could significantly accelerate the progress of reward and policy learning. Our experiments show that SENIOR outperforms other five existing methods in both human feedback-efficiency and policy convergence speed on six complex robot manipulation tasks from simulation and four real-worlds. Videos can be found on our project website: https://2025senior.github.io/ Hexian Ni, Tao Lu 0006, Haoyuan Hu, Yinghao Cai, Shuo Wang 0001 |
IROS | 5 |
| 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization
Jiangyong Yu, Sifan Zhou, Shuoyu Li, Shuo Wang 0001, Xing Hu 0010, Zukang Xu, Changyong Shu, Zhihang Yuan |
ACM Multimedia | 5 |
| 2025 | ViTac-Gripper: A Vision-Based Tactile Gripper With Enhanced Multi-Physical Field Perception for Underwater RobotsabstractTactile sensing is indispensable for underwater operations, as it provides critical feedback on contact information and environmental interactions. Although existing sensors, such as optical, piezoelectric devices, have been employed for underwater tactile perception, they exhibit some limitations in multi-physical field contact perception. To address these challenges, this study introduces the ViTac-Gripper, a vision-based tactile sensing underwater gripper designed to enhance underwater perception and grasping capabilities. The Finite Element Method (FEM) is utilized to simulate the deformation and force distribution on the tactile surface, providing a high-fidelity model for contact behavior analysis. A domain-aligned multi-physical information perception network framework is proposed, which effectively bridges the simulation-to-reality gap and enables robust extraction of tactile information, including normal/shear force, contact position, 3D Reconstruction and dense force distribution. Experimental results demonstrate the system’s ability to accurately reconstruct multi-physical contact information, while the adaptive grasping strategy ensures stable and reliable object manipulation. The ViTac-Gripper represents a significant advancement in underwater robotics, offering a cost-effective and versatile solution for underwater manipulation tasks. Hongfei Chu, Xuejian Bai, Naijun Liu, Fei Suo, Shuo Wang 0001, Min Tan 0001, Yu Wang 0062 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous ManipulationsabstractLong-horizon dexterous tasks have been a long-standing problem in robotic manipulation. Previous studies have developed task and motion planning, imitation learning, and reinforcement learning methods for long-horizon manipulations. However, these methods are hard to achieve efficient planning for new tasks. Empowered with the Vision Language Model (VLM), recent studies significantly improve the generalization of robot systems. However, these works are only verified in simple pick-and-place tasks due to limited skills. To this end, we propose the Learn-Gen-Plan (LGP), which combines the VLM and learning-based primitives to endow robots with the ability to efficiently plan and complete various long-horizon dexterous tasks. LGP contains two key phases: skill generation and task planning. In skill generation, the Skill Generator is proposed to utilize the learned key primitives and hand-crafted trivial primitives to generate adaptive robot skills. In task planning, the Multimodal Planner generates the robot plan based on image observation, generated skills, and text prompts. We set up a series of dexterous tasks (e.g., cable routing, peg-in-hole assembly) in a real-world lighting circuit wiring scenario to evaluate LGP. The experimental results show that LGP efficiently generates robot plans with learned skills, controlling the robot to complete various multi-step cable wiring tasks. Peng Hao 0003, Shaowei Cui, Junhang Wei, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Vision-Based Autonomous Robotic Arc Welding: State-of-the-Art Review and Perspectives
Yunkai Ma, Junfeng Fan, Yichen Fu, Shuo Wang 0001, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Dexterity-Guided Dimensional Synthesis and Multi-Task Control for Fingertip ManipulationabstractThe geometric parameters severely impact the performance of dexterous manipulation, but manually adjusting them is time-consuming. In this paper, we employ the dexterity-guided dimensional synthesis to explore the geometric design space of a robotic hand, and propose a multi-objective optimization framework to improve its manipulation dexterity. Specifically, we first establish a screw-based mathematical model to describe the hand-object system’s kinetostatic properties. Three objectives are specified to achieve Pareto optimality: maximizing the system’s feasible position space, feasible orientation space, and manipulation stability. Furthermore, to validate the optimal design parameters in practice, we develop a multi-task object motion controller and apply it to a letter handwriting task. Finally, the optimized results are quantitatively analyzed by simulation, yielding a Pareto front and 112 optimal solutions. The applicability of the optimal parameters is assessed through a letter “O” handwriting experiment. Results show that the optimized hand can write the “O” with a maximum radius of 30.0 mm, which is 27.12% larger than that of the non-optimized hand. This controller is also used to schedule subtasks with varying priorities to avoid constrained regions. In contrast to the neural network-based PID controller, the designed controller prioritizes the task’s critical parts, ensuring the legibility of the letters. Congjia Su, Rui Wang 0031, Shaowei Cui, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | GelStereo Tip: A Spherical Fingertip Visuotactile Sensor for Multi-Finger Screwing ManipulationabstractDexterous hands are the key element for robots to achieve human-like manipulation capabilities. An outstanding challenge is to provide fingertips of dexterous hands with precise tactile deformation sensing capabilities. In this paper, we present the GelStereo Tip, a spherical and easy-to-integrate GelStereo-type visuotactile sensor capable of sensing high-resolution 3D elastomer deformation. Previous calibration method does not take into account the impact of imaging errors caused by the sensor’s compact and high-curvature structural characteristics on the accuracy of tactile sensing. Therefore, we propose a novel self-calibration method based on the Refractive Stereo Ray Tracing model, named GTSC, and demonstrate the accuracy of less than 0.3 mm for deformation sensing. Furthermore, we also propose a Contact Retention Tactile Controller to address the issue of fingertips being unable to overcome obstructive torque during the multi-finger bottle cap screwing. After integrating GelStereo Tip into fingertips of Allegro Hand, the controller adjusts the joint positions of the given trajectory using proportional control based on the difference between the sensor’s actual deformation and the reference state for contact retention. We believe that the GelStereo Tip sensor combined with robotic dexterous hands has great application potential in the field of multi-finger fingertip manipulation. Note to Practitioners—The motivation of this paper is to design a fingertip visuotactile sensor with high-precision 3D tactile deformation sensing capabilities for multi-finger robotic hands and to validate its sensing performance. Additionally, it aims to address the issue of overcoming resistance in multi-finger screwing manipulations. Currently, most sensors do not consider the refraction effect or ignore the impact of planar imaging errors in refractive calibration. This paper proposes a visuotactile sensor along with a corresponding self-calibration method to ensure its sensing accuracy. Experiments show that our sensor possesses high-precision and robust 3D deformation sensing capabilities. On the other hand, multi-finger hands often struggle to complete screwing tasks along the given trajectory due to disturbances from torque resistance. This paper proposes a tactile controller that evaluates the contact state through aforementioned tactile sensing to improve subsequent trajectory and achieve continuous screwing. Comparative experiments highlight the necessity of this controller and the reliability of tactile sensing. We hope that the design of our sensor, the self-calibration method, and the tactile controller applied to multi-finger screwing can provide new insights for other practitioners. Boyue Zhang 0002, Shaowei Cui, Chaofan Zhang, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Hardness-Aware Metric Learning With Cluster-Guided Attention for Visual Place RecognitionabstractVisual place recognition is crucial to accurate localization in large-scale environments. Existing methods combine convolutional neural network and deep metric learning to improve performance, however, it is still challenging to promote the adaptability of features under various environmental conditions. To address the problem, this paper proposes a hardness-aware metric learning method with cluster-guided attention. By leveraging the affinity between each local feature and the corresponding scene cluster center, the model is attracted to focus on the local features proximate to their clusters, while suppressing outlier local features that deviate from their clusters. In this way, the scene-related reliable local features are concentrated to construct the global feature of the whole image with adaptability to environmental conditions. Meanwhile, a hardness-aware metric loss is designed to train the proposed network, which determines the hardness of negative samples based on their similarity to the query images and the training iterations. Subsequently, the hardness is used to reweigh each term of the loss function to promote network optimization. In addition, a condition normalization layer is also introduced to regularize the feature distributions under different environmental conditions to a canonical space, improving the feature robustness to condition variations. Our method achieves the top-10 recalls of 97.2%, 99.0%, and 94.9% on Pitts250k-test, TokyoTM-val, and Tokyo 24/7 datasets, respectively. Extensive experiments demonstrate that the proposed method learns robust global features with the adaptability to various environmental conditions. Peiyu Guan, Zhiqiang Cao 0002, Shengxuan Fan, Yuequan Yang, Junzhi Yu 0001, Shuo Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Focus-TransUnet3D: High-Precision Model for 3D Segmentation of Medical Point TargetsabstractDeep learning has been extensively applied in medical image segmentation, providing significant support for disease diagnosis. However, traditional encoder-decoder networks struggle with segmenting scale-sensitive point target lesions. To address this challenge, this paper proposes an innovative incremental fusion architecture that can integrate different models and achieve significant performance improvements through complementary fusion. Based on this architecture, we developed Focus-TransUnet3D by combining the Trans-FusionNet3D model and the 3D Unet model. This model adopts a global-to-local segmentation strategy, effectively addressing the challenges of medical point target segmentation, thereby expanding the application of deep learning in the field of medical image processing. Furthermore, we design a deep fusion strategy suitable for the transformer model to adapt to multi-scale feature learning. The integration of the transformer model with convolutional neural networks brings improvements in local and global feature extraction capabilities, enhancing the applicability of our model. We evaluate our model on three clinical datasets with different target scales: the Intracranial Artery dataset, the Intracranial Aneurysm dataset, and the LiTS17 dataset. The results indicate that in the external test for intracranial aneurysm auxiliary diagnosis, the model trained with only 47 annotated samples achieved the state-of-the-art performance, attaining a Dice coefficient of 84.14% and a sensitivity of 100%. This effectively addresses the challenges of annotation scarcity and tiny targets. Our code will be released athttps://github.com/caijilia/FTUnet3D. Dihua Zhai, Hao Li 0075, Ke Tian, Yi Yang 0009, Zhenyao Chang, Shuo Wang 0001, Yuanqing Xia |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Design and Pipeline Tracking Control of an Underwater Biomimetic Vehicle-Manipulator System With Hybrid PropulsionabstractUnderwater vehicle-manipulator systems (UVMSs) play crucial roles in the fields of underwater target monitoring and pipeline maintenance. However, achieving accurate tracking for underwater pipelines is challenging due to the complexity of UVMSs in terms of nonlinearity, strong coupling and underactuation. To solve the aforementioned problems, an underwater biomimetic vehicle-manipulator system (UBVMS) and an underwater pipeline tracking control method based on the robot vision are proposed. The UBVMS is equipped with the biomimetic undulatory fin propulsors and the biomimetic flipper propulsors, which are inspired by the median and/or paired fin propulsion mode and the body and/or caudal fin propulsion mode of fishes, respectively. The biomimetic undulatory fin propulsors provide the UBVMS with advantages of maneuverability and stability, while the biomimetic flipper propulsors enable the UBVMS to have improved acceleration ability. A tracking control algorithm with adaptive weight coefficients is designed to improve the pose stability of the UBVMS. A fuzzy rule mapping model is constructed to describe the nonlinear relationship between the biomimetic propulsors' control parameters and the propulsive force/torque. Finally, four types of pipeline tracking experiments are conducted to verify the effectiveness and feasibility of the proposed UBVMS and control algorithm. Xuejian Bai, Yu Wang 0062, Xiaolong Hui, Shuo Wang 0001, Min Tan 0001 |
IEEE Trans. Cybern. | 6 |
| 2025 | TacFlex: Multimode Tactile Imprints Simulation for Visuotactile Sensors With Coating PatternsabstractVisuotactile sensors have been shown to provide rich contact information for robots. However, how to build a high-fidelity visuotactile simulator that supports multi-mode tactile imprints and various sensor configurations (such as coating patterns) remains a challenging problem. In this paper, we present TacFlex, an efficient and flexible simulator for visuotactile sensors, which physically simulates the elastomer deformation using Finite Element Methods (FEM), and focuses on linking the deformed elastomer mesh to diverse tactile imprints, including tactile images with arbitrary coating patterns and tactile 3D point clouds. We further propose a ray tracing-based rectification method to deal with multi-medium refraction effects to make the simulated tactile images more realistic. Extensive qualitative and quantitative experiments are conducted to demonstrate the effectiveness of TacFlex on several visuotactile sensors. Furthermore, we explore the Sim2Real performance of different tactile imprints provided by TacFlex in tactile perception and manipulation tasks, such as cylindrical object pose estimation and peg-in-hole. The perception/policy models trained in simulation are successfully deployed in the real world. Finally, we present the outlook on the potential of TacFlex in visuotactile manipulation learning. The TacFlex simulator is open-sourced to the community. See supplementary video, code, and results athttps://sites.google.com/view/tacflex/. Chaofan Zhang, Shaowei Cui, Jingyi Hu, Tiandong Zhang, Rui Wang 0031, Shuo Wang 0001 |
IEEE Trans. Robotics | 7 |
| 2025 | FlowSight: Vision-Based Artificial Lateral Line Sensor for Water Flow PerceptionabstractThis paper presents a novel vision-based artificial lateral line (ALL) sensor, FlowSight, enhancing the perception capabilities of underwater robots. Through an autonomous vision system, FlowSight allows for simultaneous sensing the speed and direction of local water flow without relying on external auxiliary equipment. Inspired by the lateral line neuromast of fish, a flexible bionic tentacle is designed to sense water flow. Deformation and motion characteristics of the tentacle are modeled and analyzed using bidirectional fluid-structure interaction (FSI) simulation. Upon contact with water flow, the tentacle converts water flow information into elastic deformation information, which is captured and processed into an image sequence by the autonomous vision system. Subsequently, a water flow perception method based on deep neural networks is proposed to estimate the flow speed and direction from the captured image sequence. The perception network is trained and tested using data collected from practical experiments conducted in a controllable swim tunnel. Finally, the FlowSight sensor is integrated into the bionic underwater robot RoboDact, and a closed-loop motion control experiment based on water flow perception is conducted. Experiments conducted in the swim tunnel and water pool demonstrate the feasibility and effectiveness of FlowSight sensor and the water flow perception method. Tiandong Zhang, Rui Wang 0031, Qiyuan Cao, Shaowei Cui, Gang Zheng 0002, Shuo Wang 0001 |
IEEE Trans. Robotics | 6 |
| 2024 | Autonomous Manipulation of an Underwater Vehicle-Manipulator System by a Composite Control Scheme With Disturbance EstimationabstractThis article addresses an autonomous manipulation problem for an underwater vehicle-manipulator system (UVMS) operating in a free-floating way while subjecting to unknown continuous disturbance. More specifically, a composite control scheme composed of disturbance observer (DOB), predictor model network (PM-Net), and nonlinear model predictive control (NMPC), is devised to improve the control performance of UVMS (i.e., unicycle-like UVMS actuated only in the surge, heave, and yaw for vehicle body) in the case of disturbance, model mismatch, and input saturation. A RBF-DOB is formulated by combining a DOB and a Radial Basis Function (RBF) neural network to estimate disturbance at the current step. Then, the PM-Network, composed of a disturbance predictor network and state predictor network, is developed based on long short-term memory (LSTM) network that predicts UVMS state sequences considering model mismatch and disturbance. The NMPC is deployed as a feedback control law to endow the input saturation of the UVMS and produce optimal control action. Compared with conventional DOB control methods using feed-forward compensation of disturbance, the primary merit of the proposed approach is that the disturbance estimated by RBF-DOB is utilized in the PM-Net to predict future UVMS state sequences, which are exploited on the NMPC’s receding optimization. Finally, realistic simulation and relevant experiment are conducted to demonstrate the effectiveness of the proposed method. Note to Practitioners—The motivation behind this article is the autonomous manipulation of an underwater vehicle-manipulator system subjected to unknown disturbance. However, it is not always feasible or straightforward to obtain the external disturbance and unmodeled dynamics for designing robust controllers. On the one hand, how to manipulate the disturbance into the designed controller to generate optimal control action rather than by using feed-forward compensation. On the other hand, the control input saturation often occurs in the UVMS control, especially under the disturbance rejection conditions, where it should be considered in the controller design. Currently, the predominant methods for UVMS control lack a control scheme that provides a complete and credible control strategy that takes the aforementioned issues into consideration. Motivated by the above analysis, this study provides a composite control scheme to deal with the dynamic uncertainties, unknown disturbance, and input saturation. The results of realistic simulation and relevant experiments demonstrate the effectiveness of the proposed method. Hopefully, our control method can provide valuable theoretical and technical guidance to practicing marine engineers for controller design. Mingxue Cai, Yu Wang 0062, Shuo Wang 0001, Rui Wang 0031, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Sample-Observed Soft Actor-Critic Learning for Path Following of a Biomimetic Underwater VehicleabstractThis paper addresses a learning-based path following control scheme for a biomimetic underwater vehicle (BUV) driven by undulatory fins. A dynamic line-of-sight (DLOS) guidance system is designed, which uses a virtual ball with a dynamic radius to detect the reference path. This DLOS system guides our BUV in the path following control and extracts essential information for the Markov decision process (MDP) of the control task. A deep reinforcement learning (DRL) algorithm, sample-observed soft actor-critic (SOSAC) is proposed. The can train out control policy with greater cumulative reward and higher success rate by using two tricks: sample observation and sample diversification. Based on the DLOS system, the MDP of the control task, and a multilayer perceptron (MLP) trained by the SOSAC, our control scheme is established. Experiments show that our BUV can successfully achieve path following control in an indoor pool environment by using this control scheme.Note to Practitioners—The motivation of this paper is to design a practical end-to-end path following control scheme for the BUV driven by undulatory fins, and verify this scheme in a real-world environment. Unlike common autonomous underwater vehicles (AUVs) using axial propellers, the BUVs apply biomimetic propellers such as the undulatory fin. Multimodel wave patterns can be implemented by the undulatory fin, which generates nonlinear thrust and lateral force simultaneously. This propulsive feature makes the driving force on different directions of the BUV to be strong coupled, and it is complicated to convert the outputs of a common controller into waveform parameters of the undulatory fins to control the BUV. Therefore, in this paper, we proposed an end-to-end learning-based path following controller, which observes environmental information and directly generates waveform parameters to control our BUV. Experiments suggest that our control scheme is practical and valid. Yu Wang 0062, Shuo Wang 0001, Long Cheng 0001, Rui Wang 0031, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | ToolBot: Learning Oriented Keypoints for Tool Usage From Self-SupervisionabstractTool usage is critical for enabling robots to complete challenging tasks that exceed their innate capabilities. Task-oriented grasp and manipulation are two primitive actions in tool usage tasks. In this article, we present an end-to-end framework for jointly inferring two primitive actions through self-supervision, which can guide robots to complete tool usage tasks. We formulate primitive actions as oriented keypoint representations so that the existing pose-based policies can be easily used to achieve tool usage tasks. To address the low task completion rates in self-supervision, we propose a novel technique based on self-supervision properties and forward kinematic models to generate additional effective training samples. The resulting system,ToolBot, is evaluated with the following four different kinds of tools: hammer, knife, screwdriver, and wrench, and it achieves an average task success rate of 82.88% in simulation for four tools and 77.72% in real-world experiments. Junhang Wei, Peng Hao 0003, Shuo Wang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Learning-Based Slip Detection for Dexterous Manipulation Using GelStereo SensingabstractEndowing the robot with tactile perception can effectively improve manipulation dexterity, along with various benefits of human-like touch. Using GelStereo (GS) tactile sensing, which gives high-resolution contact geometry information, including 2-D displacement field, and 3-D point cloud of the contact surface, we present a learning-based slip detection system in this study. The results reveal that the well-trained network achieves 95.79% accuracy on the never-seen testing dataset, which surpasses the current model-based and learning-based methods using visuotactile sensing. We also propose a general framework for slip feedback adaptive control for dexterous robot manipulation tasks. The experimental results show the effectiveness and efficiency of the proposed control framework using GS tactile feedback when deployed on real-world grasping and screwing manipulation tasks on various robot setups. Shaowei Cui, Shuo Wang 0001, Rui Wang 0031, Chaofan Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | GelStereo Palm: A Novel Curved Visuotactile Sensor for 3-D Geometry SensingabstractRecently, visuotactile sensors have shown promising potential in robotics due to their high-resolution sensing ability. Unfortunately, the majority of available visuotactile sensors are limited to flat shapes, which severely limits their application possibilities. In this article, we propose a novel curved visuotactile sensor, the GelStereo Palm, which senses the 3-D contact geometry on a curved surface using a binocular vision system. Meanwhile, to solve the light refraction problem in the binocular stereo vision system under a curved medium, a refractive stereo ray tracing model for GelStereo Palm is presented. Moreover, a 3-D tactile point cloud sensing pipeline is introduced to reconstruct the 3-D contact geometry in real-time. Finally, extensive experiments are conducted to verify the accuracy and robustness of the 3-D contact geometry sensing of our GelStereo Palm sensor. Jingyi Hu, Shaowei Cui, Shuo Wang 0001, Chaofan Zhang, Rui Wang 0031, Lipeng Chen |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | An Efficient Training Approach for Very Large Scale Face RecognitionabstractFace recognition has achieved significant progress in deep learning era due to the ultra-large-scale and well- labeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is in- dispensable. The heavy computational and memory costs mainly result from the million-level dimensionality of the fully connected (FC) layer. To this end, we propose a novel training approach, termed Faster Face Classification (F2C), to alleviate time and cost without sacrificing the performance. This method adopts Dynamic Class Pool (DCP) for storing and updating the identities' features dy-namically, which could be regarded as a substitute for the FC layer. DCP is efficiently time-saving and cost-saving, as its smaller size with the independence from the whole face identities together. We further validate the proposed F2C method across several face benchmarks and private datasets, and display comparable results, meanwhile the speed is faster than state-of-the-art FC-based methods in terms of recognition accuracy and hardware costs. More-over, our method is further improved by a well-designed dual data loader including indentity-based and instance- based loaders, which makes it more efficient for updating DCP parameters. Kai Wang 0036, Shuo Wang 0001, Xiaojiang Peng, Baigui Sun, Hao Li 0030, Yang You 0001 |
CVPR | 2 |
| 2022 | CAFE: Learning to Condense Dataset by Aligning FeaturesabstractDataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intuitive motivation and promising results, such gradient-based methods, by nature, easily overfit to a biased set of samples that produce dominant gradients, and thus lack a global supervision of data distribution. In this paper, we propose a novel scheme to Condense dataset by Aligning FEatures (CAFE), which explicitly attempts to preserve the real-feature distribution as well as the discriminant power of the resulting synthetic set, lending itself to strong generalization capability to various architectures. At the heart of our approach is an effective strategy to align features from the real and synthetic data across various scales, while accounting for the classification of real samples. Our scheme is further backed up by a novel dynamic bi-level optimization, which adaptively adjusts parameter updates to prevent over-/under-fitting. We validate the proposed CAFE across various datasets, and demonstrate that it generally outperforms the state of the art: on the SVHN dataset, for example, the performance gain is up to 11%. Extensive experiments and analysis verify the effectiveness and necessity of proposed designs. Kai Wang 0036, Bo Zhao 0038, Shuo Yang 0006, Shuo Wang 0001, Guan Huang 0003, Hakan Bilen, Xinchao Wang, Yang You 0001 |
CVPR | 6 |
| 2022 | Meta-Imitation Learning by Watching Video Demonstrations
Tao Lu 0006, Xiaoge Cao, Yinghao Cai, Shuo Wang 0001 |
ICLR | 5 |
| 2022 | Learning-based Six-axis Force/Torque Estimation Using GelStereo Fingertip Visuotactile SensingabstractVisuotactile sensors have recently attracted much attention in robot communities due to the benefit of high spatial resolution sensing. However, force/torque estimation by visuotactile sensors remains a challenging problem. In this paper, we propose a learning-based six-axis force/torque estimation network using GelStereo visuotactile sensor, which can provide two-dimensional (2D) and three-dimensional (3D) displacements of markers embedded in the sensor surface. The convolutional neural networks are employed to extract multi-modal tactile deformation features; and a novel contact positional encoding method is proposed to eliminate the influence of translation invariance in convolutional operators. The well-trained model achieves the best RMSE of 0.290 N in force and 0.0084 Nm in torque. Furthermore, the proposed force/torque estimation network is integrated with a force-feedback policy for adaptive grasping tasks. The experimental results demonstrate the effectiveness of the proposed method and its potential application in robotic grasping and manipulation tasks. Chaofan Zhang, Shaowei Cui, Yinghao Cai, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001 |
IROS | 6 |
| 2022 | VGPN: 6-DoF Grasp Pose Detection Network Based on Hough VotingabstractIn this paper, we propose a novel Voting based Grasp Pose Network (VGPN) to detect 6-DoF grasps in cluttered scenes. The motivation of this paper is that local object geometry can provide useful clues about where the object can be grasped. Generated by the sampled seed points from raw point cloud, the votes allow seed points in different object regions to contribute to locations where the object can be grasped. Geometric features from various local regions are aggregated to generate grasps in a more confident and dense space, which enables grasp prediction utilizing more global context features. The search space of grasp pose detection is also greatly reduced. Experimental results on both simulation and real-world environments show that our proposed method outperforms state-of-the-art approaches in terms of both success rate and coverage of the ground truth grasps. The objects can be grasped with fewer attempts which is critical in real-world applications. Yinghao Cai, Tao Lu 0006, Shuo Wang 0001 |
IROS | 4 |
| 2022 | Modeling and analysis of an underwater biomimetic vehicle-manipulator system
Xuejian Bai, Yu Wang 0062, Shuo Wang 0001, Rui Wang 0031, Min Tan 0001, Wei Wang 0292 |
Sci. China Inf. Sci. | 3 |
| 2022 | Manipulation skill learning on multi-step complex task based on explicit and implicit curriculum learning
Naijun Liu, Tao Lu 0006, Yinghao Cai, Rui Wang 0031, Shuo Wang 0001 |
Sci. China Inf. Sci. | 5 |
| 2022 | Design and Locomotion Control of a Dactylopteridae-Inspired Biomimetic Underwater Vehicle With Hybrid PropulsionabstractThis article presents the design and implementation of an innovative biomimetic underwater vehicle (BUV) and its locomotion controller. Through mimicking a dactylopteridae, the hybrid propulsion BUV is designed with two symmetrical bio-inspired long-fins and a double-joint fishtail. The mechatronic design of the dactylopteridae-inspired BUV with the pectoral long-fins and a double-joint fishtail is first provided. The two flexible long-fins compose the median and/or paired fin (MPF) propulsion, while the fishtail acts as the body and/or caudal fin (BCF) propulsion. Through the coordination of BCF and MPF propulsion modes, the BUV obtains excellent low-speed locomotion stability and also keeps high maneuverability. Moreover, the locomotion control methods based on central pattern generators (CPGs) model and fuzzy adaptive proportion integral differential (PID) are proposed for this BUV. In the end, the experimental results of the multimode motion and closed-loop motion control demonstrate the feasibility and effectiveness of the mechanism and the locomotion control system.Note to Practitioners—The motivation behind this article is the design of a novel biomimetic underwater vehicle (BUV) that possesses low-speed locomotion stability and fast swimming ability, which is suitable for carrying relevant sensors to complete water quality monitoring, biological observation, underwater equipment inspection, underwater structure detection, and other marine tasks. Currently, BUVs are usually designed as only one propulsion mode by caudal fin or paired fins, which makes it difficult to have the advantages of both modes. In order to further study the problem, we designed a dactylopteridae-inspired BUV with the bilateral pectoral long-fins (providing low-speed locomotion stability) and a double-joint fishtail (providing fast swimming ability). A hybrid-driven motion control framework is presented for the BUV based on a central pattern generators (CPGs) model and fuzzy adaptive proportion integral differential (PID). A series of experiments suggests that the mechanism and the locomotion control system are practical and valid. Hopefully, our mechanism and control framework can provide valuable theoretical and technical support guidance to the practicing marine engineer for the codesign of propulsion mode and control. Tiandong Zhang, Rui Wang 0031, Yu Wang 0062, Long Cheng 0001, Shuo Wang 0001, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2022 | Multimodal Unknown Surface Material Classification and Its Application to Physical ReasoningabstractUnknown surface material classification (SMC) can inform a robot about material properties, enabling it to interact with environments appropriately. Recent research has leveraged multimodal data using deep learning to improve the performance of SMC. In this article, we present a deep learning model, multimodal temporal convolutional neural network (MTCNN), which integrates energy spectrum, dilated convolutions, and sequence poolings into a unified network architecture. The proposed model can learn material representations from auditory and multitactile (i.e., acceleration, normal force, and friction force) data generated by dragging a tool along surfaces, and distinguish unknown object surface materials into categories. For surface material data collection, a tool is also designed to detect different object surfaces. The performance of MTCNN is evaluated on a public dataset and the highest classification accuracy is 87.55%. A robotic curling example is provided to illustrate how the presented model helps the robot in manipulation. Junhang Wei, Shaowei Cui, Jingyi Hu, Peng Hao 0003, Shuo Wang 0001, Zheng Lou |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Target Tracking Control of a Biomimetic Underwater Vehicle Through Deep Reinforcement LearningabstractIn this article, the underwater target tracking control problem of a biomimetic underwater vehicle (BUV) is addressed. Since it is difficult to build an effective mathematic model of a BUV due to the uncertainty of hydrodynamics, target tracking control is converted into the Markov decision process and is further achieved via deep reinforcement learning. The system state and reward function of underwater target tracking control are described. Based on the actor-critic reinforcement learning framework, the deep deterministic policy gradient actor-critic algorithm with supervision controller is proposed. The training tricks, including prioritized experience replay, actor network indirect supervision training, target network updating with different periods, and expansion of exploration space by applying random noise, are presented. Indirect supervision training is designed to address the issues of low stability and slow convergence of reinforcement learning in the continuous state and action space. Comparative simulations are performed to show the effectiveness of the training tricks. Finally, the proposed actor-critic reinforcement learning algorithm with supervision controller is applied to the physical BUV. Swimming pool experiments of underwater object tracking of the BUV are conducted in multiple scenarios to verify the effectiveness and robustness of the proposed method. Yu Wang 0062, Chong Tang 0004, Shuo Wang 0001, Long Cheng 0001, Rui Wang 0031, Min Tan 0001, Zeng-Guang Hou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Development and Motion Control of Biomimetic Underwater Robots: A SurveyabstractBiomimetic underwater robots have attracted considerable research attention globally, owing to their quieter actuations, higher propulsion efficiency, and stronger maneuverability when compared with conventional underwater vehicles equipped with axial propellers. This article provides a comprehensive survey of current research in this field. First, we review the development status of biomimetic underwater robots in both body/caudal fin (BCF), median/paired fin (MPF), and their hybrid propulsion modes. Then, we outline the motion control methods employed in biomimetic underwater robots, including open-loop swimming control and typical closed-loop control strategies. In particular, we detail our latest studies on the RobCutt series underwater robots. On this basis, some critical issues and future directions are summarized. We predict that biomimetic underwater robots will have excellent prospects in underwater environment exploration and resource utilization. Rui Wang 0031, Shuo Wang 0001, Yu Wang 0062, Long Cheng 0001, Min Tan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Hierarchical Learning from Demonstrations for Long-Horizon TasksabstractAlthough reinforcement learning (RL) has achieved great success in robotic manipulation skills learning, it is still challenging for long-horizon tasks. Combining RL with demonstrations is an effective solution. In this paper, we propose a novel hierarchical learning from demonstrations method for long-horizon tasks, which leverages (i) object-centered segmentation of demonstrations to automatically segment the teaching trajectories into episodes. (ii) a bi-level hierarchical imitation learning method with a parallel training mechanism to train the two-level policies simultaneously. Experimental results on three challenging long-horizon tasks with sparse rewards show that our proposed method significantly outperforms state-of-art approaches in terms of both sample-efficiency and success rate. Moreover, our method is the only one which achieves satisfactory performance in tasks of multi-object stack and multi-object push&stack. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 5 |
| 2021 | DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action NoiseabstractExploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-inefficient. In this paper, we propose a novel fast exploration method for off-policy reinforcement learning, called Density-based Intrinsic Motivation and Self-adaptive Action Noise (DIMSAN). Our main contribution is twofold: (1) We propose a Density-based Intrinsic Motivation (DIM) method. It introduces a new intrinsic-reward generation mechanism based on samples’ density estimation during experience replay and encourages the agent to seek novel and unfamiliar states. (2) We propose a Self-adaptive Action Noise (SAN) to deal with the exploration-exploitation tradeoffs, which could automatically change the exploration step through adding adaptive action space noise. The synergy between DIM and SAN could guide the agent to search the state and action space with high efficiency. We evaluate our method on the benchmark manipulation tasks and the designed challenging ones. Empirical results show that our method outperforms the existing methods in terms of convergence speed and sample efficiency, especially in challenging tasks. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 6 |
| 2021 | Prediction-Based Seabed Terrain Following Control for an Underwater Vehicle-Manipulator SystemabstractThis article addresses a problem of seabed terrain following control (STFC) for an underwater vehicle-manipulator system (UVMS). The motivation is to perform a visual search of marine products closely to seabed in unknown environment. In terms of this issue, we propose a novel and robust STFC framework for our UVMS to maintain an appropriate height to seabed. A nonlinear model predictive control (NMPC) method is formulated to solve the STFC problem. To relieve online computational burden and system noisy influence, Ohtsuka's continuation/generalized minimal residual (C/GMRES) algorithm incorporated with a tracking differentiator (TD) is investigated. In order to improve the following accuracy, the system state prediction part of the NMPC and a long short-term memory (LSTM) network are elaborated to predict future seabed terrain using a depth gauge and an altimeter, respectively. Finally, the three different physical scenarios for STFC problem are established using ROS to demonstrate the robustness and efficiency of the proposed algorithm. Mingxue Cai, Yu Wang 0062, Shuo Wang 0001, Rui Wang 0031, Long Cheng 0001, Min Tan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Coordinated Control of Underwater Biomimetic Vehicle-Manipulator System for Free Floating Autonomous ManipulationabstractThis article presents a coordinated vehicle-manipulator control method for an underwater biomimetic vehicle-manipulator system (UBVMS) to implement floating autonomous manipulation in practice. An algorithm framework composed of adaptive tracking differentiator (ATD), extended state observer (ESO), improved nonsingular terminal sliding-mode control (I-NTSMC), fuzzy-logic controller (FLC), and estimator of manipulator disturbances, is proposed. The ATD is designed to generate desired motion state and alleviate noise. The ESO is developed to estimate the motion state, systematic uncertainties, and external disturbances. The proposed I-NTSMC method assures the finite-time convergence of the system states and alleviate chattering. The estimation of the manipulator disturbances is incorporated into the control strategy to enhance the station keeping of the vehicle. Finally, underwater autonomous free floating manipulation experiments about opening a door and grasping objects are conducted to validate the theoretical results and confirm the feasibility of the proposed control strategy. Mingxue Cai, Shuo Wang 0001, Yu Wang 0062, Rui Wang 0031, Min Tan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | PSO-based Optimal Formation of Multiple Biomimetic Underwater VehiclesabstractThis paper aims to investigate optimal formation solutions of multiple biomimetic underwater vehicles (BUVs). The BUV is propelled by undulatory fins on both sides, and can perform various locomotion patterns, especially turning in situ and diving vertically. Firstly, the optimal formation problem is formulated, followed by theoretical analysis of a special case of optimal line formation. Then, a solution is proposed from the perspective of evolutionary computation. In particularly, the coordinates and the slope of the desired line formation, together with the pairings between initial positions and target positions, are obtained based on particle swarm optimization. Furthermore, we demonstrate the validity of this method by comparing the simulation results with the results of theoretical analysis. Finally, simulations results of multiple BUVs verify the feasibility of the proposed optimal formation methods. Rui Wang 0031, Ge Bai, Shuo Wang 0001, Yu Wang 0062, Min Tan 0001 |
CEC | 3 |
| 2020 | Grasp State Assessment of Deformable Objects Using Visual-Tactile Fusion PerceptionabstractHumans can quickly determine the force required to grasp a deformable object to prevent its sliding or excessive deformation through vision and touch, which is still a challenging task for robots. To address this issue, we propose a novel 3D convolution-based visual-tactile fusion deep neural network (C3D-VTFN) to evaluate the grasp state of various deformable objects in this paper. Specifically, we divide the grasp states of deformable objects into three categories of sliding, appropriate and excessive. Also, a dataset for training and testing the proposed network is built by extensive grasping and lifting experiments with different widths and forces on 16 various deformable objects with a robotic arm equipped with a wrist camera and a tactile sensor. As a result, a classification accuracy as high as 99.97% is achieved. Furthermore, some delicate grasp experiments based on the proposed network are implemented in this paper. The experimental results demonstrate that the C3D-VTFN is accurate and efficient enough for grasp state assessment, which can be widely applied to automatic force control, adaptive grasping, and other visual-tactile spatiotemporal sequence learning problems. Shaowei Cui, Rui Wang 0031, Junhang Wei, Fanrong Li, Shuo Wang 0001 |
ICRA | 5 |
| 2020 | ACDER: Augmented Curiosity-Driven Experience ReplayabstractExploration in environments with sparse feed-back remains a challenging research problem in reinforcement learning (RL). When the RL agent explores the environment randomly, it results in low exploration efficiency, especially in robotic manipulation tasks with high dimensional continuous state and action space. In this paper, we propose a novel method, called Augmented Curiosity-Driven Experience Replay (ACDER), which leverages (i) a new goal-oriented curiosity-driven exploration to encourage the agent to pursue novel and task-relevant states more purposefully and (ii) the dynamic initial states selection as an automatic exploratory curriculum to further improve the sample-efficiency. Our approach complements Hindsight Experience Replay (HER) by introducing a new way to pursue valuable states. Experiments conducted on four challenging robotic manipulation tasks with binary rewards, including Reach, Push, Pick&Place and Multi-step Push. The empirical results show that our proposed method significantly outperforms existing methods in the first three basic tasks and also achieves satisfactory performance in multi-step robotic task learning. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 6 |
| 2020 | Grasping Marine Products With Hybrid-Driven Underwater Vehicle-Manipulator SystemabstractThis article presents the comprehensive framework for a hybrid-driven underwater vehicle-manipulator system (HD-UVMS) to grasp marine products on the seabed. The purpose of the proposed hybrid-driven propulsion system is to improve the swimming ability of the HD-UVMS by using thrusters and enhance the stability of its pose adjustment mechanism via two unique long fin propulsors. The control mode for the thrusters and long fin propulsors is based on a fuzzy logic control method. Subsequently, a lightweight manipulator is developed to grasp marine products. The open-closed angle and current controls for the gripper help to avoid damaging marine products. A vision system is installed to enable the HD-UVMS to gradually approach marine products with the aid of monocular vision and grasp them with the aid of binocular vision. A detailed method for monocular passive ranging and stereo matching, in accordance with real-time metrics, is elaborated. Finally, relevant experiments are conducted in an indoor pool and under real sea condition to assess the effectiveness of the proposed framework. Note to Practitioners-The motivation behind this article is the design of an underwater vehicle-manipulator system that can grasp marine products on the real seabed and perform other underwater intervention tasks. Currently, the predominant method of fishing for marine products relies on human divers, which has disadvantages for human divers' health due to the long periods of time spent working underwater. In order to further study the problem, this article develops a hybrid-driven underwater vehicle-manipulator system (HD-UVMS) to work in a real seabed environment. A hybrid-driven motion control framework is presented using the thrusters to achieve effective cruising and searching for marine products and long fin propulsors for the fine pose adjustment required to grasp marine products. The proposed lightweight underwater manipulator can grasp marine products on the seabed with the aid of a vision system. A series of experiments suggests that the HD-UVMS is practical and valid. Mingxue Cai, Yu Wang 0062, Shuo Wang 0001, Rui Wang 0031, Yong Ren 0001, Min Tan 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2020 | Solving Trajectory Optimization Problems in the Presence of Probabilistic ConstraintsabstractThe objective of this paper is to present an approximation-based strategy for solving the problem of nonlinear trajectory optimization with the consideration of probabilistic constraints. The proposed method defines a smooth and differentiable function to replace probabilistic constraints by the deterministic ones, thereby converting the chance-constrained trajectory optimization model into a parametric nonlinear programming model. In addition, it is proved that the approximation function and the corresponding approximation set will converge to that of the original problem. Furthermore, the optimal solution of the approximated model is ensured to converge to the optimal solution of the original problem. Numerical results, obtained from a new chance-constrained space vehicle trajectory optimization model and a 3-D unmanned vehicle trajectory smoothing problem, verify the feasibility and effectiveness of the proposed approach. Comparative studies were also carried out to show the proposed design can yield good performance and outperform other typical chance-constrained optimization techniques investigated in this paper. Runqi Chai, Al Savvaris, Antonios Tsourdos, Senchun Chai, Yuanqing Xia, Shuo Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2019 | Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network*abstractIn this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal sequence and vehicle motion, estimating the motion state of the next moment from historical data, eliminating complex CFFR modeling. The other is the tracking network. Based on the simulation data of the motion model, the relationship between the target trajectory and the current control command is learned, which simplifies the design and cumbersome tuning of the complex controller. The effectiveness of the proposed method is verified in both simulated and real-world environments. Qualitative and quantitative experimental results verify the accuracy of the tracking. Wenkai Chang, Caiyun Yang, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
IROS | 6 |
| 2019 | Parameter estimation survey for multi-joint robot dynamic calibration case study
Shuo Wang 0001, Min Tan 0001 |
Sci. China Inf. Sci. | 2 |
| 2019 | A Sensorless Hand Guiding Scheme Based on Model Identification and Control for Industrial RobotabstractMost industrial robots are not capable of teaching by hand and require path points to be specified by teaching pendants. To enable the teaching of industrial robots by hand without any force sensors, this paper proposes a scheme to minimize the external force estimation error and reduce disturbance in the guiding task by using the virtual mass and virtual friction model. In this case, the maximum velocity and acceleration of the robot end effector shall be limited to ensure safety. Thus, the operator is allowed to guide the robot by hand. The joint torque is obtained from the motor current. The inertial force and friction of the links and driving systems are analyzed. The nonlinear dynamic model of the industrial robot is built and its parameters are calibrated by a nonlinear method. The force estimation is referenced to set the virtual friction and to design the force-following controller. Hence the end effector can follow the direction of external force compliantly and suppress jitters. Finally, several experiments on a six degrees of freedom industrial robot demonstrate the validity of the proposed control scheme. Shuo Wang 0001, Min Tan 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | A Paradigm for Path Following Control of a Ribbon-Fin Propelled Biomimetic Underwater VehicleabstractThis paper addresses the problem of path following for biomimetic underwater vehicles (BUVs) propelled by undulatory ribbon-fins. First, the general kinematics and dynamics models of underwater vehicles are presented, followed by a fuzzy logic model for dealing with a nonlinear relationship between the propulsive force/torque and the control parameters of the undulatory fins of the BUV. Then the path following problem of the BUV is formulated. A path following control paradigm integrating the line-of-sight guidance system with backstepping (BP) technique is proposed to maneuver the BUV to follow a predefined parameterized curve without time constraints. The stability of the BP controller is analyzed and guaranteed by Lyapunov stability theory. Finally, simulations and experimental results illustrate the performance of the proposed path following control paradigm. Rui Wang 0031, Shuo Wang 0001, Yu Wang 0062, Min Tan 0001, Junzhi Yu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | A Navigation Framework for Mobile Robots with 3D LiDAR and Monocular CameraabstractThis paper presents a navigation framework for mobile robots. In order to traverse in different environments, mobile robots have to avoid obstacles it encountered and make a suitable decision with the consideration of global destination. On the basis of grid map provided by the 3D LiDAR, this framework is related to the SLAMs with motion planning methods, where different SLAMs and different planning solutions can be combined to achieve the navigation in dynamic environments. Concretely, the ORB SLAM2 with a slightly improved VFH+ provides a solution, and a navigation method based on Bezier curve with Cartographer SLAM is presented, specifically, we consider the influence area caused by the dynamic obstacle with probabilistic descriptions. The effectiveness is verified by the experiments. Yelan Wu, Shuang Liang 0004, Zhiqiang Cao 0002, Shuo Wang 0001 |
IECON | 6 |
| 2017 | Path following for a biomimetic underwater vehicle based on ADRCabstractThis paper addresses the problem of path following for a biomimetic underwater vehicle (BUV) propelled by undulatory fins with uncertain model and unknown disturbance. The mechanical structure of the BUV is briefly described. Moreover, the general kinematics and dynamics models of the vehicle are presented and the path following problem is formulated. The controller combining line-of-sight (LOS) guidance system with active disturbance rejection control (ADRC) technique is designed to maneuver the BUV to follow a predefined parameterized curve. Specifically, a guidance system based on LOS principle is implemented to decouple the multi-variable system to steer the surge speed and the course respectively. Furthermore, in order to deal with model uncertainty, ADRC is used in development of the surge speed controller and the course controller. Finally, simulations and experimental results validated the performance of the proposed path following control scheme. Rui Wang 0031, Shuo Wang 0001, Yu Wang 0062, Chong Tang 0004 |
ICRA | 2 |
| 2017 | Generation of temporal-spatial Bezier curve for simultaneous arrival of multiple unmanned vehicles
Shuo Wang 0001, Rui Wang 0031, Min Tan 0001 |
Inf. Sci. | 2 |
| 2013 | Motion modeling and neural networks based yaw control of a biomimetic robotic fish
Chao Zhou 0002, Zeng-Guang Hou, Zhiqiang Cao 0002, Shuo Wang 0001, Min Tan 0001 |
Inf. Sci. | 4 |
| 2013 | Backward swimming gaits for a carangiform robotic fish
Chao Zhou 0002, Zhiqiang Cao 0002, Zeng-Guang Hou, Shuo Wang 0001, Min Tan 0001 |
Neural Comput. Appl. | 4 |
| 2010 | Fuzzy logic PID based control design for a biomimetic underwater vehicle with two undulating long-finsabstractThis paper proposes a fuzzy logic PID based control method for a biomimetic underwater vehicle on swimming speed and yaw angle. The vehicle has a pair of long-fins installed symmetrically on its body. A set of inertial sensors are applied for collecting its velocity and pose information. And an embedded control system and a driving system based on FPGA are designed for generating and switching motion modes. Based on its motion discipline and system architecture, black-box identification is employed for system modeling. Therefore, according to the control system and driving system, a fuzzy logic PID control scheme is proposed for the underwater vehicle. A fuzzy logic controller is applied for the vehicle before the error is reduced to a given range. Then PID controller is applied when the error is within the given range. Here, velocity control is considered, which involves swimming speed and yaw angle. The fuzzy logic PID control scheme is used for the goals respectively. And Yaw angle control is prior to the swimming speed control. Finally, simulation results show the proposed control scheme is valid. Liuji Shang, Shuo Wang 0001, Min Tan 0001 |
IROS | 2 |
| 2008 | Kinematic modeling of a bio-inspired robotic fishabstractThis paper proposes a kinematic modeling method for a bio-inspired robotic fish based on single joint. Lagrangian function of freely swimming robotic fish is built based on a simplified geometric model. In order to build the kinematic model, the fluid force acting on the robotic fish is divided into three parts: the pressure on links, the approach stream pressure and the frictional force. By solving Lagrange's equation of the second kind and the fluid force, the movement of robotic fish is obtained. The robotic fish's motion, such as propelling and turning are simulated, and experiments are taken to verify the model. Chao Zhou 0002, Min Tan 0001, Zhiqiang Cao 0002, Shuo Wang 0001, Douglas C. Creighton, Nong Gu, Saeid Nahavandi |
ICRA | 4 |
| 2008 | The dynamic analysis of the backward swimming mode for biomimetic carangiform robotic fishabstractThe swimming backward method for biomimetic carangiform robotic fish is analyzed in this paper based on the dynamic/kinematic model. The equation of Lagrange of multi-link carangiform robotic fish and simplified fluid force are inducted to calculate the dynamic and kinematic characteristics of the motions. A specific gait is calculated to make the profile of the carangiform robotic fishpsilas undulation fit the characteristics of European eelpsilas swimming backward, which is summarized from the motion sequence of European eel. The simulated and experimental data is given to verify the method. Chao Zhou 0002, Zhiqiang Cao 0002, Shuo Wang 0001, Min Tan 0001 |
IROS | 3 |
| 2006 | The Posture Control and 3-D Locomotion Implementation of Biomimetic Robot FishabstractIn this paper, a method for the posture control of a biomimetic robot fish TPF-I is proposed. In this method, the position of the robot fish's gravity centre can be changed by a barycenter-adjustor, which leads to the pitching angle changing. Propelled by coordinating a multi-link body and a tail, the robot fish can complete the posture control and 3-D Locomotion. The 3-D locomotion and posture control are implemented by synthesizing three basic control methods speed control, orientation control and pitching control, which are described in detail respectively. Finally, the experimental results of the robot fish's motion control are given and the performance is analyzed Chao Zhou 0002, Zhiqiang Cao 0002, Shuo Wang 0001, Min Tan 0001 |
IROS | 3 |
| 2006 | Cooperative hunting by distributed mobile robots based on local interactionabstractThis paper proposes a distributed control approach called local interactions with local coordinate systems (LILCS)to multirobot hunting tasks in unknown environments, where a team of mobile robots hunts a target called evader, which will actively try to escape with a safety strategy. This robust approach can cope with accumulative errors of wheels and imperfect communication networks. Computer simulations show the validity of the proposed approach. Zhiqiang Cao 0002, Min Tan 0001, Nong Gu, Shuo Wang 0001 |
IEEE Trans. Robotics | 5 |
| 2004 | Development of a biomimetic robotic fish and its control algorithmabstractThis paper is concerned with the design of a robotic fish and its motion control algorithms. A radio-controlled, four-link biomimetic robotic fish is developed using a flexible posterior body and an oscillating foil as a propeller. The swimming speed of the robotic fish is adjusted by modulating joint's oscillating frequency, and its orientation is tuned by different joint's deflections. Since the motion control of a robotic fish involves both hydrodynamics of the fluid environment and dynamics of the robot, it is very difficult to establish a precise mathematical model employing purely analytical methods. Therefore, the fish's motion control task is decomposed into two control systems. The online speed control implements a hybrid control strategy and a proportional-integral-derivative (PID) control algorithm. The orientation control system is based on a fuzzy logic controller. In our experiments, a point-to-point (PTP) control algorithm is implemented and an overhead vision system is adopted to provide real-time visual feedback. The experimental results confirm the effectiveness of the proposed algorithms. Junzhi Yu 0001, Min Tan 0001, Shuo Wang 0001, Erkui Chen |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2003 | Formation constrained multi-robot system in unknown environmentsabstractThis paper explores the application of the behavior-based approach to path planning for multiple mobile robots performing a formation control task in unknown environments. To predict the positions of moving obstacles for the purpose of collision avoidance, parabola prediction model whose parameters are estimated by the recurrence least square algorithm with restricted scale is adopted. Then, on the basis of the task and environment, we adopt five primitive behaviors and design a series of generation functions to generate control parameters for behaviors' combination. Furthermore, as the outputs of these functions can be adjusted according to the current situation, thus robots can achieve a motion strategy by reasonably combining behaviors and the adaptability to the environment is improved. We illustrate the validity of the approach by the simulations. Zhiqiang Cao 0002, Liangjun Xie, Bin Zhang 0010, Shuo Wang 0001, Min Tan 0001 |
ICRA | 4 |
| 2002 | Relation between task-based diversity and efficiency in multi-robot foragingabstractIn this paper, we present a new foraging algorithm applied to multi-robot system. The algorithm classifies robots into two groups; one group is engaged in navigating, while the other in collecting. Due to different duties of robots in foraging, the robots team displays diversity during the foraging process. The results of simulation show that the team of greater diversity would be more efficient than the team without diversity in accomplishing the foraging task. Under the consideration of such factors as the robot team's size and the density of attractors in environment, the relationship between the duty-based diversity and the performance of multi-robot system is discussed based on simulation results. Shuo Wang 0001, Min Tan 0001 |
ICARCV | 2 |