Yu She

dblp:128/0519 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-5914-3573ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Systems, architecture and hardware · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 DiffOG: Differentiable Policy Trajectory Optimization With Generalizability
abstract
Imitation learning-based visuomotor policies excel at manipulation tasks but often produce suboptimal action trajectories compared to model-based methods. Directly mapping camera data to actions via neural networks can result in jerky motions and difficulties in meeting critical constraints, compromising safety and robustness in real-world deployment. For tasks that require high robustness or strict adherence to constraints, ensuring trajectory quality is crucial. However, the lack of interpretability in neural networks makes it challenging to generate constraint-compliant actions in a controlled manner. This paper introduces differentiable policy trajectory optimization with generalizability (DiffOG), a learning-based trajectory optimization framework designed to enhance visuomotor policies. By leveraging the proposed differentiable formulation of trajectory optimization with transformer, DiffOG seamlessly integrates policies with a generalizable optimization layer. DiffOG refines action trajectories to be smoother and more constraint-compliant while maintaining alignment with the original demonstration distribution, thus avoiding degradation in policy performance. We evaluated DiffOG across 11 simulated tasks and 2 real-world tasks. The results demonstrate that DiffOG significantly enhances the trajectory quality of visuomotor policies while having minimal impact on policy performance, outperforming trajectory processing baselines such as greedy constraint clipping and penalty-based trajectory optimization. Furthermore, DiffOG achieves superior performance compared to existing constrained visuomotor policy.
Zhengtong Xu, Zichen Miao, Qiang Qiu 0001, Yu She
IEEE Trans. Robotics5
2026 Canonical Policy: Learning Canonical 3-D Representation for $\text{SE}(3)$-Equivariant Policy
abstract
Visual imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy architectures to exploit spatial symmetries. However, existing equivariant approaches often lack interpretability and rigor due to unstructured integration of equivariant components. We introduce canonical policy, a principled framework for 3D equivariant imitation learning that unifies 3D point cloud observations under a canonical representation. We first establish a theory of 3D canonical representations, enabling equivariant observation-to-action mappings by grouping both seen and novel point clouds to a canonical representation. We then propose a flexible policy learning pipeline that leverages geometric symmetries from canonical representation and the expressiveness of modern generative models. We validate canonical policy on 12 diverse simulated tasks and 4 real-world manipulation tasks across 16 configurations, involving variations in object color, shape, camera viewpoint, and robot platform. Compared to state-of-the-art imitation learning policies, canonical policy achieves an average improvement of 18.0% in simulation and 39.7% in real-world experiments, demonstrating superior generalization capability and sample efficiency.
Zhiyuan Zhang 0014, Zhengtong Xu, Jai Nanda Lakamsani, Yu She
IEEE Trans. Robotics4
2025 A Magnetic-Actuated Vision-Based Whisker Array for Contact Perception and Grasping
abstract
Tactile sensing and the manipulation of delicate objects are critical challenges in robotics. This study presents a vision-based magnetic-actuated whisker array sensor that integrates these functions. The sensor features eight whiskers arranged circularly, supported by an elastomer membrane and actuated by electromagnets and permanent magnets. A camera tracks whisker movements, enabling high-resolution tactile feedback. The sensor's performance was evaluated through object classification and grasping experiments. In the classification experiment, the sensor approached objects from four directions and accurately identified five distinct objects with a classification accuracy of 99.17% using a Multi-Layer Perceptron model. In the grasping experiment, the sensor tested configurations of eight, four, and two whiskers, achieving the highest success rate of 87% with eight whiskers. These results highlight the sensor's potential for precise tactile sensing and reliable manipulation.
Zhixian Hu, Juan P. Wachs, Yu She
ICRA3
2025 DartBot: Overhand Throwing of Deformable Objects With Tactile Sensing and Reinforcement Learning
abstract
Object transfer through throwing is a classic dynamic manipulation task that necessitates precise control and perception capabilities. However, developing dynamic models for unstructured environments using analytical methods presents challenges. In this study, we present DartBot, a robot that integrates tactile exploration and reinforcement learning to achieve robust throwing skills for nonrigid relatively small objects under the influence of moment of inertia which cause the object to spin in the air. Unlike traditional sim-to-real transfer methods, our approach involves direct training of the agent on a real hardware robot equipped with a high-resolution tactile sensor, enabling reinforced learning in a realistic and dynamic environment. By leveraging tactile perception, we incorporate pseudo-embeddings of the physical properties of objects into the learning process through tilting actions at two distinct angles. This tactile information enables the agent to infer and adapt its throwing strategy, resulting in improved accuracy when handling various objects and targeting distant locations. Furthermore, we demonstrate that the quality of a grasp significantly impacts the success rate of the throwing task. We evaluate the effectiveness of our method through extensive experiments, demonstrating superior performance and generalization capabilities in real-world throwing scenarios. We achieved a success rate of 95% for unseen objects with a mean error of 3.15 cm from the goal. A high-resolution video demo of our work is available athttps://youtu.be/KNFgDeLt-0g. Note to Practitioners—The industrial demand for precise and accurate object transfer beyond the robot’s maximum kinematic range is rapidly growing, necessitating advancements in robotic throwing manipulation to efficiently utilize resources. In this context, this paper contributes to the field by presenting a method for the transfer of deformable objects through overhand throwing, addressing the challenges of achieving precise control and perception capabilities by learning the complex physics of the task using raw data. The approach offers valuable insights and techniques for enhancing throwing manipulation tasks. The integration of high-resolution tactile sensing and reinforcement learning, while considering the influence of moment of inertia, opens up new possibilities for handling deformable objects in throwing tasks and developing dynamic models for such unstructured environments. The learned throwing policy is applied to a variety of previously unseen objects, demonstrating its effectiveness and ability to generalize in real-world throwing scenarios. The presented work holds applicability in various domains such as medical rehabilitation, logistic warehouses for packaging, and handling of urban waste. It shows promise in improving the precision, accuracy and efficiency of the throwing task. In the future, we aim to expand the current overhand throwing framework by incorporating rapid grasping in cluttered environments and enabling throwing to far distances with diverse target locations.
Shoaib Aslam, Krish Kumar, Pokuang Zhou, Hongyu Yu, Michael Yu Wang, Yu She
IEEE Trans Autom. Sci. Eng.6
2025 LeTO: Learning Constrained Visuomotor Policy With Differentiable Trajectory Optimization
abstract
This paper introduces LeTO, a method for learning constrained visuomotor policy with differentiable trajectory optimization. Our approach integrates a differentiable optimization layer into the neural network. By formulating the optimization layer as a trajectory optimization problem, we enable the model to end-to-end generate actions in a safe and constraint-controlled fashion without extra modules. Our method allows for the introduction of constraint information during the training process, thereby balancing the training objectives of satisfying constraints, smoothing the trajectories, and minimizing errors with demonstrations. This “gray box” method marries optimization-based safety and interpretability with powerful representational abilities of neural networks. We quantitatively evaluate LeTO in simulation and in the real robot. The results demonstrate that LeTO performs well in both simulated and real-world tasks. In addition, it is capable of generating trajectories that are less uncertain, higher quality, and smoother compared to existing imitation learning methods. Therefore, it is shown that LeTO provides a practical example of how to achieve the integration of neural networks with trajectory optimization. We release our code athttps://github.com/ZhengtongXu/LeTO. Note to Practitioners—LeTO is driven by the goal of developing an imitation learning algorithm capable of generating safe and constraint-satisfying robotic behaviors. The idea of imitation learning is to enable the robot to learn from human demonstrations of certain tasks. Subsequently, the robot is able to autonomously perform the learned tasks on its own. Thanks to the powerful representational and fitting capabilities of neural networks, imitation learning can let robots perform complex manipulation tasks. However, neural networks often exhibit a certain level of uncertainty and lack theoretical safety guarantees. For robotic systems, it is crucial that robot behaviors meet specific constraints; otherwise, the system may not be sufficiently reliable. Therefore, we introduce LeTO, an approach that integrates trajectory optimization with neural networks to generate actions that not only achieve manipulation tasks, but also comply with constraints. This improves the interpretability, safety, and reliability of robot policies acquired through imitation learning, facilitating their deployment in scenarios with high safety requirements.
Zhengtong Xu, Yu She
IEEE Trans Autom. Sci. Eng.2
2024 Stick Roller: Precise In-hand Stick Rolling with a Sample-Efficient Tactile Model
abstract
In-hand manipulation is challenging in robotics due to the intricate contact dynamics and high degrees of control freedom. Precise manipulation with high accuracy often requires tactile perception, which adds further complexity to the system. Despite the challenges in perception and control, the rolling stick problem is an essential and practical motion primitive with many demanding industrial applications. This work aims to learn the high-resolution tactile dynamics of the rolling stick. Specifically, we try manipulating a small stick using the Allegro hand equipped with the Digit vision-based tactile sensor. The learning framework includes an action filtering module, tactile perception module, and learning with uncertainty module, all designed to operate in low data regimes. With only 2.3% amount of data and 5.7% model complexity of previous similar work, our learned contact dynamics model achieves better grasp stability, sub-millimeter precision, and promising zero-shot generalizability across novel objects. The proposed framework demonstrates the potential for precise in-hand manipulation with tactile feedback on real hardware. The project source code is available at: https://github.com/duyipai/Allegro_Digit. A video presentation is available here.
Yipai Du, Pokuang Zhou, Michael Yu Wang, Wenzhao Lian, Yu She
IROS5
2024 SeeBelow: Sub-dermal 3D Reconstruction of Tumors with Surgical Robotic Palpation and Tactile Exploration
abstract
Surgical scene understanding in Robot-assisted Minimally Invasive Surgery (RMIS) is highly reliant on visual cues and lacks tactile perception. Force-modulated surgical palpation with tactile feedback is necessary for localization, geometry/depth estimation, and dexterous exploration of abnormal stiff inclusions in subsurface tissue layers. Prior works explored surface-level tissue abnormalities or single layered tissue-tumor embeddings with more than 300 palpations for dense 2D stiffness mapping. Our approach focuses on 3D reconstructions of sub-dermal tumor surface profiles in multi-layered tissue (skin-fat-muscle) using a visually-guided novel tactile navigation policy. A robotic palpation probe with triaxial force sensing was leveraged for tactile exploration of the phantom. From a surface mesh of the surgical region initialized from a depth camera, the policy explores a surgeon’s region of interest through palpation, sampled from bayesian optimization. Each palpation includes contour following using a contact-safe impedance controller to trace the sub-dermal tumor geometry, until the underlying tumor-tissue boundary is reached. Projections of these contour following palpation trajectories allows 3D reconstruction of the subdermal tumor surface profile in less than 100 palpations. Our approach generates high-fidelity 3D surface reconstructions of rigid tumor embeddings in tissue layers with isotropic elasticities, although soft tumor geometries are yet to be explored. For more details, please refer to our open-source codebase1and project website2.
Raghava Uppuluri, Abhinaba Bhattacharjee, Sohel Anwar, Yu She
IROS4
2024 Feelit: Combining Compliant Shape Displays with Vision-Based Tactile Sensors for Real-Time Teletaction
abstract
Teletaction, the transmission of tactile feedback or touch, is a crucial aspect in the field of teleoperation. High-quality teletaction feedback allows users to remotely manipulate objects and increase the quality of the humanmachine interface between the operator and the robot, making complex manipulation tasks possible. Advances in the field of teletaction for teleoperation however, have yet to make full use of the high-resolution 3D data provided by modern vision-based tactile sensors. Existing solutions for teletaction lack in one or more areas of form or function, such as fidelity or hardware footprint. In this paper, we showcase our design for a low-cost teletaction device that can utilize real-time high-resolution tactile information from vision-based tactile sensors, through both physical 3D surface reconstruction and shear displacement. We present our device, the Feelit, which uses a combination of a pin-based shape display and compliant mechanisms to accomplish this task. The pin-based shape display utilizes an array of 24 servomotors with miniature Bowden cables, giving the device a resolution of 6x4 pins in a 15x10 mm display footprint. Each pin can actuate up to 3 mm in 200 ms, while providing 80 N of force and 1.5 um of depth resolution. Shear displacement and rotation is achieved using a compliant mechanism design, allowing a minimum of 1 mm displacement laterally and 10 degrees of rotation. This real-time 3D tactile reconstruction is achieved with the use of a vision-based tactile sensor, the GelSight [1], along with an algorithm that samples the depth data and marker tracking to generate actuator commands. Through a series of experiments including shape recognition and relative weight identification, we show that our device has the potential to expand teletaction capabilities in the teleoperation space.
Oscar Yu, Yu She
IROS2
2024 In-Hand Singulation and Scooping Manipulation with a 5 DOF Tactile Gripper
abstract
Manipulation tasks often require a high degree of dexterity, typically necessitating grippers with multiple degrees of freedom (DoF). While a robotic hand equipped with multiple fingers can execute precise and intricate manipulation tasks, the inherent redundancy stemming from its extensive DoF often adds unnecessary complexity. In this paper, we introduce the design of a tactile sensor-equipped gripper with two fingers and five DoF. We present a novel design integrating a GelSight tactile sensor, enhancing sensing capabilities and enabling finer control during specific manipulation tasks. To evaluate the gripper’s performance, we conduct experiments involving two challenging tasks: 1) retrieving, singularizing, and classification of various objects embedded in granular media, and 2) executing scooping manipulations of credit cards in confined environments to achieve precise insertion. Our results demonstrate the efficiency of the proposed approach, with a high success rate for singulation and classification tasks, particularly for spherical objects at high as 94.3%, and a 100% success rate for scooping and inserting credit cards.
Pokuang Zhou, Shaoxiong Wang, Yu She
IROS4
2024 LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping
abstract
Grasping is a crucial task in robotics, necessitating tactile feedback and reactive grasping adjustments for robust grasping of objects under various conditions and with differing physical properties. In this article, we introduce LeTac-MPC, a learning-based model predictive control (MPC) for tactile-reactive grasping. Our approach enables the gripper to grasp objects with different physical properties on dynamic and force-interactive tasks. We utilize a vision-based tactile sensor, GelSight (Yuan et al. 2017), which is capable of perceiving high-resolution tactile feedback that contains information on the physical properties and states of the grasped object. LeTac-MPC incorporates a differentiable MPC layer designed to model the embeddings extracted by a neural network from tactile feedback. This design facilitates convergent and robust grasping control at a frequency of 25 Hz. We propose a fully automated data collection pipeline and collect a dataset only using standardized blocks with different physical properties. However, our trained controller can generalize to daily objects with different sizes, shapes, materials, and textures. The experimental results demonstrate the effectiveness and robustness of the proposed approach. We compare LeTac-MPC with two purely model-based tactile-reactive controllers (MPC and PD) and open-loop grasping. Our results show that LeTac-MPC has optimal performance in dynamic and force-interactive tasks and optimal generalizability.
Zhengtong Xu, Yu She
IEEE Trans. Robotics2
2021 GelSight Wedge: Measuring High-Resolution 3D Contact Geometry with a Compact Robot Finger
abstract
Vision-based tactile sensors have the potential to provide important contact geometry to localize the objective with visual occlusion. However, it is challenging to measure high-resolution 3D contact geometry for a compact robot finger, to simultaneously meet optical and mechanical constraints. In this work, we present the GelSight Wedge sensor, which is optimized to have a compact shape for robot fingers, while achieving high-resolution 3D reconstruction. We evaluate the 3D reconstruction under different lighting configurations, and extend the method from 3 lights to 1 or 2 lights. We demonstrate the flexibility of the design by shrinking the sensor to the size of a human finger for fine manipulation tasks. We also show the effectiveness and potential of the reconstructed 3D geometry for pose tracking in the 3D space.
Shaoxiong Wang, Yu She, Branden Romero, Edward H. Adelson
ICRA2
2020 Exoskeleton-covered soft finger with vision-based proprioception and tactile sensing
abstract
Soft robots offer significant advantages in adaptability, safety, and dexterity compared to conventional rigid-body robots. However, it is challenging to equip soft robots with accurate proprioception and tactile sensing due to their high flexibility and elasticity. In this work, we describe the development of a vision-based proprioceptive and tactile sensor for soft robots called GelFlex, which is inspired by previous GelSight sensing techniques. More specifically, we develop a novel exoskeleton-covered soft finger with embedded cameras and deep learning methods that enable high-resolution proprioceptive sensing and rich tactile sensing. To do so, we design features along the axial direction of the finger, which enable high-resolution proprioceptive sensing, and incorporate a reflective ink coating on the surface of the finger to enable rich tactile sensing. We design a highly underactuated exoskeleton with a tendon-driven mechanism to actuate the finger. Finally, we assemble 2 of the fingers together to form a robotic gripper and successfully perform a bar stock classification task, which requires both shape and tactile information. We train neural networks for proprioception and shape (box versus cylinder) classification using data from the embedded sensors. The proprioception CNN had over 99% accuracy on our testing set (all six joint angles were within 1° of error) and had an average accumulative distance error of 0.77 mm during live testing, which is better than human finger proprioception. These proposed techniques offer soft robots the high-level ability to simultaneously perceive their proprioceptive state and peripheral environment, providing potential solutions for soft robots to solve everyday manipulation tasks. We believe the methods developed in this work can be widely applied to different designs and applications.
Yu She, Sandra Q. Liu, Peiyu Yu, Edward H. Adelson
ICRA1
2017 On the impact force of human-robot interaction: Joint compliance vs. link compliance
abstract
In this paper, we study the effect of mechanical compliance on the impact force of human-robot interactions, more specifically the maximum impact force during a collision. Here we consider two methods of introducing compliance to industrial manipulators: joint compliance and link compliance. To compare their effect on the maximum impact force, we study two designs of a 2D robot link: a rigid link with torsion spring at the joint and a uniform compliant link. The dynamic impact model is based on the Hertz contact model. The results show that the compliant joint solution could produce a larger impact force than that of the compliant link solution if the arm mass is larger than that of the end mass, given the same lateral stiffness and all other inertial parameters (e.g. mass). Simulations and experiment have been done and verified this conclusion. The research demonstrates that the compliant link solution could be a promising approach for addressing safety concerns of human robot interactions.
Yu She, Deshan Meng, Junxiao Cui
ICRA1
2015 A transformable wheel robot with a passive leg
abstract
In this paper, we present a novel wheeled robot that transforms from a circled configuration to a spoke-like legged configuration. Wheeled mobile robots are able to quickly and efficiently move on flat surfaces. However they may sink into a terrain with dynamic surface such as snow, sand, dirt, moss, and small gravel. The transformable wheel robot presented in this paper is able to overcome obstacles and navigate these dynamics surfaces, yet can still move quickly over flat ground. The wheel is comprised of five spokes or legs, four of which are actively driven by a motor via four slider-crank linkages. The last one is designed be passive in order to significantly decrease the actuation force of the transformer mechanism. Analysis shows that the maximum driving force required to actuate the passive leg design is 1/187.5 times that of the active leg design. Experiments demonstrate that the passive leg design allows the mobile robot to open the transformation mechanism with a maximum external loading of 1.809 times the original robot weight. The diameter of the legged configuration wheel is 1.576 times that of the circled configuration.
Yu She, Carter J. Hurd
IROS1
2015 Dynamic modeling of a 2D compliant link for safety evaluation in human-robot interactions
abstract
In this article, we build dynamic models of 2D compliant links to evaluate injury level in a human-robot interaction. Safety is a premium concern for co-robotic systems. It has been studied that using compliant links in a robot can greatly reduce the injury level. Since most safety criteria are based on tolerance of acceleration of the operator's head during the impact, an efficient and yet accurate dynamic model of compliant links is needed. In this paper, we compare three dynamic models for calculating acceleration and head injury criterion: the compliant Beam-Spring-Mass (BSM) model, the Mass-Spring-Mass (MSM) model and the Link-Spring-Mass (LSM) model. For the MSM model and LSM model, we obtain analytic expressions of acceleration. While numerical results are achieved for the compliant BSM model. To develop the compliant BSM model, we compared three different methods: the Pseudo-Rigid-Body (PRB) model, the Finite-Segment-Model (FSM), and the Assumed-Mode-Method (AMM). Finally, all these models are validated by human-robot impact simulation programs built in Matlab. The acceleration from these simulations can be used to quantitatively measure the injury level during an impact.
Yu She, Deshan Meng, Hongliang Shi
IROS1