Zhengtong Xu

dblp:367/5671 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-2789-1910ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Motion planning and robot control · 45% Robot manipulation · 24% 3D vision · 19%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
1.322026
Canonical Policy: Learning Canonical 3-D Representation for $\text{SE}(3)$-Equivariant Policy · IEEE Trans. Robotics 2026
DiffOG: Differentiable Policy Trajectory Optimization With Generalizability · IEEE Trans. Robotics 2026
Computer vision › 3D vision
3d shape representation
1.012026
Canonical Policy: Learning Canonical 3-D Representation for $\text{SE}(3)$-Equivariant Policy · IEEE Trans. Robotics 2026
Computer vision › 3D vision › 3d shape representation
canonical representation
1.012026
Canonical Policy: Learning Canonical 3-D Representation for $\text{SE}(3)$-Equivariant Policy · IEEE Trans. Robotics 2026
Robotics › Motion planning and robot control › trajectory optimization
constrained trajectory optimization
1.012026
DiffOG: Differentiable Policy Trajectory Optimization With Generalizability · IEEE Trans. Robotics 2026
Robotics › Motion planning and robot control › trajectory optimization
gradient-based trajectory optimization
1.012026
DiffOG: Differentiable Policy Trajectory Optimization With Generalizability · IEEE Trans. Robotics 2026
Robotics › Motion planning and robot control
trajectory optimization
1.012026
DiffOG: Differentiable Policy Trajectory Optimization With Generalizability · IEEE Trans. Robotics 2026
Robotics › Motion planning and robot control › robot learning
visuomotor policy
1.012026
DiffOG: Differentiable Policy Trajectory Optimization With Generalizability · IEEE Trans. Robotics 2026
Robotics › Robot manipulation
grasping
0.812024
LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping · IEEE Trans. Robotics 2024
Robotics › Motion planning and robot control › robot control › model predictive control
learning-based model predictive control
0.812024
LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping · IEEE Trans. Robotics 2024
Robotics › Robot manipulation
tactile sensing
0.812024
LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping · IEEE Trans. Robotics 2024
Robotics › Robot manipulation › tactile sensing
vision-based tactile sensing
0.812024
LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping · IEEE Trans. Robotics 2024
Robotics › Robot manipulation › grasping
grasp control
0.212024
LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping · IEEE Trans. Robotics 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.0point cloud · 1.0generative model · 1.0differentiable optimization layers · 1.0SE(3) equivariance · 1.0neural network · 0.8model predictive control · 0.8differentiable MPC layer · 0.8
YearPublicationVenuePosition
2026 DiffOG: Differentiable Policy Trajectory Optimization With Generalizability
abstract
Imitation learning-based visuomotor policies excel at manipulation tasks but often produce suboptimal action trajectories compared to model-based methods. Directly mapping camera data to actions via neural networks can result in jerky motions and difficulties in meeting critical constraints, compromising safety and robustness in real-world deployment. For tasks that require high robustness or strict adherence to constraints, ensuring trajectory quality is crucial. However, the lack of interpretability in neural networks makes it challenging to generate constraint-compliant actions in a controlled manner. This paper introduces differentiable policy trajectory optimization with generalizability (DiffOG), a learning-based trajectory optimization framework designed to enhance visuomotor policies. By leveraging the proposed differentiable formulation of trajectory optimization with transformer, DiffOG seamlessly integrates policies with a generalizable optimization layer. DiffOG refines action trajectories to be smoother and more constraint-compliant while maintaining alignment with the original demonstration distribution, thus avoiding degradation in policy performance. We evaluated DiffOG across 11 simulated tasks and 2 real-world tasks. The results demonstrate that DiffOG significantly enhances the trajectory quality of visuomotor policies while having minimal impact on policy performance, outperforming trajectory processing baselines such as greedy constraint clipping and penalty-based trajectory optimization. Furthermore, DiffOG achieves superior performance compared to existing constrained visuomotor policy.
Zhengtong Xu, Zichen Miao, Qiang Qiu 0001, Yu She
IEEE Trans. Robotics1
2026 Canonical Policy: Learning Canonical 3-D Representation for $\text{SE}(3)$-Equivariant Policy
abstract
Visual imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy architectures to exploit spatial symmetries. However, existing equivariant approaches often lack interpretability and rigor due to unstructured integration of equivariant components. We introduce canonical policy, a principled framework for 3D equivariant imitation learning that unifies 3D point cloud observations under a canonical representation. We first establish a theory of 3D canonical representations, enabling equivariant observation-to-action mappings by grouping both seen and novel point clouds to a canonical representation. We then propose a flexible policy learning pipeline that leverages geometric symmetries from canonical representation and the expressiveness of modern generative models. We validate canonical policy on 12 diverse simulated tasks and 4 real-world manipulation tasks across 16 configurations, involving variations in object color, shape, camera viewpoint, and robot platform. Compared to state-of-the-art imitation learning policies, canonical policy achieves an average improvement of 18.0% in simulation and 39.7% in real-world experiments, demonstrating superior generalization capability and sample efficiency.
Zhiyuan Zhang 0014, Zhengtong Xu, Jai Nanda Lakamsani, Yu She
IEEE Trans. Robotics2
2025 LeTO: Learning Constrained Visuomotor Policy With Differentiable Trajectory Optimization
abstract
This paper introduces LeTO, a method for learning constrained visuomotor policy with differentiable trajectory optimization. Our approach integrates a differentiable optimization layer into the neural network. By formulating the optimization layer as a trajectory optimization problem, we enable the model to end-to-end generate actions in a safe and constraint-controlled fashion without extra modules. Our method allows for the introduction of constraint information during the training process, thereby balancing the training objectives of satisfying constraints, smoothing the trajectories, and minimizing errors with demonstrations. This “gray box” method marries optimization-based safety and interpretability with powerful representational abilities of neural networks. We quantitatively evaluate LeTO in simulation and in the real robot. The results demonstrate that LeTO performs well in both simulated and real-world tasks. In addition, it is capable of generating trajectories that are less uncertain, higher quality, and smoother compared to existing imitation learning methods. Therefore, it is shown that LeTO provides a practical example of how to achieve the integration of neural networks with trajectory optimization. We release our code athttps://github.com/ZhengtongXu/LeTO. Note to Practitioners—LeTO is driven by the goal of developing an imitation learning algorithm capable of generating safe and constraint-satisfying robotic behaviors. The idea of imitation learning is to enable the robot to learn from human demonstrations of certain tasks. Subsequently, the robot is able to autonomously perform the learned tasks on its own. Thanks to the powerful representational and fitting capabilities of neural networks, imitation learning can let robots perform complex manipulation tasks. However, neural networks often exhibit a certain level of uncertainty and lack theoretical safety guarantees. For robotic systems, it is crucial that robot behaviors meet specific constraints; otherwise, the system may not be sufficiently reliable. Therefore, we introduce LeTO, an approach that integrates trajectory optimization with neural networks to generate actions that not only achieve manipulation tasks, but also comply with constraints. This improves the interpretability, safety, and reliability of robot policies acquired through imitation learning, facilitating their deployment in scenarios with high safety requirements.
Zhengtong Xu, Yu She
IEEE Trans Autom. Sci. Eng.1
2024 LeTac-MPC: Learning Model Predictive Control for Tactile-Reactive Grasping
abstract
Grasping is a crucial task in robotics, necessitating tactile feedback and reactive grasping adjustments for robust grasping of objects under various conditions and with differing physical properties. In this article, we introduce LeTac-MPC, a learning-based model predictive control (MPC) for tactile-reactive grasping. Our approach enables the gripper to grasp objects with different physical properties on dynamic and force-interactive tasks. We utilize a vision-based tactile sensor, GelSight (Yuan et al. 2017), which is capable of perceiving high-resolution tactile feedback that contains information on the physical properties and states of the grasped object. LeTac-MPC incorporates a differentiable MPC layer designed to model the embeddings extracted by a neural network from tactile feedback. This design facilitates convergent and robust grasping control at a frequency of 25 Hz. We propose a fully automated data collection pipeline and collect a dataset only using standardized blocks with different physical properties. However, our trained controller can generalize to daily objects with different sizes, shapes, materials, and textures. The experimental results demonstrate the effectiveness and robustness of the proposed approach. We compare LeTac-MPC with two purely model-based tactile-reactive controllers (MPC and PD) and open-loop grasping. Our results show that LeTac-MPC has optimal performance in dynamic and force-interactive tasks and optimal generalizability.
Zhengtong Xu, Yu She
IEEE Trans. Robotics1