Jianshu Hu

dblp:337/1942 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0006-2837-425XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 52% Motion planning and robot control · 16% Learning theory · 12%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
sample efficiency
1.622025
Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement Learning · NeurIPS 2025
Revisiting Data Augmentation in Deep Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
symmetry exploitation
0.912025
Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.812024
Revisiting Data Augmentation in Deep Reinforcement Learning · ICLR 2024
Machine learning › Learning theory
generalization
0.812024
Revisiting Data Augmentation in Deep Reinforcement Learning · ICLR 2024
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion
0.812024
Beyond Inverted Pendulums: Task-Optimal Simple Models of Legged Locomotion · IEEE Trans. Robotics 2024
Robotics › Motion planning and robot control › dynamic modeling
reduced-order model
0.812024
Beyond Inverted Pendulums: Task-Optimal Simple Models of Legged Locomotion · IEEE Trans. Robotics 2024
Machine learning › Efficient and distributed learning
model optimization
0.212024
Beyond Inverted Pendulums: Task-Optimal Simple Models of Legged Locomotion · IEEE Trans. Robotics 2024
Robotics › Motion planning and robot control
trajectory optimization
0.212024
Beyond Inverted Pendulums: Task-Optimal Simple Models of Legged Locomotion · IEEE Trans. Robotics 2024

Methods — techniques the papers use, named apart from their topics

trajectory reversal augmentation · 0.9reward shaping · 0.9tangent-prop · 0.8simulation · 0.8regularization · 0.8model optimization algorithm · 0.8data augmentation · 0.8
YearPublicationVenuePosition
2025 Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement Learning
abstract
Symmetry is pervasive in robotics and has been widely exploited to improve sample efficiency in deep reinforcement learning (DRL). However, existing approaches primarily focus on spatial symmetries—such as reflection, rotation, and translation—while largely neglecting temporal symmetries. To address this gap, we explore time reversal symmetry, a form of temporal symmetry commonly found in robotics tasks such as door opening and closing. We propose Time Reversal symmetry enhanced Deep Reinforcement Learning (TR-DRL), a framework that combines trajectory reversal augmentation and time reversal guided reward shaping to efficiently solve temporally symmetric tasks. Our method generates reversed transitions from fully reversible transitions, identified by a proposed dynamics-consistent filter, to augment the training data. For partially reversible transitions, we apply reward shaping to guide learning, according to successful trajectories from the reversed task. Extensive experiments on the Robosuite and MetaWorld benchmarks demonstrate that TR-DRL is effective in both single-task and multi-task settings, achieving higher sample efficiency and stronger final performance compared to baseline methods.
Yunpeng Jiang, Jianshu Hu, Paul Weng, Yutong Ban
NeurIPS2
2025 State-novelty guided action persistence in deep reinforcement learning
Jianshu Hu, Paul Weng, Yutong Ban
Mach. Learn.1
2024 Revisiting Data Augmentation in Deep Reinforcement Learning
abstract
Various data augmentation techniques have been recently proposed in image-based deep reinforcement learning (DRL). Although they empirically demonstrate the effectiveness of data augmentation for improving sample efficiency or generalization, which technique should be preferred is not always clear. To tackle this question, we analyze existing methods to better understand them and to uncover how they are connected. Notably, by expressing the variance of the Q-targets and that of the empirical actor/critic losses of these methods, we can analyze the effects of their different components and compare them. We furthermore formulate an explanation about how these methods may be affected by choosing different data augmentation transformations in calculating the target Q-values. This analysis suggests recommendations on how to exploit data augmentation in a more principled way. In addition, we include a regularization term called tangent prop, previously proposed in computer vision, but whose adaptation to DRL is novel to the best of our knowledge. We evaluate our proposition and validate our analysis in several domains. Compared to different relevant baselines, we demonstrate that it achieves state-of-the-art performance in most environments and shows higher sample efficiency and better generalization ability in some complex environments.
Jianshu Hu, Yunpeng Jiang, Paul Weng
ICLR1
2024 Beyond Inverted Pendulums: Task-Optimal Simple Models of Legged Locomotion
abstract
Reduced-order models (ROM) are popular in online motion planning due to their simplicity. A good ROM for control captures critical task-relevant aspects of the full dynamics while remaining low dimensional. However, planning within the reduced-order space unavoidably constrains the full model, and hence we sacrifice the full potential of the robot. In the community of legged locomotion, this has lead to a search for better model extensions, but many of these extensions require human intuition, and there has not existed a principled way of evaluating the model performance and discovering new models. In this work, we propose a model optimization algorithm that automatically synthesizes reduced-order models, optimal with respect to a user-specified distribution of tasks and corresponding cost functions. To demonstrate our work, we optimized models for a bipedal robot Cassie. We show in simulation that the optimal ROM reduces the cost of Cassie's joint torques by up to 23% and increases its walking speed by up to 54%. We also show hardware result that the real robot walks on flat ground with 10% lower torque cost. All videos and code can be found athttps://sites.google.com/view/ymchen/research/optimal-rom.
Yu-Ming Chen 0003, Jianshu Hu, Michael Posa
IEEE Trans. Robotics2