Shangke Lyu

dblp:226/6149 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8302-6630ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Activation-wise Propagation: A One-Timestep Strategy for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have demonstrated significant potential in real-time multi-sensor perception tasks due to their event-driven and parameter-efficient characteristics. A key challenge is the timestep-wise iterative update of neuronal hidden states (membrane potentials), which complicates the trade-off between accuracy and latency. SNNs tend to achieve better performance with longer timesteps, inevitably resulting in higher computational overhead and latency compared to artificial neural networks (ANNs). Moreover, many recent advances in SNNs rely on architecture-specific optimizations, which, while effective with fewer timesteps, often limit generalizability and scalability across modalities and models. To address these limitations, we propose Activation-wise Membrane Potential Propagation (AMP2), a unified hidden state update mechanism for SNNs. Inspired by the spatial propagation of membrane potentials in biological neurons, AMP2 enables dynamic transmission of membrane potentials among spatially adjacent neurons, facilitating spatiotemporal integration and cooperative dynamics of hidden states, thereby improving efficiency and accuracy while reducing reliance on extended temporal updates. This simple yet effective strategy significantly enhances SNN performance across various architectures, including MLPs and CNNs for point cloud and event-based data. Furthermore, ablation studies integrating AMP2 into Transformer-based SNNs for classification tasks demonstrate its potential as a general-purpose and efficient solution for spiking neural networks.
Xiangfei Yang, Shangke Lyu
AAAI3
2025 CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
abstract
In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from complex constraints. In this paper, we introduce Coarse-to-Fine AutoRegressive Policy (CARP), a novel paradigm for visuomotor policy learning that redefines the autoregressive action generation process as a coarse-to-fine, next-scale approach. CARP decouples action generation into two stages: first, an action autoencoder learns multi-scale representations of the entire action sequence; then, a GPT-style transformer refines the sequence prediction through a coarse-to-fine autoregressive process. This straightforward and intuitive approach produces highly accurate and smooth actions, matching or even surpassing the performance of diffusion-based policies while maintaining efficiency on par with autoregressive policies. We conduct extensive evaluations across diverse settings, including single-task and multi-task scenarios on state-based and image-based simulation benchmarks, as well as real-world tasks. CARP achieves competitive success rates, with up to a 10% improvement, and delivers 10x faster inference compared to state-of-the-art policies, establishing a high-performance, efficient, and flexible paradigm for action generation in robotic tasks.
Zhefei Gong, Pengxiang Ding, Shangke Lyu, Siteng Huang, Zhaoxin Fan
ICCV3
2025 GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
abstract
With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the inevitable external perturbations encountered during deployment. These perturbations introduce unforeseen state information to the VLA, resulting in inaccurate actions and consequently, a significant decline in generalization performance. The classic internal model control (IMC) principle demonstrates that a closed-loop system with an internal model that includes external input signals can accurately track the reference input and effectively offset the disturbance. We propose a novel closed-loop VLA method GEVRM that integrates the IMC principle to enhance the robustness of robot visual manipulation. The text-guided video generation model in GEVRM can generate highly expressive future visual planning goals. Simultaneously, we evaluate perturbations by simulating responses, which are called internal embeddings and optimized through prototype contrastive learning. This allows the model to implicitly infer and distinguish perturbations from the external environment. The proposed GEVRM achieves state-of-the-art performance on both standard and perturbed CALVIN benchmarks and shows significant improvements in realistic robot tasks.
Hongyin Zhang 0001, Pengxiang Ding, Shangke Lyu
ICLR3
2025 Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning
abstract
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUARTOnline, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference at 50 Hz in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65 %. Our project page is https://quart-online.github.io.
Xinyang Tong, Pengxiang Ding, Yiguo Fan, Can Cui 0008, Han Zhao 0008, Hongyin Zhang 0001, Yonghao Dang, Siteng Huang, Shangke Lyu
ICRA12
2025 Integrating Trajectory Optimization and Reinforcement Learning for Quadrupedal Jumping with Terrain-Adaptive Landing
abstract
Jumping constitutes an essential component of quadruped robots’ locomotion capabilities, which includes dynamic take-off and adaptive landing. Existing quadrupedal jumping studies mainly focused on the stance and flight phase by assuming a flat landing ground, which is impractical in many real world cases. This work proposes a safe landing framework that achieves adaptive landing on rough terrains by combining Trajectory Optimization (TO) and Reinforcement Learning (RL) together. The RL agent learns to track the reference motion generated by TO in the environments with rough terrains. To enable the learning of compliant landing skills on challenging terrains, a reward relaxation strategy is synthesized to encourage exploration during landing recovery period. Extensive experiments validate the accurate tracking and safe landing skills benefiting from our proposed method in various scenarios.
Shangke Lyu, Xin Lang
IROS2
2025 Toward Air Operation Aerial Manipulator Control With a Refined Anti-Disturbance Architecture
abstract
Due to the presence of strong inner dynamic coupling and changes in center of mass (CoM) during tasks execution, the precise tracking control problem of aerial manipulator systems becomes challenging. When the mounted manipulator is performing a task, the movements of the manipulator will cause the inherent dynamic coupling force/torque disturbances. Such disturbances are quickly acted on the position loop and the orientation loop of the UAV body, to the detriment of the control accuracy. On the other hand, the fluctuation of the UAV body leads to the base floating, resulting in a significantly adverse influence on the precision of the manipulator end-effector. Since different disturbances have distinct mathematical properties, i.e., norm bounded and rate bounded, currently, there is no control framework that is able to tackle the dynamic coupling, model uncertainty and base floating simultaneously. In this paper, a refined anti-disturbance control architecture is proposed for the decentralized aerial manipulator model, where various disturbances with different mathematical properties are well explored and tackled according to their positions and effects acting on the system. The stability of the proposed control framework is ensured by using Lyapunov-like analysis. Experimental results are presented to illustrate the performance of the proposed control framework.Note to Practitioners—One of the key challenges that hinder the aerial manipulator potential applications is its stability and accuracy due to the existence of various disturbances. Most of disturbance rejection control methods in the literature for aerial manipulator always deal with the various disturbances as the lumped one. Their distinct mathematical properties and different impacts on the system are not well explored, which may result in the performance degradation and limit its practical implementation. In this article, a refine anti-disturbance architecture is proposed, which is able to handle various disturbances in a more systematical way according to their mathematical properties and effects acting on the system. Physical experiments suggest that finely tackling the different disturbances enjoys a better performance in aerial manipulator trajectory tracking control problem and thus can promote the aerial manipulator to be deployed in the tasks demanding on the accuracy. In addition, such control strategy can also be extended to other robotic systems suffering from various disturbances.
Shangke Lyu, Chien Chern Cheah, Xiang Yu 0003
IEEE Trans Autom. Sci. Eng.1
2025 Precise End-Effector Control for an Aerial Manipulator Under Composite Disturbances: Theory and Experiments
abstract
One of inescapable challenges in facilitating the application of aerial manipulators is to achieve the high precision control performance of the end-effector. The manipulator motions beneath the UAV platform constantly contend with composite disturbances, such as floating base, strong inner coupling effects, and model uncertainties. These factors collectively contribute to an inadequate control performance. In this paper, a composite control scheme is presented to tackle this issue. Specifically, a joint velocity planner is proposed to handle the base-floating disturbance in kinematic loop. By virtue of the generated joint reference signal, the base-floating disturbance can be effectively alleviated. The tracking error of the end-effector can be ensured within a small set. Moreover, in a complementary manner, neural network (NN) approximation and nonlinear disturbance observer (NDO) compensation are combined to track the joint references. The NN is adopted to estimate composite dynamic model including inner coupling effects and model uncertainties, while the NDO is designed to handle the remaining uncompensated part. The stability of the closed-loop system including the manipulator kinematics and dynamics is guaranteed using the Lyapunov-like method. Experimental results are reported to manifest the effectiveness of the proposed composite control scheme.Note to Practitioners—This work is driven by the precise end-effector control problem of an aerial manipulator subject to base-floating, strong dynamic coupling, and model uncertainties. Most of existing approaches implicitly address this issue by improving the flight performance of the aerial platform. However, the composite disturbances acted on the manipulator, which would deteriorate the operation accuracy of the end-effector, are not systematically addressed. In this work, a composite control scheme is constructed, which consists of the manipulator joint velocity planner and the dynamic controller. The idea is intuitive. The joint velocity is generated to counteract the fluctuation of the aerial platform. Furthermore, the dynamic controller is developed to accurately track the planned joint velocity in the presence of strong coupling and model uncertainties. The proposed scheme guarantees the stability of the close loop system, making our approach especially promising solution for aerial manipulation under composite disturbances.
Meng Wang 0044, Shangke Lyu, Qianyuan Liu, Kexin Guo 0001, Xiang Yu 0003
IEEE Trans Autom. Sci. Eng.2
2024 GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
abstract
Multi-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in collecting training datasets. In this paper, we propose GeRM (Generalist Robotic Model). We utilize offline reinforcement learning to optimize data utilization strategies to learn from both demonstrations and sub-optimal data, thus surpassing the limitations of human demonstrations. Thereafter, we employ a transformer-based VLA network to process multi-modal inputs and output actions. By introducing the Mixture-of-Experts structure, GeRM allows faster inference speed with higher whole model capacity, and thus resolves the issue of limited RL parameters, enhancing model performance in multi-task learning while controlling computational costs. Through a series of experiments, we demonstrate that GeRM outperforms other methods across all tasks, while also validating its efficiency in both training and inference processes. Additionally, we uncover its potential to acquire emergent skills. Additionally, we contribute the QUARD-Auto dataset, collected automatically to support our training approach and foster advancements in multi-task quadruped robot learning. This work presents a new paradigm for reducing the cost of collecting robot data and driving progress in the multi-task learning community.You can reach our project and video through the link: https://songwxuan.github.io/GeRM/.
Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Can Cui 0008, Shangke Lyu, Yaning Fan
IROS5
2024 Contact Force Estimation of Robot Manipulators With Imperfect Dynamic Model: On Gaussian Process Adaptive Disturbance Kalman Filter
abstract
This paper is concerned with the contact force estimation problem of robot manipulators based on imperfect dynamic models of the manipulator and the contact force. To handle the imperfect dynamic information of the manipulator, a hybrid model, consisting of the nominal model and the residual dynamics, is established for the manipulator, and the Gaussian process regression (GPR) technique is employed to learn the mean and covariance of the residual dynamics. On this basis, a virtual measurement equation is established for contact force estimation and a Gaussian process adaptive disturbance Kalman filter (GPADKF) is developed where the variational Bayes technique is employed to achieve online identification of the noise statistics in the force dynamics. The GPADKF is capable of decoupling the contact force from residual dynamics and system noises, thereby reducing the dependence on accurate dynamic models of the manipulator and the contact force. Simulation and experimental results demonstrate that the proposed scheme outperforms the state-of-art methods.Note to Practitioners—Contact force estimation for robot manipulators can be achieved by fusing the dynamic models of the manipulator and the contact force. When both models are imprecise, the traditional inverse dynamics-based and disturbance Kalman filter-based approaches can no longer provide accurate force estimates. To handle this challenge, a computationally efficient hybrid dynamic model is established for the manipulator, which consists of the nominal model and a residual dynamics compensation term learned from offline data via the GPR. On this basis, an adaptive disturbance Kalman filter is constructed by using the variational Bayes technique to deal with the inaccurate noise covariance matrix in the force dynamic model. Compared with the existing approaches, the force estimate obtained via the proposed scheme is more accurate and reliable, as refined noise covariance matrices (provided by both the GPR and the variational Bayes procedure) have been adopted in the Kalman gain calculation. The proposed GPADKF method is the extension of the composite disturbance filtering (CDF) framework. With the proposed scheme, the dependency on the perfect dynamic models in contact force estimation can be significantly reduced, and this makes our approach especially suitable for contact force estimation problems under unfamiliar and complicated environments.
Yanran Wei, Shangke Lyu, Xiang Yu 0003, Zidong Wang 0001, Lei Guo 0003
IEEE Trans Autom. Sci. Eng.2
2024 A New Observer for Perspective Vision Systems With Partially Uncertain Linear Motion Parameters
abstract
Depth estimation problem for the perspective vision systems has been extensively studied in the literature and the depth estimation can be achieved under the assumption that the exact camera motion parameters are available. However, in practice, it is hard to obtain the camera motion parameters exactly. For instance, the exact camera velocity is always unavailable and instead only the roughly estimated one can be obtained. This introduces the significant difficulties to achieve the depth estimation with partially uncertain camera motion parameters. In this article, we consider the depth estimation problem for the perspective vision system in the case that the camera angular velocities are accurate, while partial camera linear velocities are contaminated by some disturbances. This problem is theoretically formulated and solved for the first time in this article by proposing a new depth observer in the presence of the partially uncertain camera linear velocities. Local exponential convergence of the depth and disturbance estimates is achieved in presence of the partially uncertain camera linear velocities such that the estimation errors of states and disturbances converge to zero for constant disturbances or to a small error bound for bound time-varying disturbances. Simulations and experiments are carried out to verify the performance of the proposed observer.
Shangke Lyu, Zhitao Wang, Jianzhong Qiao, Yukai Zhu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2023 A Composite Control Strategy for Quadruped Robot by Integrating Reinforcement Learning and Model-Based Control
abstract
Locomotion in the wild requires the quadruped robot to have strong capabilities in adaptation and robustness. The deep reinforcement learning (DRL) exhibits the huge potential in environmental adaptability, while its stability issues remain open. On the other hand, the quadruped robot dynamic model contains a lot of useful information that is beneficial to the robust control. The combination of DRL with model-based control may take both strengths and hold promises in better robustness. In this paper, the DRL and the proposed model-based controller are firmly integrated in a novel manner such that the proposed model-based controller is able to rectify the gait commands generated by DRL based on the system dynamic model so as to enhance the robustness of the quadruped robot against the external disturbances. Besides, a potential energy function is introduced to achieve the compliant contact. The stability of the proposed method is ensured in terms of passivity analysis. Several physical experiments are carried out to verify the performance of the proposed method.
Shangke Lyu, Han Zhao 0008
IROS1
2022 Interaction Task Motion Learning for Human-Robot Interaction Control
abstract
Traditional applications of robot manipulators are mainly limited to tracking control tasks, in which the desired objectives are specified as desired positions or trajectories. In such applications, a convenient and easy way of programming robots is the traditional teach-and-playback method. However, advances in sensing and robotic technologies have led to the requirements of more demanding tasks, in which robots may need to interact with a human or follow the human’s instructions in performing a sequence of more complex tasks. In such applications, it is not sufficient to just learn the positions or motion and play it back using a robot controller. In this article, a task learning approach is proposed for human–robot interaction systems, where a set of interaction behaviors is formulated and solved by specifying the task requirements in terms of potential energy. The motion behaviors demonstrated by humans can thus be acquired by the robot by seeking the appropriate task parameters of the dynamic potential energy function. To play back and combine the tasks in a sequential way by using a single controller, a new robot controller is also proposed. Lyapunov-like analysis is adopted to guarantee the stability of the control system, and experimental results are given to validate the performance of the proposed learning and control strategies.
Shangke Lyu, Nithish Muthuchamy Selvaraj, Chien Chern Cheah
IEEE Trans. Hum. Mach. Syst.1
2018 Human-guided Optical Manipulation of Multiple Microscopic Objects
abstract
Existing control systems for optical manipulation of multiple micro-objects do not allow users to interact with the systems during manipulation to deal with unexpected events, while ensuring stability of the overall control systems. In this paper, we propose a robotic control technique for human-guided optical manipulation of multiple micro-objects using robotic tweezers. Humans, when necessary, are able to interact or intervene with an automated optical manipulation system in a stable manner, and thus guiding a group of micro-objects to be manipulated towards a desired region while ensuring collision avoidance during manipulation. Both the ability of humans, in term of decision making, when needed, and the advantages of an automated optical manipulation system which provides precise and productive manipulation of micro-objects, are consolidated into one single technique. A theoretical foundation is developed and investigated to achieve the control objective. Experimental results are presented to illustrate the effectiveness of the proposed control technique.
Quang Minh Ta, Shangke Lyu, Chien Chern Cheah
ICRA2