EDBT 2026 Demo / reviewers in the wild / expert
Jiyuan Shi
dblp:148/7197
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Forward KL Regularized Preference Optimization for Aligning Diffusion PoliciesabstractDiffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous methods conduct return-conditioned policy generation or Reinforcement Learning (RL)-based policy optimization, while they both rely on pre-defined reward functions. In this work, we propose a novel framework, Forward KL regularized Preference optimization for aligning Diffusion policies, to align the diffusion policy with preferences directly. We first train a diffusion policy from the offline dataset without considering the preference, and then align the policy to the preference data via direct preference optimization. During the alignment phase, we formulate direct preference learning in a diffusion policy, where the forward KL regularization is employed in preference optimization to avoid generating out-of-distribution actions. We conduct extensive experiments for MetaWorld manipulation and D4RL tasks. The results show our method exhibits superior alignment with preferences and outperforms previous state-of-the-art algorithms. Zhao Shan, Chenyou Fan, Jiyuan Shi, Chenjia Bai |
AAAI | 4 |
| 2025 | Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement LearningabstractHumans possess delicate dynamic balance mechanisms that enable them to maintain stability across diverse terrains and under extreme conditions. However, despite significant advances recently, existing locomotion algorithms for humanoid robots are still struggle to traverse extreme environments, especially in cases that lack external perception (e.g., vision or LiDAR). This is because current methods often rely on gait-based or perception-condition rewards, lacking effective mechanisms to handle unobservable obstacles and sudden balance loss. To address this challenge, we propose a novel whole-body locomotion algorithm based on dynamic balance and Reinforcement Learning (RL) that enables humanoid robots to traverse extreme terrains, particularly narrow pathways and unexpected obstacles, using only proprioception. Specifically, we introduce a dynamic balance mechanism by leveraging a novel Zero Moment Point (ZMP)-driven reward and task-driven rewards in a whole-body actor-critic framework, aiming to achieve coordinated actions of the upper and lower limbs for robust locomotion. Experiments conducted on a full-sized Unitree H1-2 robot verify the ability of our method to maintain balance on extremely narrow terrains and under external disturbances, demonstrating its effectiveness in enhancing the robot's adaptability to complex environments. The videos are given at https://whole-body-loco.github.io. Weiji Xie, Chenjia Bai, Jiyuan Shi, Yunfei Ge, Weinan Zhang 0001, Xuelong Li 0001 |
IROS | 3 |
| 2025 | Adversarial Locomotion and Motion Imitation for Humanoid Policy LearningabstractHumans exhibit diverse and expressive whole-body movements. However, attaining human-like whole-body coordination in humanoid robots remains challenging, as conventional approaches that mimic whole-body motions often neglect the distinct roles of upper and lower body. This oversight leads to computationally intensive policy learning and frequently causes robot instability and falls during real-world execution. To address these issues, we propose Adversarial Locomotion and Motion Imitation (ALMI), a novel framework that enables adversarial policy learning between upper and lower body. Specifically, the lower body aims to provide robust locomotion capabilities to follow velocity commands while the upper body tracks various motions. Conversely, the upper-body policy ensures effective motion tracking when the robot executes velocity-based movements. Through iterative updates, these policies achieve coordinated whole-body control, which can be extended to loco-manipulation tasks with teleoperation systems. Extensive experiments demonstrate that our method achieves robust locomotion and precise motion tracking in both simulation and on the full-size Unitree H1-2 robot. Additionally, we release a large-scale whole-body motion control dataset featuring high-quality episodic trajectories from MuJoCo simulations. The project page is https://almi-humanoid.github.io. Jiyuan Shi, Ouyang Lu, Sören Schwertfeger, Chi Zhang 0012, Chenjia Bai, Xuelong Li 0001 |
NeurIPS | 1 |
| 2025 | KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic SkillsabstractHumanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of tracking smooth, low-speed human motions, even with delicate reward and curriculum design. This paper presents a physics-based humanoid control framework, aiming to master highly-dynamic human behaviors such as Kungfu and dancing through multi-steps motion processing and adaptive motion tracking. For motion processing, we design a pipeline to extract, filter out, correct, and retarget motions, while ensuring compliance with physical constraints to the maximum extent. For motion imitation, we formulate a bi-level optimization problem to dynamically adjust the tracking accuracy tolerance based on the current tracking error, creating an adaptive curriculum mechanism. We further construct an asymmetric actor-critic framework for policy training. In experiments, we train whole-body control policies to imitate a set of highly dynamic motions. Our method achieves significantly lower tracking errors than existing approaches and is successfully deployed on the Unitree G1 robot, demonstrating stable and expressive behaviors. The project page is https://kungfubot.github.io. Weiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li 0014, Jiyuan Shi, Weinan Zhang 0001, Chenjia Bai, Xuelong Li 0001 |
NeurIPS | 6 |
| 2024 | Hardware Assist for Linux IPC on an FPGA PlatformabstractSpecialized hardware units often accelerate compute-intensive or memory-heavy functions. In previous publications, we proposed concepts to assist Linux with a hardware unit for managing waiting threads to improve blocking inter-process communication (IPC) mechanisms. This paper assesses the effectiveness of this hardware support on a Zynq platform. Although main memory accesses by our hardware unit are time-consuming, a consumer-producer application achieved an up to 220% increased message rate. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 4 |
| 2024 | HASIIL: Hardware-Assisted Scheduling to Improve IPC Latency in LinuxabstractInter-processes communication (IPC) is essential for multi-threaded applications to achieve efficient execution. Synchronization through IPC can become a bottleneck for these applications. The effectiveness of IPC is determined by both its latency and CPU utilization needed for the associated functions. Our research has revealed that for blocking IPC mechanisms, the thread scheduling functions within the Linux operating system significantly contribute to the notification latency. To address this issue, we propose a novel concept called HASIIL, which combines offloading IPC functionality with hardware-assisted scheduling to enhance IPC latency. Through this approach, we can improve the latency of blocking IPC mechanisms by up to 36% in Linux, while also improving CPU utilization by 40%. Tim Twardzik, Lars Nolte, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
CF | 4 |
| 2024 | Robust Quadrupedal Locomotion via Risk-Averse Policy LearningabstractThe robustness of legged locomotion is crucial for quadrupedal robots in challenging terrains. Recently, Reinforcement Learning (RL) has shown promising results in legged locomotion and various methods try to integrate privileged distillation, scene modeling, and external sensors to improve the generalization and robustness of locomotion policies. However, these methods are hard to handle uncertain scenarios such as abrupt terrain changes or unexpected external forces. In this paper, we consider a novel risk-sensitive perspective to enhance the robustness of legged locomotion. Specifically, we employ a distributional value function learned by quantile regression to model the aleatoric uncertainty of environments, and perform risk-averse policy learning by optimizing the worst-case scenarios via a risk distortion measure. Extensive experiments in both simulation environments and a real Aliengo robot demonstrate that our method is efficient in handling various external disturbances, and the resulting policy exhibits improved robustness in harsh and uncertain situations in legged locomotion. Jiyuan Shi, Chenjia Bai, Haoran He, Lei Han 0001, Dong Wang 0008, Bin Zhao 0001, Mingguo Zhao, Xiu Li 0001, Xuelong Li 0001 |
ICRA | 1 |
| 2024 | HW-FUTEX: Hardware-Assisted Futex SyscallabstractEfficient thread synchronization primitives are crucial in modern computer systems for the performant execution of interdependent code segments. In Linux, the futex() syscall is used to construct blocking synchronization primitives such as mutexes or conditional variables. When using futex, the uncontended case is efficiently handled entirely in user space. In the event of contention, the kernel is called to put the waiting thread to sleep until the state of the primitive changes to uncontended. The kernel must be notified of this change by a futex() syscall to wake-up the sleeping thread. This syscall must be issued by the thread that changes the primitive, which is a significant burden on this thread. To remove this burden, we introduce HW-FUTEX to offload the futex wake functionality to a hardware unit (HW Unit) that asynchronously initiates wake-ups of the sleeping threads. This reduces the time required to issue the futex wake functionality by at least 90% to 350 cycles, with no additional overhead in the uncontended case. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Thomas Wild, Andreas Herkersdorf |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | HAWEN: Hardware Accelerator for Thread Wake-Ups in Linux Event NotificationabstractThe performance of multi-threaded applications relies on efficient inter-process communication. One common practice is putting a thread asleep while waiting for a certain condition. Exemplary Linux kernel mechanisms that use this practice include futex, sockets, epoll, eventfd and pipe. Once the condition is met, i.e., the associated event has occurred, the waiting thread is notified. Optimizations for event notification mechanisms in Linux mostly target the thread which receives events. Contrarily, we identified high potential in relieving the event-generating thread and propose HAWEN, a hardware accelerator for thread wake-up support. HAWEN has been integrated into Linux event notification in a minimally intrusive manner. Gem5-based multi-core architecture simulations revealed up to 80% faster thread wake-up times and a 53% shorter event-generating syscall. Lars Nolte, Tim Twardzik, Camille Jalier, Jiyuan Shi, Clara Kowalsky, Thomas Wild, Andreas Herkersdorf |
DAC | 5 |
| 2020 | Accelerating Skycube Computation with Partial and Parallel Processing for Service SelectionabstractRecently researchers use skyline techniques to optimize service selection procedure, where they can filter those low-quality web services from the large amount of candidates and return a much smaller high-quality service set. The skycube concept is adopted for quickly responding to the skyline queries with different combinations of Quality of Web Service (QoWS) parameters. As the skycube computation is quite time-consuming, it is a compelling challenge to accelerate this procedure. However, the current solutions usually have a number of redundant computations which will significantly affect the efficiency. To address such drawbacks, after an in-depth analysis of skycube computation procedure, we introduce a partial skycube, which only consists of the skylines with frequently used combinations of QoWS. Then the computational relationships between the skyline on one subspace and its parent-space are studied. Based on the relationships, we develop ParCube algorithm to speedup partial skycube computation by reusing the intermediate comparison results. Meanwhile, at the execution phase, ParCube can be further optimized with parallel execution mode and optimized scheduling strategy. Finally, we evaluate the efficiency and scalability of ParCube on both single machine and cluster environment. The results show that ParCube can efficiently compute partial skycube and scale well in cluster environment. Fang Dong 0001, Junzhou Luo, Jiahui Jin 0001, Jiyuan Shi, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2016 | Resource provisioning optimization for service hosting on cloud platformabstractWith the popularity of cloud computing technology, service hosting is used as a typical model to deploy different kinds of services on cloud platform. In recent years, how to effectively provide resources for service hosting has attracted more and more attention. However, most of the existing works only focused on how to effectively provide virtual machines for service hosting. They ignored how to efficiently place these virtual machines into physical servers, when considering multidimensional resource requirements. This may result in unreasonable virtual machine placement in servers, thereby causing the underutilization of resource. To address this problem, we propose a novel resource provisioning method including virtual machine provisioning for hosting service and virtual machine placement in servers. The proposed method decides how many virtual machines should be provided for each service by utilizing queuing theory. Then based on the virtual machines to be provided, the proposed method models the virtual machine placement problem as a variant of cutting stock problem, and decides how many servers should be provided by solving this problem. The proposed method is evaluated by simulations. Experimental results show the proposed method achieves a better performance than these baseline methods. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
CSCWD | 1 |
| 2015 | Two-Phase Online Virtual Machine Placement in Heterogeneous Cloud Data CenterabstractWith the rapid development and popularity of cloud computing technology, more and more Collaborative Virtual Environment (CVE) systems are migrated to cloud computing environment to improve the effectiveness of resource usage. Virtual Machine (VM) placement in cloud data center is a key issue of providing high-efficient cloud platform for CVE system. However, most existing VM placement algorithms ignore the following characteristics of actual cloud environment: (1) VMs deployment requests arrive and leave dynamically, (2) Cloud data center usually consists of many heterogeneous Physical Machines (PMs). Ignoring these two characteristics result in an inefficient and unbalanced use of multiple resources of PMs. Thus using these algorithms directly will lead to a poor resource utilization. In this article, we propose a two-phase online VM placement algorithm, which helps the cloud data center to minimize different resource usages and aims at a more efficient use of multiple resources. Our algorithm selects the most suitable PM type for VM based on Cosine Similarity, and adaptively maps VMs to PMs by using an approximation algorithm. The proposed algorithm is evaluated by simulations. Experimental results show our proposed algorithm ensures a more efficient use of multiple resources over the existing approaches. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Junzhou Luo, Ding Ding 0002 |
SMC | 1 |
| 2014 | A budget and deadline aware scientific workflow resource provisioning and scheduling mechanism for cloudabstractCurrently in large-scale scientific experiments, scientists often submit scientific workflow jobs at different time. From the view of system, the entire workload is a stream of jobs submitted at an unpredictable time and different job has different priority and deadline. Moreover the cost of performing these jobs cannot exceed a certain budget constraint. Therefore how to perform scientific workflow applications efficiently in cloud has become the urgent problem. However most of existing work didn't consider unpredictable submission time of jobs, as well as budget and deadline constrains. In this paper, we design an elastic resource provisioning and task scheduling mechanism to perform scientific workflows in cloud. Our goal is to complete as many high-priority workflows as possible under budget and deadline constrains. This mechanism consists of three phases: workflow preprocessing, elastic resource provisioning and task scheduling. We perform evaluation with real AMS experiment scientific computing data under different budget constraints. We also consider inaccurate task execution time, VM provisioning delays and task failures in evaluation. The results show that our mechanism achieves a better performance than these reference mechanisms. In addition, the inaccurate task execution time, VM provisioning delays, and task failures do not bring significant impact to mechanism's performance. Jiyuan Shi, Junzhou Luo, Fang Dong 0001, Jinghui Zhang 0001 |
CSCWD | 1 |