EDBT 2026 Demo / reviewers in the wild / expert
Yongpeng Liu
dblp:96/7237
· DBLP profile ↗
10ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-4544-4217ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Deep Learning Molecular Dynamics to 500M Atoms on 4096-Node ARMv8 ClustersabstractMolecular dynamics (MD) simulations are essential tools for investigating large-scale molecular systems, yet achieving high performance and scalability on CPU-based architectures remains challenging. In this study, we present a highly optimized framework based on DeepMD-kit for conducting 500 millionatom MD simulations on an ARMv8 SVE high-performance computing (HPC) system. Key optimizations include leveraging OpenMP for multi-threaded acceleration of DeepMD-kit and utilizing the ARMv8 SVE instruction set to optimize doubleprecision matrix multiplication in PyTorch. These enhancements enable single ARMv8 SVE 64-core processors to achieve 1.3x the training performance of NVIDIA V100 GPU, and two ARMv8 SVE 64-core processors to achieve 1.05x the inference performance of NVIDIA V100 GPU. Leveraging this optimized framework, we achieve large-scale MD simulations across 4,096 computing nodes. Qi Du, Feng Wang 0050, Chengkun Wu, Han Wang 0006, Yongpeng Liu, Zhaoyin Zhou, Kenli Li 0001 |
CLUSTER | 5 |
| 2025 | Efficient TEE-Based DNN Inference on Edge Devices: A PyTorch-Compatible DesignabstractThe uncontrollable nature of deployment environments and supply chains for edge devices makes the security of deployed deep neural network (DNN) models a key concern. Leveraging the hardware isolation provided by Trusted Execution Environments (TEEs) to protect sensitive layers of the model is considered a practical solution. However, current studies face the following challenges: (1) The dynamic memory management of mainstream deep learning frameworks is incompatible with the static memory allocation mechanism of TEEs, hindering the integration of an efficient intelligent computing ecosystem throughout the model lifecycle. (2) TEEs lack native support for parallel computing, leading to a performance gap between the TEE and the Rich Execution Environment (REE), which increases overall inference latency. To address these issues, this paper proposes a TEE-based model inference scheme integrated with PyTorch for TrustZone-enabled edge devices. The solution adopts a pre-allocated operator management strategy to eliminate the incompatibility between dynamic graph features and the static memory allocation in the TEE. A semantic-segmentationbased parallel optimization is also introduced to reduce the inference latency of sensitive layers within the TEE. We implement a prototype system on Phytium and evaluate it using four well-known DNN models. Experimental results show that, compared to existing TEE-based DNN inference solutions, our design reduces the lines of code (LoC) for model construction scripts by an average of of 82.2%, lowers secure memory overhead by up to 58.5%, and improves inference performance by an average of$5.38 \times$via octa-core parallel optimization. Yuanhang Yu, Yongpeng Liu, Yipin Sun, Zihao Guan, Wei Wang 0250 |
HPCC | 3 |
| 2024 | Tipcc: TEE-based Integrity Protection of Consortium Blockchain ContractsabstractSecure execution of smart contracts on consortium blockchains is essential. Existing solutions commonly use trusted execution environments (TEEs) to ensure isolated and confidential contract execution, enhancing security. However, these methods have limitations, such as resource constraints in TEEs, difficulties in supporting various smart contract programming languages, and potential expansion of the TEE’s attack surface. The use of TEE-based integrity measurement for smart contracts offers a viable solution to this problem. However, the challenges of analyzing the scope of integrity attacks must be addressed, determining the object and timing of integrity measurements, and ensuring trust transfer within the contract invocation chain. We study the integrity of consortium blockchain contracts using TEE and establish a smart contract integrity model for the endorsement policy. Furthermore, we propose Tipcc, an integrity measurement framework that combines static system components and dynamic user contracts using TEE. Tipcc ensures integrity measurement and verification throughout the contract lifecycle, providing trusted transmission along the invocation chain. Prototype system validation and simulation experiments using Hyperledger Fabric revealed that this approach improves transaction security while maintaining availability, exhibiting a performance overhead of approximately 7%. Zihao Guan, Wei Wang 0250, Yongpeng Liu, Liaoliao Feng |
ISPA | 6 |
| 2018 | An Effective Deep Learning Based Scheme for Network Intrusion DetectionabstractIntrusion detection systems (IDS) play an important role in the protection of network operations and services. In this paper, we propose an effective network intrusion detection scheme based on deep learning techniques. The proposed scheme employs a denoising autoencoder (DAE) with a weighted loss function for feature selection, which determines a limited number of important features for intrusion detection to reduce feature dimensionality. The selected data is then classified by a compact multilayer perceptron (MLP) for intrusion identification. Extensive experiments are conducted on the UNSW-NB dataset to demonstrate the effectiveness of the proposed scheme. With a small feature selection ratio of 5.9%, the proposed scheme is still able to achieve a superior performance in terms of different evaluation criteria. The strategic selection of a reduced set of features yields satisfactory detection performance with low memory and computing power requirements, making the proposed scheme a promising solution to intrusion detection in high-speed networks. Hongpo Zhang, Chase Qishi Wu, Zongmin Wang, Yuxiao Xu, Yongpeng Liu |
ICPR | 6 |
| 2014 | Iaso: an autonomous fault-tolerant management system for supercomputers
Kai Lu 0001, Gen Li 0002, Ruibo Wang, Wanqing Chi, Yongpeng Liu, Hong-Wei Tang, Yinghui Gao |
Frontiers Comput. Sci. | 6 |
| 2013 | Low-latency asynchronous duty-cycle MAC protocol for burst traffic in wireless sensor networksabstractMany energy-efficient asynchronous duty-cycle media access control (MAC) protocols for wireless sensor networks (WSNs) have been proposed in recent years. However, for burst traffic, most of them suffer from significant performance degradation due to randomly waking up to communicate with each other. In this paper, we propose a new asynchronous duty-cycle receiver-initiated MAC protocol called HKMAC. In proposed HKMAC, by adaptively adjusting beacon time of the receiver and scheduling the sender's listening time during scheduled period, it can achieve low end-to-end packet delivery latency and high energy efficiency under burst traffic. We have evaluated the performance of HKMAC through detailed ns-2 simulation. The simulation results show that HKMAC can always reduce end-to-end packet delivery latency and energy consumption under various data rates in different topologies compared with RI-MAC - a state-of-the-art MAC protocol in WSNs. Hong-Wei Tang, Caixia Sun, Yongpeng Liu, Baohua Fan |
IWCMC | 3 |
| 2012 | Self-adaptive management of the sleep depths of idle nodes in large scale systems to balance between energy consumption and response timesabstractDue to the time-varying nature of real workload, a large scale computer system has quite a number of idle nodes in most time of operation. They consume energy, but do nothing useful. To save the huge energy waste caused by such active idle nodes, most modern compute nodes provide multiple level dynamic sleep mechanisms to reduce power consumption. However, awaking sleeping nodes takes time, thus affects the response times and performance of the system. A node is deeper in sleep, it consumes less energy, but has longer wakeup latency. This paper proposes a sleep state management model to balance the system's energy consumption and response times. In this model, idle nodes are classified into different groups according to their sleep states. Each group contains nodes of same level of sleep depth and forms a reserve pool of a certain readiness level. In a resource allocation process, nodes in the pool of highest level of readiness are preferentially provided to the application. When the nodes in the pool of the highest readiness level are not sufficient, the nodes in the pool(s) of next level(s) of readiness are allocated. After each allocation and reclaim of nodes, the numbers of nodes in each level of pools are adjusted by changing the sleep depth of the nodes up and down. Thus, the reserve pools can be maintained at all times. Obviously, a key factor that affects the effectiveness of the idle node management is the sizes of the reserve pools. This paper proposes and investigates a self-adaptive approach to this problem so that the sizes of reserve pools are dynamically adjusted according to the applications. Our experiments demonstrated that, by applying our self-adaptive management, the power consumption of idle nodes can be reduced by 84.12% with the cost of slowdown rate being only 8.85%. Yongpeng Liu, Hong Zhu 0002, Kai Lu 0001 |
CloudCom | 1 |
| 2011 | Parallel Compression Checkpointing for Socket-Level Heterogeneous SystemsabstractCheck pointing is an effective fault tolerant technique to improve the reliability of large scale parallel computing systems. However, check pointing causes a large number of computation nodes to store a huge amount of data into file system simultaneously. It does not only require a huge storage space to store system state, but also brings a tremendous pressure on the communication network and I/O subsystem because a massive demand of accesses are concentrated in a short period of time. Data compression can reduce the size of checkpoint data to be saved in the file system and to go through the communication network. However, compression induces a huge time overhead especially in large scale parallel systems, which is the main technical barrier of its practical usability. In this paper, we propose a parallel compression check pointing technique to reduce the time overhead in socket-level heterogeneous architectures. It integrates a number of parallel processing techniques, including transmitting checkpoint data between CPU, GPU and file system in double buffered pipelines, aggregating file write operations, SIMD parallel compression algorithm running on GPU, etc. The paper also reports an implementation of the technique on the Tianhe-1 supercomputer system and the evaluation experiments with the system. The experiment data show that the technique is efficient and practically usable. Yongpeng Liu, Hong Zhu 0002, Yongyan Liu, Baohua Fan |
HPCC | 1 |
| 2010 | A survey of the research on power management techniques for high-performance systemsabstractAbstract This paper surveys the research on power management techniques for high‐performance systems. These include both commercial high‐performance clusters and scientific high‐performance computing (HPC) systems. Power consumption has rapidly risen to an intolerable scale. This results in both high operating costs and high failure rates so it is now a major cause for concern. It has imposed new challenges to the development of high‐performance systems. In this paper, we first review the basic mechanisms that underlie power management techniques. Then we survey two fundamental techniques for power management: metrics and profiling. After that, we review the research for the two major types of high‐performance systems: commercial clusters and supercomputers. Based on this, we discuss the new opportunities and problems presented by the recent adoption of virtualization techniques, and again we present the most recent research on this. Finally, we summarize and discuss the future research directions. Copyright © 2010 John Wiley & Sons, Ltd. Yongpeng Liu, Hong Zhu 0002 |
Softw. Pract. Exp. | 1 |
| 2009 | HPVZ: A High Performance Virtual Computing Environment for Super Computers
Kai Lu 0001, Wanqing Chi, Yongpeng Liu, Hong-Wei Tang |
APPT | 3 |