VLDB 2026 Research / reviewers in the wild / expert
Hongyi Lu
dblp:64/6723
· DBLP profile ↗
16ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 2 since 2021Security and privacy · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Secure BPF Kernel Extension With Hardware-Enhanced Memory IsolationabstractThe Linux kernel extensively uses the Berkeley Packet Filter (BPF) to allow user-written BPF applications to execute in the kernel space. The BPF employs a verifier to check the security of user-supplied BPF code statically. Recent attacks show that BPF programs can evade security checks and gain unauthorized access to kernel memory, indicating that the verification process is not flawless. In this paper, we present MOAT, a novel hardware-assisted, cross-platform isolation framework designed to protect the kernel from malicious BPF programs. MOAT introduces a two-layer memory isolation scheme that leverages hardware features such as Intel MPK and Arm Stage-2 translation to enforce isolation. Our design overcomes several key challenges, including the limited scalability of available hardware isolation mechanisms and the risk of helper function abuse. We implement MOAT for Intel x86 and Arm on Linux (ver. 6.1.38), and our evaluation shows that MOAT delivers low-cost isolation of BPF programs under mainstream use cases, such as isolating a BPF packet filter with only 3% throughput loss. Lijian Huang, Hongyi Lu, Shuai Wang 0011, Fengwei Zhang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Hardware-assisted Memory IsolationabstractModern computing systems increasingly rely on hardware-assisted memory isolation to secure critical data and execution contexts without the overhead of purely software-based mechanisms. While features like Intel MPK, Arm POE, and RISC-V PMP offer promising support, they often suffer from a limited number of available isolation domains and primarily focus on CPU memory, leaving interactions with peripheral devices unprotected. Hongyi Lu |
CCS | 1 |
| 2025 | MOLE: Breaking GPU TEE with GPU-Embedded MCUabstractGraphics Processing Units (GPUs) are extensively used for applications such as machine learning, scientific computing, and graphics rendering. To protect sensitive data processed by GPUs, Trusted Execution Environments (TEEs) for GPUs have been proposed. GPU TEEs, built with hardware-based isolation primitives, can defend against high-privilege attackers like OS kernels. However, in this paper, we present MOLE, a novel attack that compromises the security of GPU TEEs on Arm Mali GPUs by exploiting the GPU-embedded Microcontroller Unit (MCU). By injecting malicious firmware into the MCU, an attacker can bypass GPU TEEs' security guarantees. We evaluated MOLE with state-of-the-art GPU TEE proposals under multiple real-world attack scenarios, such as in-GPU AES encryption and object detection tasks. Our evaluation shows that MOLE can successfully extract sensitive data or manipulate the computation results of GPU TEEs. We responsibly disclosed our findings to the authors of the affected GPU TEE proposals and received acknowledgments from all of them. Moreover, our findings prompted Arm to enhance the security of its GPU firmware supply chains. Hongyi Lu, Yunjie Deng 0001, J. Sukarno Mertoguno, Shuai Wang 0011, Fengwei Zhang |
CCS | 1 |
| 2025 | Air-Ground Cooperative Multitarget Hierarchical Tracking Method Based on Aerial Fisheye ViewabstractThe multitarget tracking of multirobot systems has a wide range of applications, including urban security, search, and rescue. However, these tasks often take place in obstacle-rich environments, posing significant challenges, such as an unknown global map, unclear target locations, and inefficient resource allocation. In this article, we propose an air–ground cooperative hierarchical method for multitarget tracking. The uncrewed aerial vehicle (UAV) is equipped with a fisheye camera to provide a wide field of view (FOV) for multitarget detection and an initial global map generated from segmented aerial images. Multiple targets are assigned to uncrewed ground vehicles (UGVs) using the task allocation method, which is based on the PIO and auction mechanism. UGV track targets through motion planning to avoid obstacles and keep their assigned targets in the FOV. Experimental validation confirmed the effectiveness of our method, demonstrating that UGVs can autonomously and reliably track targets in obstacle-rich environments with UAV assistance. Yangjie Cui, Hongyi Lu, Xin Dong 0020, Jinwu Xiang, Daochun Li, Zhan Tu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | DAWN: Matrix Operation-Optimized Algorithm for Shortest Paths Problem on Unweighted GraphsabstractThe shortest paths problem is a fundamental challenge in graph theory, with a broad range of potential applications. The algorithms based on matrix multiplication exhibits excellent parallelism and scalability, but is constrained by high memory consumption and algorithmic complexity. Traditional shortest paths algorithms are limited by priority queues, such as BFS and Dijkstra algorithm, making the improvement of their parallelism a focal issue. We propose a matrix operation-optimized algorithm, which offers improved parallelism, reduced time complexity, and lower memory consumption. The novel algorithm requires O(Ewcc(i)) and O(Swcc · Ewcc) times for single-source and all-pairs shortest paths problems, respectively, where Swcc and Ewcc denote the number of nodes and edges included in the largest weakly connected component in graph. To evaluate the effectiveness of the novel algorithm, we tested it using graphs from SuiteSparse Matrix Collection and Gunrock benchmark dataset. Our algorithm outperformed the BFS implementations from Gunrock and GAP (the previous state-of-the-art solution), achieving an average speedup of 3.769 × and 9.410 ×, respectively. Yelai Feng, Huaixi Wang, Yining Zhu, Xiandong Liu, Hongyi Lu |
ICS | 5 |
| 2024 | MOAT: Towards Safe BPF Kernel Extension
Hongyi Lu, Shuai Wang 0011, Yechang Wu, Wanning He, Fengwei Zhang |
USENIX Security Symposium | 1 |
| 2022 | Raven: a novel kernel debugging tool on RISC-VabstractDebugging is an essential part of kernel development. However, debugging features are not available on RISC-V without the use of external hardware. In this paper, we leverage a security feature called Physical Memory Protection (PMP) as a debugging primitive to address this issue. Based on this debugging primitive, we design Raven, a novel kernel debugging tool with the standard functionalities (breakpoints, watchpoints, stepping, introspection). A prototype of Raven is implemented on a SiFive Unmatched development board. Our experiments show that Raven imposes a moderate but acceptable overhead to the kernel. Moreover, a real-world debugging scenario is set up to test its effectiveness. Hongyi Lu, Fengwei Zhang |
DAC | 1 |
| 2018 | Adaptive VC Partitioning for NoCs in GPGPUsabstractThe design of efficient Networks-on-Chip (NoCs) is essential for GPGPUs. The asymmetry of GPGPU traffic has a significant effect on the overall system performance. An existing VC partitioning design statically assigns more VCs to the heavier reply traffic. Yet, its static partitioning cannot adapt to the dynamic variation of NoC traffic. Thus, we propose an adaptive VC partitioning (A-VCP) mechanism , which dynamically chooses the optimal VC partitioning by sampling the traffic status. Compared with the static configuration, A-VCP averagely improves the system performance by 10.5%, and reduces the energy-delay product by 9.1%. Sheng Ma, Hongyi Lu, Libo Huang 0002, Li Shen 0007, Yang Guo 0003, Zhiying Wang 0003, Wenliang Xue |
ISCAS | 2 |
| 2015 | Adaptive remaining hop count flow control: Consider the interaction between packetsabstractThe interaction between packets affects performance and global fairness of Network-on-Chip. Preferentially transferring packets with small remaining hop counts (PPSR) can reduce the flying packet amount to improve the performance. Yet, the global fairness is negatively affected. In contrast, preferentially transferring packets with large remaining hop counts (PPLR) can achieve better global fairness with a poorer performance. In this paper, we propose adaptive remaining hop count flow control, which dynamically switches between PPSR and PPLR. In this way, we can achieve higher performance and better global fairness. Peng Wang 0036, Sheng Ma, Hongyi Lu, Zhiying Wang 0003, Chen Li 0015 |
ASP-DAC | 3 |
| 2011 | A specialized low-cost vectorized loop buffer for embedded processorsabstractCurrent loop buffer has been mainly explored as an effective architectural technique for low-power execution in embedded processor. Another avenue, however, for exploiting loop buffer is to obtain its performance benefit. In this paper, we propose an application specific loop buffer organization for vectorized processing kernels, to achieve low-power and high-performance goals. The vectorized loop buffer (VLB) is simplified with single loop support for SIMD devices. Since significant data rearrangement overhead is required in order to use the SIMD capabilities, the VLB is specialized for zero-overhead implicit data permutation. We extend several instructions to the baseline ISA for programming and integrate it into an embedded processor for evaluation. Our results show that VLB improves the performance and power measures significantly compared to conventional SIMD devices. Libo Huang 0002, Zhiying Wang 0003, Li Shen 0007, Hongyi Lu, Nong Xiao 0001, Cong Liu 0009 |
DATE | 4 |
| 2010 | DSS: Applying asynchronous techniques to architectures exploiting ILP at compile timeabstractEmbedded application environments require both high performance and low power. Architectures exploiting instruction-level parallelism (ILP) at compile time, such as very long instruction word (VLIW) and transport triggered architecture (TTA), may satisfy the requirements. They can be further enhanced by using asynchronous circuits to significantly reduce power consumption. As such, we are interested in asynchronous processors with architectures exploiting ILP at compile time. However, most of the current asynchronous processors are based on RISC-like architectures. When designing asynchronous VLIW or TTA processors, the distribution of control introduces some serious problems, and errors may occur because of the variable latencies of operations. This paper investigates the asynchronous processor with architecture exploiting ILP at compile time. In order to overcome these problems, we propose a data source selecting (DSS) scheme to guarantee instructions run correctly on asynchronous VLIW and TTA processors. Concretely, an asynchronous pipelined processor based on TTA is designed. The micro-architecture of the proposed asynchronous TTA processor is presented and an asynchronous processor named Tengyue is implemented using 180nm technology. The experimental results, for a range of benchmarks and working modes, show that the implemented asynchronous TTA processor with DSS scheme support runs correctly and power dissipation is reduced to about 43% to 65% of the equivalent synchronous processor. Zhiying Wang 0003, Hongguang Ren, Wei Chen 0009, Hongyi Lu |
ICCD | 7 |
| 2009 | A Light-weight Code Cache Design for Dynamic Binary TranslationabstractInterpretation and basic block translation (BBT) are two typical strategies for cold code emulation in a dynamic binary translation (DBT) system. More and more DBT systems employ BBT as the generated native code runs more efficient than the interpretation routines. We observe that BBT's high efficiency is based on those special hardware assists. With certain simple hardware techniques, interpretation could outperform BBT. In our pervious work, we proposed a hardware interpreted code cache (Pcache) mechanism to speedup interpretation by saving the decoded instruction information during interpretation. This light-weight code cache design could be extended to assist the hotspots translation, thus further reduce the DBT systems' overhead. We add the translation entry into the Pcache design thus saving most decoding operations during translation. We use eight SPEC 2000 integer benchmarks on our DBT simulator. Results show that the modified Pcache design causes a speedup of 1.94 according to the referenced DBT with basic interpretation and the interpretation based DBT system assisted by the modified Pcache performs more efficiently than the DBT system which employs BBT for the cold code. Wei Chen 0009, Li Shen 0007, Hongyi Lu, Zhiying Wang 0003, Nong Xiao 0001 |
ICPADS | 3 |
| 2009 | Using Pcache to Speedup Interpretation in Dynamic Binary TranslationabstractAbstract— Dynamic binary translation (DBT) converts codes written for a source instruction set architecture (ISA) into optimized code for a target ISA. DBT has emerged as an important tool with real world applications. Interpretation is always adopted to handle the non-hotspot code in a two-stage DBT system. An important consideration in such DBT systems is the interpretation overhead. We investigate that repeated redecoding operations are the bottleneck of interpretation overhead. We propose interpreted code cache (Pcache), a hardware assist to save the information of the decoded instruction for reuse. We analyze and model Pcache performance via simulation on a DBT system simulator. Results from SPEC2000 integer benchmarks show that Pcache could significantly reduce redecoding operations and the overhead of interpretation in a DBT system. The speedup of interpretation is up to 17.12 on average with assist of Pcache. We also analyze the extra overhead caused by Pcache, which is neglectable compared to the performance gains. Wei Chen 0009, Hongyi Lu, Li Shen 0007, Zhiying Wang 0003, Nong Xiao 0001 |
ISPA | 2 |
| 2008 | A dynamically-allocated virtual channel architecture with congestion awareness for on-chip routersabstractIn this paper, the dynamically-allocated virtual channels (VCs) architecture with congestion awareness is introduced. All the buffers are shared among VCs whose structure varies with traffic condition. In low rate, this structure extends VC depth for continual transfers to reduce packet latencies. In high rate, it dispenses many VCs and avoids congestion situations to improve the throughput. We modify the VC controller and VC allocation modules, while designing simple congestion avoidance logic. The experiment shows that the proposed routers outperform conventional ones under different traffic patterns. They provide 8.3% throughput increase and 19.6% latency decrease while saving 27.4% of area and 28.6% of power. Zhiying Wang 0003, Hongyi Lu, Kui Dai |
DAC | 4 |
| 2006 | Designing Power Analysis Resistant and High Performance Block Cipher Coprocessor Using WDDL and Wave-Pipelining
Yuan-man Tong, Zhiying Wang 0003, Kui Dai, Hongyi Lu |
Inscrypt | 4 |
| 2006 | A 6.35Mbps 1024-bit RSA crypto coprocessor in a 0.18um CMOS technologyabstractIn this paper a RSA crypto coprocessor that is fabricated using a 0.18mum CMOS technology is presented. This processor combines a new version of high radix Montgomery multiplication algorithm with a super-pipeline design. With this algorithm, modular exponentiation can be decomposed into a series of primitive operation (PO) matrixes. All the POs are scheduled on the pipeline by employing column-sharing strategy, and inside the PO all the partial results are compressed first by Wallace tree to assure only one carry propagation in the critical path. With these optimizations, a decryption rate of 6.35 Mbps can be achieved for 1024-bit RSA Xue-mi Zhao, Zhiying Wang 0003, Hongyi Lu, Kui Dai |
VLSI-SoC | 3 |