VLDB 2026 Research / reviewers in the wild / expert
Zhengyu He
dblp:79/7449
· DBLP profile ↗
24ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 4 since 2021Security and privacy · 8 · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SKernel: An Elastic and Efficient Secure Container System at Scale with a Split-Kernel ArchitectureabstractSecure containers leverage hardware virtualization to isolate container sandboxes, enabling dedicated guest kernels to mitigate shared kernel attacks prevalent in traditional systems. However, existing approaches struggle with a fundamental trade-off: VM-based solutions (e.g., Kata) prioritize performance but lack elasticity and on-demand usage for volatile and bursty workloads, while lightweight methods (e.g., gVisor) rely on the host kernel for dynamic resource management at the cost of significant performance degradation due to guest-host dependencies. Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Guotao Tan, Anqi Shen, Dawei Shen, Xinyao Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
EuroSys | 13 |
| 2026 | SoK: Analysis of Accelerator TEE Designs
Chenxu Wang 0005, Yujun Liang, Xuanyao Peng, Yuqun Zhang, Fengwei Zhang, Jiannong Cao 0001, Rui Hou 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 12 |
| 2026 | Building Confidential Accelerator Computing Environment for Arm CCA
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 10 |
| 2026 | Complementing Confidential Computing Environment for Applications on Arm CCA
Yiming Zhang 0030, Zhenyu Ning, Fengwei Zhang, Xiapu Luo, Haoyang Huang, Shoumeng Yan, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2025 | ccAI: A Compatible and Confidential System for AI ComputingabstractConfidential xPU computing has emerged as a prominent technique for effectively securing users' AI computing workloads on heterogeneous systems equipped with xPUs.Although the industry adopts this technology in cutting-edge hardware (e.g.NVIDIA H100 GPU) to safeguard high-performance AI computing, most clouds still rely on legacy xPUs and suffer from data leakage problems. Chenxu Wang 0005, Danqing Tang, Changxu Ci, Yankai Xu, Fengwei Zhang, Jiannong Cao 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
MICRO | 11 |
| 2025 | SCRUTINIZER: Towards Secure Forensics on Compromised TrustZone
Yiming Zhang 0030, Fengwei Zhang, Xiapu Luo, Rui Hou 0001, Xuhua Ding, Zhenkai Liang, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 9 |
| 2025 | Fork in the Road: Reflections and Optimizations for Cold Start Latency in Production Serverless Systems
Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Anqi Shen, Dawei Shen, Qi Xing, Shun Song, Tongkai Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
OSDI | 13 |
| 2024 | Verifying Rust Implementation of Page Tables in a Software Enclave HypervisorabstractAs trusted execution environments (TEE) have become the corner stone for secure cloud computing, it is critical that they are reliable and enforce proper isolation, of which a key ingredient is spatial isolation. Many TEEs are implemented in software such as hypervisors for flexibility, and in a memory-safe language, namely Rust to alleviate potential memory bugs. Still, even if memory bugs are absent from the TEE, it may contain semantic errors such as mis-configurations in its memory subsystem which breaks spatial isolation. Zhenyang Dai, Vilhelm Sjöberg, Xupeng Li, Yu Chen 0004, Wenhao Wang 0001, Yuekai Jia, Sean Noble Anderson, Laila Elbeheiry, Shubham Sondhi, Yu Zhang 0313, Zhaozhong Ni, Shoumeng Yan, Ronghui Gu, Zhengyu He |
ASPLOS (2) | 15 |
| 2024 | CAGE: Complementing Arm CCA with GPU Extensions
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Zhengyu He |
NDSS | 8 |
| 2024 | Building a Lightweight Trusted Execution Environment for Arm GPUsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation. However, Arm GPU security has not been explored by the community. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, but there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. There is a need for generalizable and efficient Arm-based GPU security mechanisms. To address these problems, we presentStrongBox, the first GPU TEE for secured general computation on Arm endpoints.StrongBoxprovides an isolated execution environment by ensuring exclusive access to GPU. Our approach is based in part on a dynamic, fine-grained memory protection policy as Arm-based GPUs typically share a unified memory with the CPU. Furthermore,StrongBoxreduces runtime overhead from the redundant security introspection operations. We also design an effective defense mechanism withinsecure worldto protect the confidential GPU computation. Our design leverages the widely-deployed Arm TrustZone and generic Arm features, without hardware modification or architectural changes. We prototypeStrongBoxusing an off-the-shelf Arm Mali GPU and perform an extensive evaluation. Results show thatStrongBoxsuccessfully ensures GPU computation security with a low (4.70%–15.26%) overhead. Chenxu Wang 0005, Yunjie Deng 0001, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native EnvironmentabstractIn cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor. Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001 |
SOSP | 11 |
| 2023 | SHELTER: Extending Arm CCA with Isolation in User Space
Yiming Zhang 0030, Zhenyu Ning, Fengwei Zhang, Xiapu Luo, Haoyang Huang, Shoumeng Yan, Zhengyu He |
USENIX Security Symposium | 8 |
| 2022 | StrongBox: A GPU TEE on Arm EndpointsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation such as image processing and numerical processing applications. However, in spite of these important use cases, Arm GPU security has yet to be scrutinized by the community. By exploiting vulnerabilities in the kernel, attackers can directly access sensitive data used during GPU computing, such as personally-identifiable image data in computer vision tasks. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, while there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. In addition, extant Arm-based GPU defenses are intended for secure machine learning, and lack generality. There is a need for generalizable and efficient Arm-based GPU security mechanisms. Yunjie Deng 0001, Chenxu Wang 0005, Shunchang Yu, Shiqing Liu, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
CCS | 9 |
| 2022 | HyperEnclave: An Open and Cross-platform Trusted Execution Environment
Yuekai Jia, Wenhao Wang 0001, Yu Chen 0004, Zhengde Zhai, Shoumeng Yan, Zhengyu He |
USENIX ATC | 7 |
| 2022 | Unified Enclave Abstraction and Secure Enclave Migration on Heterogeneous Security Architectures
Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Chenggang Qin, Zhengyu He |
J. Comput. Sci. Technol. | 6 |
| 2013 | A queuing model-based approach for the analysis of transactional memory systemsabstractSUMMARY In this paper, we develop an analytical model of the execution efficiency of transactional memory (TM) systems. This model employs queuing theory to analyze the impact of an essential set of TM design parameters including the conflict rate, number of conflict detection/resolution points, and implementation overhead. The model is validated via extensive experiments. To demonstrate the effectiveness of the model, we further study the performance impact of two factors. Our study shows that, for a given TM‐based program, the frequency of performing conflict detection can be carefully chosen to minimize the mean transaction completion time. Our study also demonstrated the importance of reducing implementation overhead. We expect our study to be useful for designing TM systems and applications. Copyright © 2012 John Wiley & Sons, Ltd. Xiao Yu 0001, Zhengyu He |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | Profiling-based Adaptive Contention Management for Software Transactional MemoryabstractIn software transactional memory (STM) systems, the contention management (CM) policy decides what action to take when a conflict occurs. CM is crucial to the performance of STM systems. However, the performance of existing CMs is sensitive to transaction workload and system platforms. A static policy is therefore unsatisfactory. In this paper, we argue that adaptive contention management is necessary and feasible. We further present a profiling-based method that can choose a suitable CM for a given workload and system platform during run-time. We also propose to use logic-time (transactional commit or abort events) to measure the profiling length and compare it with the traditional physical-time-based method. Experimental results demonstrate that our proposed adaptive contention manager (ACM) outperforms static CMs across benchmarks and platforms. In particular, the ACM that uses the number of aborts for the profiling length performs better than others. Zhengyu He, Xiao Yu 0001 |
IPDPS | 1 |
| 2012 | On Adaptive Contention Management Strategies for Software Transactional MemoryabstractSoftware Transaction Memory (STM) is an alternative synchronization method to the traditional lock-based schemes. In an STM system, the contention manager(CM) decides what action to take when a conflict occurs. CM is crucial to the performance of STM systems. However, the performance of existing CMs is sensitive to the transaction workloads and STM configurations. A static policy is therefore unsatisfactory. In this paper, we argue that adaptive contention manager (ACM) is necessary and feasible. We further present an ACM policy that can adaptively choose a suitable CM during run-time. We prove that our adaptation strategy preserves live-lock (starvation) freedom as long as the pool of CMs to adapt from contains at least one live-lock free (starvation free) CM. Experimental results demonstrate that our approach can choose proper CMs and achieves higher average throughput than existing static CM strategies. Xiao Yu 0001, Zhengyu He |
ISPA | 2 |
| 2011 | An Asynchronous Multithreaded Algorithm for the Maximum Network Flow Problem with Nonblocking Global Relabeling HeuristicabstractIn this paper, we present a novel asynchronous multithreaded algorithm for the maximum network flow problem. The algorithm is based on the classical push-relabel algorithm, which is essentially sequential and requires intensive and costly lock usages to parallelize it. The novelty of the algorithm is in the removal of lock and barrier usages, thereby enabling a much more efficient multithreaded implementation. The newly designed push and relabel operations are executed completely asynchronously and each individual process/thread independently decides when to terminate itself. We further propose an asynchronous global relabeling heuristic to speed up the algorithm. We prove that our algorithm finds a maximum flow with O(\vert V\vert^2\Vert E\vert ) operations, where \vert V\vert is the number of vertices and \vert E\vert is the number of edges in the graph. We also prove the correctness of the relabeling heuristic. Extensive experiments show that our algorithm exhibits better scalability and faster execution speed than the lock-based parallel push-relabel algorithm. Zhengyu He |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | Dynamically tuned push-relabel algorithm for the maximum flow problem on CPU-GPU-Hybrid platformsabstractThe maximum flow problem is a fundamental graph theory problem with many important applications. Max-flow algorithms based on the push-relabel method are known to have better complexity bound and faster practical execution speed than others. However, existing push-relabel algorithms are designed for uniprocessors or parallel processors that support locking primitives, thus making it very difficult to apply the push-relabel technique to CUDA-based GPUs. In this paper, we present a first generic parallel push-relabel algorithm for CUDA devices. We model the parallelization efficiency of the algorithm, which reveals that, for a given input graph, the level of parallelism varies during the execution of the algorithm. To maximize the execution efficiency, we develop a dynamically tuned algorithm that utilizes both CPU and GPU by adaptively switching between the two computing units during run time. We show that algorithm finds the maximum flow with O(|V|2|E|) operations (summed over both the CPU and the GPU). Extensive experimental results show that the new algorithm is up to 2 times faster than the push-relabel algorithm by Goldberg et al. Zhengyu He |
IPDPS | 1 |
| 2010 | Modeling the Run-time Behavior of Transactional MemoryabstractIn this paper, we develop a queuing theory based analytical model to evaluate the performance of transactional memory. Based on the statistical characteristics observed on actual experiments, we model each transaction as a client requesting services from the computing system. Continuous Time Markov chain is used to describe the start and completion (commit or abort) of the transactions. We analyze the mean transaction execution time to evaluate the performance of target transactional memory systems. Experimental results based on STAMP benchmarks show that our model can predict the performance of real transactional memory systems with an average error rate of 7.9%. Zhengyu He |
MASCOTS | 1 |
| 2010 | An Analytical Model on the Execution of Transactional MemoryabstractIn this paper, we develop an analytical model of the execution of transactional memory (TM) systems. This model employs queuing theory to analyze the impact of an essential set of TM design parameters including the conflict rate, number of checkpoints, and implementation overhead, etc. The model is validated via extensive experiments. To demonstrate the effectiveness of the model, we further study the performance impact of two factors. Our study shows that, for a given TM-based program, the frequency of performing checkpoint can be carefully chosen to minimize the mean transaction completion time. Our study also demonstrated the importance of reducing implementation overhead. Xiao Yu 0001, Zhengyu He |
SBAC-PAD | 2 |
| 2009 | Impact of early abort mechanisms on lock-based software transactional memoryabstractSoftware transactional memory (STM) is an emerging concurrency control mechanism for shared memory accesses. Early abort is one of the important techniques to improve the execution speed of STMs and has been explored intensively via experimental studies. This paper presents a theoretical analysis characterizing the properties of early abort and its impact on the performance of lock-based STMs. Queuing theory is adopted to model the behaviors of transactional execution. Analytical results are obtained for STMs with and without early abort. The analysis is validated through extensive experiments. Our results reveal that although early abort helps improve the performance of lock-based STMs especially when the contention level is low, the gain is often marginal. We expect our theoretical results to provide useful guidance towards the design and selection of appropriate lock-based STM schemes. Zhengyu He |
HiPC | 1 |
| 2009 | On the Performance of Commit-Time-Locking Based Software Transactional MemoryabstractCompared with lock-based synchronization techniques, software transactional memory (STM) can significantly improve the programmability of multithreaded applications. Existing research results have demonstrated through experiments that current STM designs have slower execution speed than the locks. This paper develops a theoretical explanation for the performance difference. In particular, commit-time-locking (CTL) based STMs are analyzed. A queuing theory based statistical model is developed to quantify the performance of lock-based and STM-based schemes. Analytical results obtained from the model are validated by simulations. Our study shows that (1) lock-based synchronization outperforms CTL-based STMs, and (2) when the contention level becomes low, locks and CTL-based STMs exhibit similar performance. Furthermore, we show that the performance of CTL-based STMs is sensitive to the number of threads, transaction issue rate, and bandwidth of the interconnect. Our results are expected to be useful in the early stages of designing parallel programs, especially on the selection of design schemes for STMs. Zhengyu He |
HPCC | 1 |