Zhichao Hua 0001

dblp:190/4771-1 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-2211-9120ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
abstract
Large Language Models (LLMs) deployed on mobile devices offer benefits like user privacy and reduced network latency, but introduce a significant security risk: the leakage of proprietary models to end users.
Xunjie Wang, Jiacheng Shi 0002, Yang Yu 0002, Zhichao Hua 0001, Jinyu Gu 0001
EuroSys5
2026 EIDS: A Cloud Intrusion Detection System with High Performance and Maintainability
abstract
Intrusion Detection Systems (IDSes) are widely employed to identify potential attacks in guest virtual machines (VMs). Nonetheless, traditional IDSes fall short of the demands of high-performance clouds. First, monitoring VM events increases the tail latency of guest services. Second, the throughput of traditional IDSes cannot meet high-performance cloud requirements, leading to event loss and reduced detection accuracy. Finally, cloud providers typically run complex IDS tools within the VM. Updating IDS functionality requires modifying guest VMs, which hurts maintainability. To overcome these challenges, this article presents EIDS, a cloud IDS framework with high performance and good maintainability. We observe that the main bottleneck is collecting VM status, and the collected status can be divided into fundamental and supplementary status. EIDS then splits the status collection procedure spatially and temporally. First, we provide a status monitor with a separate architecture that isolates the status collection logic in a microVM, thus minimizing the code in guest VMs and improving maintainability. Second, EIDS introduces a two-phase status collection method to handle multiple events in batches, asynchronously, for high IDS throughput. A tiny tracer, implemented with eBPF, operates inside the user VM to collect fundamental status. The complex status collector runs in an isolated microVM. It utilizes Virtual Machine Introspection (VMI) to gather supplementary status, using the fundamental status to bridge the semantic gap. The status collector batches the collection for multiple events to amortize the fixed overhead of microVM switching and improve event tracing throughput. Finally, to minimize tail latency overhead, a fine-grained and workload-aware scheduler executes IDS logic with small time slices during user VM idle periods. We implemented a prototype of EIDS in Linux-KVM and conducted a comprehensive evaluation. We compared EIDS’s performance with Falco, an open-source IDS widely used by Kubernetes and AWS for runtime security monitoring. The results demonstrate that, compared to Falco, EIDS reduces the 99 th -percentile latency overhead by 97% and achieves a 13.8X improvement in IDS event handling throughput.
Xiaokang Hu, Zhichao Hua 0001, Naixuan Guan, Yibin Shen, Yang Yu 0002, Zeyu Mi, Yubin Xia, Jiesheng Wu
ACM Trans. Comput. Syst.2
2025 DPCapsule: A Decentralized Private Computing System With Self-Controlled Data
abstract
Machine learning and data analysis algorithms leverage massive datasets to deliver powerful functionalities.However, these datasets are often distributed across multiple parties and individual users.Decentralized private computing systems enable data consumers to execute algorithms on third-party data in a secure manner, eliminating the need for trust in any single node.Nevertheless, existing systems, primarily derived from blockchain technology, target small-scale workloads.Sharing large-scale datasets presents new challenges.First is data-oriented access control, enabling data providers to enforce flexible access policies throughout the complex utilization of their datasets.Second is high performance, which is essential for data analysis applications involving large-scale datasets and sophisticated computational logic.Third is whole-lifecycle privacy, ensuring comprehensive protection of both data and algorithmic privacy from task initiation through result delivery.To address these challenges, we present DPCapsule, a highperformance decentralized computing system that maintains wholelifecycle privacy.DPCapsule introduces a novel data abstraction called Capsule, which encapsulates dataset and access policies within a trusted execution environment (TEE)-based shell, enabling data providers to control their datasets throughout subsequent computations.A Capsule reborn mechanism is provided for automated access policy updates.Additionally, we design a two-layer execution architecture and consensus protocol to facilitate private computation with scalable performance.Furthermore, a secure execution protocol is designed to guarantee whole-lifecycle privacy for both dataset, algorithms, and metadata.We have developed a prototype of DPCapsule and conducted evaluations across various configurations, with networks scaling up to 32 nodes.Experimental results show that DPCapsule effectively scales to 32 nodes, achieving 116× and 2.2 * 10 7 × latency speedup for database and machine learning applications, respectively, compared with Ethereum.
Yitong Cheng, Yang Yu 0002, Zhichao Hua 0001
Internetware3
2025 DeFS: A Decentralized and High-Performance File System for Consortium Systems
abstract
Consortium decentralized systems, also known as consortium systems, enable consensus and availability among limited untrusted participants.Given the growing necessity for inter-organizational data sharing, consortium systems have gained significant prominence in cross-enterprise collaboration.File systems, which play a fundamental role in data sharing, face new challenges within consortium systems.The consortium system involves characteristics of both decentralized and centralized systems.The first requirement is decentralization.The system's functionality, availability, and security must not depend on any individual node.The second requirement is characteristic-awareness, which necessitates optimal data placement across nodes based on policy constraints, performance requirements, and security considerations.The third requirement is high performance and flexible access control.Neither centralized nor existing decentralized file systems can satisfy all three requirements.This paper presents DeFS, a novel decentralized file system designed for consortium systems.DeFS implements a two-layer architecture that incorporates public nodes into the consortium system, thereby enhancing decentralization.Additionally, we propose a Multi-Ring Distributed Hash Table (MR-DHT) protocol to facilitate characteristic-aware data block distribution.To optimize data routing efficiency, we introduce the Location Cache (L-Cache) mechanism.We have implemented a DeFS prototype and conducted comprehensive performance evaluations across three distinct network configurations, with the largest one having over 1,500 nodes.Results show that DeFS successfully achieves characteristic-aware data placement while delivering 10.32X lower latency compared to IPFS on average.
Yitong Cheng, Shenglong Zhao, Yang Yu 0002, Zhichao Hua 0001
Internetware4
2025 XpuTEE: A High-Performance and Practical Heterogeneous Trusted Execution Environment for GPUs
abstract
AI applications are employed in diverse scenarios, including data centers, personal computers, smart cars, and so on. Their privacy is threatened by the intricate software stacks and the potential malfeasance of system maintainers. The Trusted Execution Environment (TEE) has become popular for safeguarding applications from untrusted system software. However, AI applications are always speeded up with heterogeneous accelerators, e.g., GPU, which requires the TEE to be heterogeneous. A heterogeneous TEE should satisfy three requirements: (1) the joint heterogeneous abstraction that covers the CPU and GPUs and minimizes cooperation overhead among enclaves on them; (2) the high performance for supporting high-speed GPUs and introducing limited performance overhead; and (3) the compatibility with existing CPUs and GPUs so that existing machines can directly benefit from it. To meet the above requirements, this article introduces XpuTEE, a practical and high-performance heterogeneous TEE system. XpuTEE provides a new abstraction called XpuEnclave, comprising the CEnclave to protect CPU-side logic and numerous XEnclaves to guard GPU tasks. XpuEnclave is a joint TEE crossing the CPU and connected GPUs, and it removes all cryptographic operations and extra memory copies for CPU-GPU communication, which allows XpuTEE to achieve high performance. The results demonstrate that XpuTEE has an average performance overhead of 2.48% for common AI applications.
Shulin Fan, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001
ACM Trans. Comput. Syst.2
2023 CPS: A Cooperative Para-virtualized Scheduling Framework for Manycore Machines
abstract
Today's cloud platforms offer large virtual machine (VM) instances with multiple virtual CPUs (vCPU) on manycore machines. These machines typically have a deep memory hierarchy to enhance communication between cores. Although previous researches have primarily focused on addressing the performance scalability issues caused by the double scheduling problem in virtualized environments, they mainly concentrated on solving the preemption problem of synchronization primitives and the traditional NUMA architecture. This paper specifically targets a new aspect of scalability issues caused by the absence of runtime hypervisor-internal states (RHS). We demonstrate two typical RHS problems, namely the invisible pCPU (physical CPU) load and dynamic cache group mapping. These RHS problems result in a collapse in VM performance and low CPU utilization because the guest VM lacks visibility into the latest runtime internal states maintained by the hypervisor, such as pCPU load and vCPU-pCPU mappings. Consequently, the guest VM makes inefficient scheduling decisions.
Yuxuan Liu 0019, Tianqiang Xu, Zeyu Mi, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001
ASPLOS (4)4
2023 SecDFS: A Secure and Decentralized File System
abstract
Consortium networks are usually constructed among enterprises and organizations. A consortium network is a specific type of decentralized network which only has a limited number of untrusted nodes. Consortium systems are designed to achieve consensus in such a network and provide stable and trusted services. Meanwhile, secure data sharing plays a crucial role in the consortium network, introducing new requirements to the underlying file system. Firstly, a consortium file system should be decentralized since any party of the consortium network is untrusted. Secondly, it should enforce data confidentiality and integrity and provide access control. Finally, the system should offer high performance. Existing file systems are designed for public decentralized or centralized networks and cannot meet all the above requirements. To address this problem, we present SecDFS, a decentralized, secure, high-performance file system for consortium networks. SecDFS first combines trusted execution environment and consistent hashing to design a consistent consensus protocol, through which SecDFS distributes files across consortium nodes and achieves data consensus. After that, SecDFS introduces a secure storage constructing method to protect the confidentiality and integrity of both the data and metadata, as well as a two-level cache to speed up the file system performance. We have implemented a prototype of SecDFS and performed a detailed evaluation with it. The results show that SecDFS has 74.83X latency speedup compared with the InterPlanetary File System on average.
Shenglong Zhao, Zhichao Hua 0001, Yubin Xia
ICPADS2
2023 ISA-Grid: Architecture of Fine-grained Privilege Control for Instructions and Registers
abstract
Isolation is a critical mechanism for enhancing the security of computer systems. By controlling the access privileges of software and hardware resources, isolation mechanisms can decouple software into multiple isolated components and enforce the principle of least privilege. While existing isolation systems primarily focus on memory isolation, they overlook the isolation of instruction and register resources, which we refer to as ISA (Instruction Set Architecture) resources. However, previous works have shown that exploiting ISA resources can lead to serious security problems, such as breaking the system's memory isolation property by abusing x86's CR3 register. Furthermore, existing hardware only provides privilege-level-based access control for ISA resources, which is too coarse-grained for software decoupling. For example, ARM Cortex A53 has several hundred system instructions/registers, but only four exception levels (EL0 to EL3) are provided. Additionally, more than 100 instructions/registers for system control are available in only EL1 (the kernel mode). To address this problem, this paper proposes ISA-Grid, an architecture of fine-grained privilege control for instructions and registers. ISA-Grid is a hardware extension that enables the creation of multiple ISA domains, with each domain having different privileges to access instructions and registers. The ISA domain can provide bit-level fine-grained privilege control for registers. We implemented prototypes of ISA-Grid based on two different CPU cores: 1) a RISC-V CPU core on an FPGA board and 2) an x86 CPU core on a simulator. We applied ISA-Grid to different cases, including Linux kernel decomposition and enhancing existing security systems, to demonstrate how ISA-Grid can isolate ISA resources and mitigate attacks based on abusing them. The performance evaluation results on both x86 and RISC-V platforms with real-world applications showed that ISA-Grid has negligible runtime overhead (less than 1%).
Shulin Fan, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang
ISCA2
2022 Colony: A Privileged Trusted Execution Environment With Extensibility
abstract
The code base of system software is growing fast, which results in a large number of vulnerabilities: for example, 296 CVEs have been found in Xen hypervisor and 2195 CVEs in Linux kernel. To reduce the reliance on the trust of system software, many researchers try to provide trusted execution environments (TEEs), which can be categorized into two types: non-privileged TEEs and privileged TEEs. Non-privileged TEEs (e.g., Intel SGX) are extensible, but cannot protect security services like virtual machine introspection (VMI) due to the lack of system-level semantics. On the contrary, privileged TEEs (e.g., the secure world of ARM TrustZone) have system-level semantics, but any additional service implemented in the privileged TEE directly increases the TCB of the entire system. In this article, we propose a new design of TEE to support system-level security services and achieve better extensibility with a small TCB. Each TEE instance of the proposed design is named aColony. Specifically, we introduce asecure monitorfor isolation and capability management. EachColonyis assigned capabilities to access only necessary system-level semantics. We use the new TEE to build four security services, including secure device accessing, VMI tools, a system call tracer, and a much more complex service to virtualize ARM TrustZone with multipleColonies. We have implemented the system on ARMv7 and ARMv8 platforms, in Xen hypervisor and Linux kernel, and perform a detailed evaluation to show its efficiency.11.This paper is an extended version of the conference paper published in USENIX Security’17: vTZ: Virtualizing ARM TrustZone[29]. A brief summary of differences is in Section8.
Yubin Xia, Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan
IEEE Trans. Computers2
2021 TZ-Container: protecting container from untrusted OS with ARM TrustZone
Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang
Sci. China Inf. Sci.1
2021 Boosting Inter-process Communication with Architectural Support
abstract
IPC (inter-process communication) is a critical mechanism for modern OSes, including not only microkernels such as seL4, QNX, and Fuchsia where system functionalities are deployed in user-level processes, but also monolithic kernels like Android where apps frequently communicate with plenty of user-level services. However, existing IPC mechanisms still suffer from long latency. Previous software optimizations of IPC usually cannot bypass the kernel that is responsible for domain switching and message copying/remapping across different address spaces; hardware solutions such as tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt to the new hardware primitives. In this article, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for efficient and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel and supports secure message passing across multiple processes without copying. We have implemented a prototype of XPC based on the ARM AArch64 with Gem5 simulator and RISC-V architecture with FPGA boards. The evaluation shows that XPC can reduce IPC call latency from 664 to 21 cycles, 14×–123× improvement on Android Binder (ARM), and improve the performance of real-world applications on microkernels by 1.6× on Sqlite3.
Yubin Xia, Dong Du 0003, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001, Haibing Guan
ACM Trans. Comput. Syst.3
2020 HCloud: A Serverless Platform for JointCloud Computing
abstract
Most of today's cloud providers, including Amazon, Google and Microsoft, usually offer computing services with the abstraction of virtual machines (VMs), which are also known as instances. Even though these public clouds share similar infrastructures, they provide different service quality standards and price models for users, which fluctuate according to the dynamic loads. The JointCloud computing model empowers the cooperation among multiple public clouds, which can provide the users better service quality with reasonable prices. However, it is difficult to dynamically cooperate among clouds based on the traditional VM abstraction. Any computation migration among clouds will trigger costly VM states transfer, which makes it extremely hard for users to adjust their deployment models swiftly. Fortunately, serverless computing model is getting popular recently and has been supported by major cloud vendors. This model allows users to upload their computing tasks to the cloud in the unit of function. Users do not need to consider all of the tedious works like virtual server maintenance; instead, the cloud automatically instantiates function workers to handle incoming requests on the fly. In this paper, we propose a new JointCloud platform called HCloud, which efficiently manages computing resources of multiple clouds while offering the server-less model to the users. The fine-grained granularity of function enables HCloud to flexibly migrate workloads among clouds, thus brings better service qualities and lower prices at the same time. For example, HCloud can reduce the cost by requesting resource from a cheaper cloud for latency-insensitive jobs while routing the requests of latency-sensitive ones to the nearest and performant clouds. We have implemented a prototype of HCloud and evaluated it by simulating multiple cloud providers. The evaluation results show that HCloud can greatly improve the performance of serverless workloads with small costs.
Zeyu Mi, Zhichao Hua 0001, Yubin Xia
JCC4
2019 XPC: architectural support for secure and efficient cross process call
abstract
Microkernel has many intriguing features like security, fault-tolerance, modularity and customizability, which recently stimulate a resurgent interest in both academia and industry (including seL4, QNX and Google's Fuchsia OS). However, IPC (inter-process communication), which is known as the Achilles' Heel of microkernels, is still the major factor for the overall (poor) OS performance. Besides, IPC also plays a vital role in monolithic kernels like Android Linux, as mobile applications frequently communicate with plenty of user-level services through IPC. Previous software optimizations of IPC usually cannot bypass the kernel which is responsible for domain switching and message copying/remapping; hardware solutions like tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt the new hardware primitives. In this paper, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for fast and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel, and supports message passing across multiple processes through the invocation chain without copying. The primitive is compatible with the traditional address space based isolation mechanism and can be easily integrated into existing microkernels and monolithic kernels. We have implemented a prototype of XPC based on a Rocket RISC-V core with FPGA boards and ported two microkernel implementations, seL4 and Zircon, and one monolithic kernel implementation, Android Binder, for evaluation. We also implement XPC on GEM5 simulator to validate the generality. The result shows that XPC can reduce IPC call latency from 664 to 21 cycles, up to 54.2x improvement on Android Binder, and improve the performance of real-world applications on microkernels by 1.6x on Sqlite3 and 10x on an HTTP server with minimal hardware resource cost.
Dong Du 0003, Zhichao Hua 0001, Yubin Xia, Binyu Zang, Haibo Chen 0001
ISCA2
2018 EPTI: Efficient Defence against Meltdown Attack for Unpatched VMs
Zhichao Hua 0001, Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang
USENIX ATC1
2017 Secure Live Migration of SGX Enclaves on Untrusted Cloud
abstract
The recent commercial availability of Intel SGX (Software Guard eXtensions) provides a hardware-enabled building block for secure execution of software modules in an untrusted cloud. As an untrusted hypervisor/OS has no access to an enclave's running states, a VM (virtual machine) with enclaves running inside loses the capability of live migration, a key feature of VMs in the cloud. This paper presents the first study on the support for live migration of SGX-capable VMs. We identify the security properties that a secure enclave migration process should meet and propose a software-based solution. We leverage several techniques such as two-phase checkpointing and self-destroy to implement our design on a real SGX machine. Security analysis confirms the security of our proposed design and performance evaluation shows that it incurs negligible performance overhead. Besides, we give suggestions on the future hardware design for supporting transparent enclave migration.
Jinyu Gu 0001, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan
DSN2
2017 vTZ: Virtualizing ARM TrustZone
Zhichao Hua 0001, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan
USENIX Security Symposium1