VLDB 2026 Research / reviewers in the wild / expert
Jinyu Gu 0001
dblp:197/6723-1
· DBLP profile ↗
27ranked-venue papers
8as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 12 since 2021Security and privacy · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZoneabstractLarge Language Models (LLMs) deployed on mobile devices offer benefits like user privacy and reduced network latency, but introduce a significant security risk: the leakage of proprietary models to end users. Xunjie Wang, Jiacheng Shi 0002, Yang Yu 0002, Zhichao Hua 0001, Jinyu Gu 0001 |
EuroSys | 6 |
| 2026 | OneSidedMW: Managing Disaggregated Memory Efficiently, Flexibly, and Securely with RNIC Offloading
Jinyu Gu 0001, Xingda Wei, Yubin Xia |
NSDI | 2 |
| 2026 | Experiences with the ChCore Experimental Operating System KernelabstractChCore is an experimental operating system kernel built from the ground up for both education and research. Compared to other similar kernels, ChCore is unique in three aspects: 1) ARM-centric in light of a recent shift of computing from x86 to ARM; 2) microkernelbased given a resurgent interest in microkernel from industry; 3) bimodal for both education and research such that students can easily switch from learning to researching. This paper will describe our five-year experiences in teaching and experimenting operating systems with ChCore. Haibo Chen 0001, Yubin Xia, Jinyu Gu 0001 |
SIGCSE (1) | 3 |
| 2025 | A Hardware-Software Co-Design for Efficient Secure ContainersabstractVM-level containers provide strong isolation by running each container with its own kernel in a VM. However, they rely on virtualization hardware designed for general-purpose VMs, causing non-negligible performance overhead compared to OS-level containers. This performance gap widens dramatically in nested virtualization scenarios, where secure containers run inside a VM. Jiacheng Shi 0002, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia |
EuroSys | 3 |
| 2025 | ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative VerificationabstractDiffusion in transformer is central to advances in high-quality multimodal generation
but suffer from high inference latency due to their iterative nature.
Inspired by speculative decoding's success in accelerating large language models,
we propose Approximate Speculative Diffusion with Speculative Verification (ASDSV),
a novel method to enhance the efficiency of diffusion models.
Adapting speculative execution to diffusion processes presents unique challenges.
First, the substantial computational cost of verifying numerous speculative steps
for continuous, high-dimensional outputs makes traditional full verification prohibitively expensive.
Second, determining the optimal number of speculative steps $K$
involves a trade-off between potential acceleration and verification success rates.
To address these, ASDSV introduces two key innovations:
1) A speculative verification technique, which leverages the observed temporal correlation between draft and target model outputs,
efficiently validates $K$ speculative steps by only checking the alignment of the initial and final states, significantly reducing verification overhead.
2) A multi-stage speculative strategy that adjusts $K$ according to the denoising phase—employing smaller $K$ during volatile early stages
and larger $K$ during more stable later stages to optimize the balance between speed and quality.
We apply ASDSV to state-of-the-art diffusion transformers,
including Flux.1-dev for image generation and Wan2.1 for video generation.
Extensive experiments demonstrate that ASDSV achieves up to 1.77$\times$-3.01$\times$ speedup
in model inference with a minimal 0.3\%-0.4\% drop in VBench score,
showcasing its effectiveness in accelerating multimodal diffusion models without significant quality degradation.
The code will be publicly available once the acceptance of the paper. Kaijun Zhou 0001, Xingda Wei, Xijun Li, Jinyu Gu 0001 |
NeurIPS | 5 |
| 2025 | ODRP: On-Demand Remote Paging with Programmable RDMA
Xingda Wei, Jinyu Gu 0001, Hongrui Xie, Rong Chen 0001, Haibo Chen 0001 |
NSDI | 3 |
| 2025 | PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated SpeculationabstractPhoenixOS (PhOS) is the first OS service that can concurrently checkpoint and restore (C/R) GPU processes—a fundamental capability for critical tasks such as fault tolerance, process migration, and fast startup. While concurrent C/R is well-established on CPUs, it poses unique challenges on GPUs due to their lack of essential features for efficiently tracing concurrent memory reads and writes, such as specific hardware capabilities (e.g., dirty bits) and OS-mediated data paths (e.g., copy-on-write). Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen 0001, Mingcong Han, Jinyu Gu 0001, Haibo Chen 0001 |
SOSP | 7 |
| 2025 | μEFI: A Microkernel-Style UEFI with Isolation and Transparency
Yiyang Wu, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001 |
USENIX ATC | 3 |
| 2025 | SAVE: Software-Implemented Fault Tolerance for Model Inference against GPU Memory Bit Flips
Wenxin Zheng, Jinyu Gu 0001, Haibo Chen 0001 |
USENIX ATC | 3 |
| 2025 | Serverless Functions Made Confidential and Efficient with Split Containers
Jiacheng Shi 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001 |
USENIX Security Symposium | 2 |
| 2025 | Harmonizing Security and Performance in Microkernel File Servers
Wentai Li, Jinyu Gu 0001, Yubin Xia, Binyu Zang |
J. Comput. Sci. Technol. | 3 |
| 2025 | HTLL: Latency-Aware Scalable Blocking MutexabstractThis paper finds that existing mutex locks suffer from throughput collapses or latency collapses, or both, in the oversubscribed scenarios where applications create more threads than the CPU core number, e.g., database applications like mysql use per thread per connection. We make an in-depth performance analysis on existing locks and then identify three design rules for the lock primitive to achieve scalable performance in oversubscribed scenarios. First, to achieve ideal throughput, the lock design should keep adequate number of active competitors. Second, the active competitors should be arranged carefully to avoid the lock-holder preemption problem. Third, to meet latency requirements, the lock design should track the latency of each competitor and reorder the competitors according to the latency requirement. We propose a new lock library called HTLL that satisfies these rules and achieves both high throughput and low latency even when the cores are oversubscribed. HTLL only requires minimal human effort (e.g., add several lines of code) to annotate the latency requirement. Evaluation results show that HTLL achieves scalable performance in the oversubscribed scenarios. Specifically, for the real-world database, LMDB, HTLL can reduce the tail latency by up to 97% with only an average 5% degradation in throughput, compared with state-of-the-art alternatives such as Malthusian, CST, and Mutexee locks; In comparison to the widely used pthread mutex lock, it can increase the throughput by up to 22% and decrease the latency by up to 80%. Meanwhile, for the under-subscribed scenarios, it also shows comparable performance than state-of-the-art blocking locks. Ziqu Yu, Jinyu Gu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Understanding the IO Performance Gap Between OS-Level and VM-Level Containers in High-Density DeploymentabstractContainers are widely deployed in clouds. There are two common container architectures: operating system-level (OS-level) container and virtual machine-level (VM-level) container. Typical examples are runc and Kata. It is well known that VM- level containers provide better isolation than OS-level containers, but at a higher overhead. Although there are quantitative analyses of the performance gap between these two container architectures, they rarely discuss the performance gap under the constrained resources nrovisioned to containers. Since the high-density deployment of containers is demanding in the cloud, each container is provisioned with limited resources specified by the cgroup mechanism. In this paper, we provide an in-depth analysis of the storage and network (two key aspects) performance differences between runc and Kata under varying resource constraints. We identify configuration implications that are crucial to performance and find that some of them are not exposed by the Kata interfaces. Based on that, we propose a profiling tool to automatically offer configuration suggestions for optimizing container performance. Our evaluation shows that the auto-generated configuration can improve the performance of MySQL by up to 107% in the TPCC benchmark compared with the default Kata setup. Wentai Li, Kaijun Zhou 0001, Jiacheng Shi 0002, Xingman Chen, Luyuan Wang, Jinyu Gu 0001 |
ICDCS | 7 |
| 2023 | No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless Computing
Xingda Wei, Fangming Lu, Tianxia Wang, Jinyu Gu 0001, Rong Chen 0001, Haibo Chen 0001 |
OSDI | 4 |
| 2023 | Understanding and Mitigating Twin Function Misuses in Operating System KernelabstractMajor operating system kernels expose twin functions, which are groups of internal primitives that have mostly common but slightly diverging semantics, to kernel modules and subsystems. They are created to make the basic primitives work well in various scenarios. Unfortunately, though being expected as solutions, twin functions may turn to problem-makers in practice. As we have observed from over 500 patches applied to upstream Linux and FreeBSD, developers choose an improper one from the twins, leaving the kernel with stability and security bugs as well as error-prone code. In this paper, we aim to understand and mitigate the twin function misuse problem. First, we provide an informative discussion on the misuse-fix patches. We find that violating the constraints from calling context, missing the primitives with better performance, lacking the necessary security enhancements, and breaking the kernel coding style are the four major factors that lead to misuse. We then identify the programming rules from the patches and apply them with a static program analysis tool extended from Coccinelle, including callgraph tainting and type-based function pointer resolving. We have 136 patches accepted by the Linux community and fix 320 new misuses in the upstream Linux kernel. Jinyu Gu 0001, Jiacheng Shi 0002, Haroran Su, Wentai Li, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
IEEE Trans. Computers | 1 |
| 2022 | Asymmetry-aware scalable lockingabstractThe pursuit of power-efficiency is popularizing asymmetric multicore processors (AMP) such as ARM big.LITTLE, Apple M1 and recent Intel Alder Lake with big and little cores. However, we find that existing scalable locks fail to scale on AMP and cause collapses in either throughput or latency, or both, because their implicit assumption of symmetric cores no longer holds. To address this issue, we propose the first asymmetry-aware scalable lock named LibASL. LibASL provides a new lock ordering guided by applications' latency requirements, which allows big cores to reorder with little cores for higher throughput under the condition of preserving applications' latency requirements. Using LibASL only requires linking the applications with it and, if latency-critical, inserting few lines of code to annotate the coarse-grained latency requirement. We evaluate LibASL in various benchmarks including five popular databases on Apple M1. Evaluation results show that LibASL can improve the throughput by up to 5 times while precisely preserving the tail latency designated by applications. Jinyu Gu 0001, Dahai Tang, Kenli Li 0001, Binyu Zang, Haibo Chen 0001 |
PPoPP | 2 |
| 2022 | EPK: Scalable and Efficient Memory Protection Keys
Jinyu Gu 0001, Wentai Li, Yubin Xia, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2022 | A Hardware-Software Co-design for Efficient Intra-Enclave Isolation
Jinyu Gu 0001, Bojun Zhu, Wentai Li, Yubin Xia, Haibo Chen 0001 |
USENIX Security Symposium | 1 |
| 2022 | Unified Enclave Abstraction and Secure Enclave Migration on Heterogeneous Security Architectures
Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Chenggang Qin, Zhengyu He |
J. Comput. Sci. Technol. | 1 |
| 2022 | Colony: A Privileged Trusted Execution Environment With ExtensibilityabstractThe code base of system software is growing fast, which results in a large number of vulnerabilities: for example, 296 CVEs have been found in Xen hypervisor and 2195 CVEs in Linux kernel. To reduce the reliance on the trust of system software, many researchers try to provide trusted execution environments (TEEs), which can be categorized into two types: non-privileged TEEs and privileged TEEs. Non-privileged TEEs (e.g., Intel SGX) are extensible, but cannot protect security services like virtual machine introspection (VMI) due to the lack of system-level semantics. On the contrary, privileged TEEs (e.g., the secure world of ARM TrustZone) have system-level semantics, but any additional service implemented in the privileged TEE directly increases the TCB of the entire system. In this article, we propose a new design of TEE to support system-level security services and achieve better extensibility with a small TCB. Each TEE instance of the proposed design is named aColony. Specifically, we introduce asecure monitorfor isolation and capability management. EachColonyis assigned capabilities to access only necessary system-level semantics. We use the new TEE to build four security services, including secure device accessing, VMI tools, a system call tracer, and a much more complex service to virtualize ARM TrustZone with multipleColonies. We have implemented the system on ARMv7 and ARMv8 platforms, in Xen hypervisor and Linux kernel, and perform a detailed evaluation to show its efficiency.11.This paper is an extended version of the conference paper published in USENIX Security’17: vTZ: Virtualizing ARM TrustZone[29]. A brief summary of differences is in Section8. Yubin Xia, Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan |
IEEE Trans. Computers | 4 |
| 2021 | Efficiently Recovering Stateful System Components of Multi-server MicrokernelsabstractMicrokernel OSes provide OS services through mutually-isolated system servers running in different user processes, which brings stronger fault isolation than monolithic OSes. Nevertheless, considering the fault recovery capability of system servers, most existing microkernel OSes usually do no more than restarting a fault server, which will cause a server to lose all its running states and then may affect all the applications relying on it. In this paper, we present a mechanism named TxIPC that can efficiently recover stateful system servers on microkernel OSes. Since a system server provides the service by inter-process communication (IPC), TxIPC makes it fault resilient by handling each IPC in a transaction-like manner. Specifically, if a fault happens in a server (during one IPC handling procedure), TxIPC aborts all the updates made by the IPC and thus recovers the server from that fault. Evaluations show that TxIPC can enable servers to recover from 99.8% (injected) faults with 3%-45 % performance overhead on application benchmarks, which significantly outperforms existing counterparts. Wentai Li, Jinyu Gu 0001, Binyu Zang |
ICDCS | 2 |
| 2021 | TZ-Container: protecting container from untrusted OS with ARM TrustZone
Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang |
Sci. China Inf. Sci. | 3 |
| 2021 | Enclavisor: A Hardware-Software Co-Design for Enclaves on Untrusted CloudabstractThe releases of Intel SGX and AMD SEV mark the transition of hardware-based enclaves from research prototypes to mainstream products. These two paradigms of secure enclaves are attractive to both the cloud providers and tenants, since security is one of the key pillars of cloud computing. However, it is found that current hardware-defined enclaves are not flexible and efficient enough for the cloud. For example, although SGX can provide strong memory protection with both confidentiality and integrity, the size of secure memory is tightly restricted. On the contrary, SEV enables enclaves to use more memory but has critical security flaws due to no memory integrity protection. Meanwhile, both types of enclaves have relatively long booting latency, which makes them not suitable for short-term tasks like serverless workloads. After an in-depth analysis, we find that there are some intrinsic tradeoffs between security and performance due to the limitation of architectural designs. In this article, we investigate a novel hardware-software co-design of enclaves to meet the requirements of cloud by placing a part of the logic of the enclave mechanism into a lightweight software layer, named Enclavisor, to achieve a balance between security, performance, and flexibility. Specifically, our implementation is based on AMD's SEV and, Enclavisor is placed in the guest kernel mode of SEV's secure virtual machines. Enclavisor inherently supports memory encryption with no memory limitation and also achieves efficient booting, multiple enclave granularities, and post-launch remote attestation. Meanwhile, we also propose hardware/software solutions to mitigate the security flaws caused by the lack of memory integrity. We implement a prototype of Enclavisor on an AMD SEV server. The experiments on both micro-benchmarks and application benchmarks show that enclaves on Enclavisor can have close-to-native performance. Jinyu Gu 0001, Xinyue Wu, Bojun Zhu, Yubin Xia, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
IEEE Trans. Computers | 1 |
| 2020 | Harmonizing Performance and Isolation in Microkernels with Efficient Intra-kernel Isolation and Communication
Jinyu Gu 0001, Xinyue Wu, Wentai Li, Zeyu Mi, Yubin Xia, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2019 | Pisces: A Scalable and Efficient Persistent Transactional Memory
Jinyu Gu 0001, Xiayang Wang, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
USENIX ATC | 1 |
| 2017 | Secure Live Migration of SGX Enclaves on Untrusted CloudabstractThe recent commercial availability of Intel SGX (Software Guard eXtensions) provides a hardware-enabled building block for secure execution of software modules in an untrusted cloud. As an untrusted hypervisor/OS has no access to an enclave's running states, a VM (virtual machine) with enclaves running inside loses the capability of live migration, a key feature of VMs in the cloud. This paper presents the first study on the support for live migration of SGX-capable VMs. We identify the security properties that a secure enclave migration process should meet and propose a software-based solution. We leverage several techniques such as two-phase checkpointing and self-destroy to implement our design on a real SGX machine. Security analysis confirms the security of our proposed design and performance evaluation shows that it incurs negligible performance overhead. Besides, we give suggestions on the future hardware design for supporting transparent enclave migration. Jinyu Gu 0001, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan |
DSN | 1 |
| 2017 | vTZ: Virtualizing ARM TrustZone
Zhichao Hua 0001, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan |
USENIX Security Symposium | 2 |