EDBT 2026 Demo / reviewers in the wild / expert
Dong Du 0003
dblp:48/331-3
· DBLP profile ↗
34ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-7945-8430ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 4 first-author · 21 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 11 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRLabstractWith the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlike traditional pre-training approaches, RL encompasses multiple stages: rollout, reward, and training, which necessitates collaboration among various worker types. However, current RL systems continue to grapple with substantial GPU underutilization, due to two primary factors: (1) The rollout stage dominates the overall RL process due to test-time scaling; (2) Imbalances in rollout lengths (within the same batch) result in GPU bubbles. While prior solutions like asynchronous execution and truncation offer partial relief, they may compromise training accuracy for efficiency. Our key insight stems from a previously overlooked observation: rollout responses exhibit remarkable similarity across adjacent training epochs. Based on the insight, we introduce RhymeRL, an LLM RL system designed to accelerate RL training with two key innovations. First, to enhance rollout generation, we present HistoSpec, a speculative decoding inference engine that utilizes the similarity of historical rollout token sequences to obtain accurate drafts. Second, to tackle rollout bubbles, we introduce HistoPipe, a two-tier scheduling strategy that leverages the similarity of historical rollout distributions to balance workload among rollout workers. Experimental results demonstrate that RhymeRL achieves up to a 2.6x performance improvement over existing methods, without compromising accuracy or modifying the RL paradigm. Jingkai He, Tianjian Li, Erhu Feng, Dong Du 0003, Qian Liu 0033, Yubin Xia, Haibo Chen 0001 |
ASPLOS (2) | 4 |
| 2026 | SKernel: An Elastic and Efficient Secure Container System at Scale with a Split-Kernel ArchitectureabstractSecure containers leverage hardware virtualization to isolate container sandboxes, enabling dedicated guest kernels to mitigate shared kernel attacks prevalent in traditional systems. However, existing approaches struggle with a fundamental trade-off: VM-based solutions (e.g., Kata) prioritize performance but lack elasticity and on-demand usage for volatile and bursty workloads, while lightweight methods (e.g., gVisor) rely on the host kernel for dynamic resource management at the cost of significant performance degradation due to guest-host dependencies. Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Guotao Tan, Anqi Shen, Dawei Shen, Xinyao Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
EuroSys | 14 |
| 2026 | Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
Mo Zou, Hengbin Zhang, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
FAST | 4 |
| 2026 | CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Yanning Yang, Dong Du 0003, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen 0001 |
ISCA | 5 |
| 2026 | miniK8s: A Pedagogical Cloud-Native SystemabstractThe rapid adoption of cloud-native technologies, particularly containerization and orchestration with systems like Kubernetes, necessitates their effective integration into undergraduate computer science curricula. However, the complexity of production-grade cloud-native systems present a steep learning curve. Traditional pedagogical approaches often involve either oversimplified toy projects built from scratch, which lack real-world relevance, or the direct use of complex cloud systems, which can obscure fundamental concepts. To overcome these barriers, we introduce miniK8s, a lightweight, Kubernetes-like platform designed for an undergraduate cloud computing course. miniK8s distinguishes itself by promoting the pedagogical vision of teaching students to build a substantial and realistic system by integrating existing, robust open-source components with a hand-written, simplified cornerstone component. With miniK8s, students can both learn the key concepts inside the cloud architecture (e.g., resource scaling) and build practical systems with open-source building blocks. We have successfully used miniK8s as a project in a cloud computing course for hundreds of undergraduate students, which can significantly enhance students' understanding of how real-world cloud-native systems work and their ability to build a complex system. Dong Du 0003, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang |
SIGCSE (1) | 1 |
| 2026 | LayerTEE: Decoupled Memory Protection for Scalable Multilayer Communication on RISC-VabstractThe Trusted Execution Environment (TEE) has been widely implemented by modern hardware vendors to protect security and privacy-sensitive applications and data, such as Intel SGX/TDX, ARM TrustZone, AMD SEV, and RISC-V Penglai. However, existing TEE systems face challenges in balancing memory isolation among security, performance, and scalability requirements. Segment-based memory isolation mechanisms, like RISC-V PMP, struggle to scale effectively to the large number of segments needed for confidential cloud and data center environments. On the other hand, table-based isolation methods, such as page tables, combine address translation with memory protection, leading to inefficient cross-enclave communication and potential security vulnerabilities like Rowhammer attacks. This paper introduces a novel TEE system, which decouples memory protection (to segments) from address translation (to page tables). This design improves communication performance by dynamically adjusting memory protection capabilities, without sacrificing application compatibility. LayerTEE enhances enclave security and scalability by designing a multi-layer segment-based isolation mechanism. We have built a prototype of based on FPGA, incorporating hardware extensions and software support. The evaluation demonstrates that significantly surpasses existing TEE solutions, achieving three orders of magnitude lower communication latency and 10x greater scalability while maintaining robust security guarantees. Shangjie Pan, Yinghao Yang 0001, Xuanyao Peng, Xiquan Zhao, Dong Du 0003, Yubin Xia, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and PracticabilityabstractCurrent serverless platforms struggle to optimize resource utilization for both CPU and GPU functions due to their dynamic and fine-grained nature. Conventional techniques like overcommitment and autoscaling fall short, often sacrificing utilization for practicability or incurring performance tradeoffs. Overcommitment requires predicting performance to prevent QoS violation, introducing tradeoff between prediction accuracy and overheads. Autoscaling requires scaling instances in response to load fluctuations quickly to reduce resource wastage, but more frequent scaling also leads to more cold start overheads. The rich concurrency of GPU resources further complicates GPU instance orchestration, such as setting right batch sizes. This article introduces Jiagu to harmonize efficiency with practicability through the following novel techniques. First, pre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and scheduling. Second, dual-staged scaling achieves frequent adjustment of instances with minimum overhead. Third, Jiagu conducts an in-depth analysis about the complexity of the relationship between GPU function configuration and execution. It then proposes batch-aware scaling that achieves optimal configurations for both batch size setting and autoscaling, addressing all the challenges according to the analysis. We have implemented a prototype and evaluated it using real-world applications and traces from the public cloud platform. Our evaluation shows an improvement in deployment density over commercial clouds (with Kubernetes) while maintaining QoS for both CPU and GPU functions (54.8% and 18% respectively), and 81.0%–93.7% lower scheduling costs and a 57.4%–69.3% reduction in cold start latency compared to existing QoS-aware schedulers. Yanning Yang, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
ACM Trans. Comput. Syst. | 3 |
| 2025 | Dep-TEE: Decoupled Memory Protection for Secure and Scalable Inter-enclave Communication on RISC-VabstractTrusted Execution Environment (TEE) has been widely implemented by modern hardware vendors to protect security and privacy-sensitive applications and data, such as Intel SGX/TDX, ARM TrustZone, AMD SEV, and RISC-V Penglai. However, existing TEE systems face challenges in balancing memory isolation among security, performance, and scalability requirements. This paper introduces a novel TEE system, Dep-TEE, which decouples memory protection (to segments) from address translation (to page tables). This design improves communication performance by dynamically adjusting memory protection capabilities, without sacrificing application compatibility, and enhances security by safeguarding against attacks on page tables. We have built a prototype of Dep-TEE based on FPGA, incorporating hardware extensions and software support. The evaluation demonstrates that Dep-TEE significantly surpasses existing TEE solutions, achieving three orders of magnitude lower communication latency and 10x greater scalability while maintaining robust security guarantees. Shangjie Pan, Xuanyao Peng, Zeyuan Man, Xiquan Zhao, Dongrong Zhang, Bicheng Yang, Dong Du 0003, Yubin Xia, Xiaowei Li 0001 |
ASP-DAC | 7 |
| 2025 | D-VSync: Decoupled Rendering and Displaying for Smartphone GraphicsabstractRendering service, which typically orchestrates screen display and UI through Vertical Synchronization (VSync), is an indispensable system service for user experiences of smartphone OSes (e.g., Android, OpenHarmony, and iOS). The recent trend of large high-frame-rate screens, stunning visual effects, and physics-based animations has placed unprecedented pressure on the VSync-based rendering architecture, leading to higher frame drops and longer rendering latency. Yuanpei Wu, Dong Du 0003, Yubin Xia, Ming Fu, Binyu Zang, Haibo Chen 0001 |
ASPLOS (1) | 2 |
| 2025 | Offloading Cloud-Native Infrastructure with XpuPodabstractCloud-native systems increasingly rely on infrastructure services (e.g., service meshes, monitoring agents), which compete for resources with user applications, thereby degrading performance and scalability. We propose XpuPod, a new abstraction that offloads these services to Data Processing Units (DPUs) to enforce strict isolation while reducing host resource contention and operational costs. XpuPod enables a cross-processing-unit (XPU) network system. This system features two key components: (1) transparent XPU networking, which provides a unified network abstraction for processes spanning both the CPU and DPU, and (2) elastic and efficient XPU communication, a mechanism that achieves shared-memory performance without the costs of pinned resources. By leveraging the XPU network system and the compositional nature of cloud-native workloads, XpuPod can optimally offload infrastructure containers to DPUs. We implement XpuPod on Linux and a corresponding cloud-native system, XpuK8s, on Kubernetes. We evaluate our system using NVIDIA BlueField-2 DPUs and a simulator backed by a CXL memory device. The results show that XpuK8s effectively supports complex, unmodified, commodity cloud-native applications. Compared to a kernel-bypass design, XpuK8s provides up to 31.9x lower latency and consumes 64x fewer resources. Furthermore, compared to state-of-the-art systems, XpuK8s can achieve 60% lower end-to-end latency and 55% higher scalability. Bicheng Yang, Jingkai He, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
SoCC | 3 |
| 2025 | Topology-Aware Virtualization over Inter-Core Connected Neural Processing UnitsabstractWith the rapid development of artificial intelligence (AI) applications, an emerging class of AI accelerators, termed Inter-core Connected Neural Processing Units (NPU), has been adopted in both cloud and edge computing environments, like Graphcore IPU, Tenstorrent, etc.Despite their innovative design, these NPUs often demand substantial hardware resources, leading to suboptimal resource utilization due to the imbalance of hardware requirements across various tasks.To address this issue, prior research has explored virtualization techniques for monolithic NPUs, but has neglected inter-core connected NPUs with the hardware topology.This paper introduces vNPU, the first comprehensive virtualization design for inter-core connected NPUs, integrating three novel techniques: (1) NPU route virtualization, which redirects instruction and data flow from virtual NPU cores to physical ones, creating a virtual topology; (2) NPU memory virtualization, designed to minimize translation stalls for SRAM-centric and NoC-equipped NPU cores, thereby maximizing the memory bandwidth; and (3) Besteffort topology mapping, which determines the optimal mapping from all candidate virtual topologies, balancing resource utilization with end-to-end performance.We have developed a prototype of vNPU on both an FPGA platform (Chipyard+FireSim) and a simulator (DCRA).Evaluation results demonstrate that when executing multiple NPU workloads on virtual NPUs, vNPU achieves performance improvements of up to 1.92x and 1.28x for the Transformer and ResNet models, respectively, in comparison to the MIG-based virtualization method.Furthermore, the hardware performance * Both authors contributed equally to this research. Dahu Feng, Erhu Feng, Dong Du 0003, Pinjie Xu, Yubin Xia, Haibo Chen 0001 |
ISCA | 3 |
| 2025 | Sprig: Low-Latency Startup for Knative Serverless Platforms with Async-CforkabstractServerless computing has emerged as a transformative paradigm in the cloud-native era, widely adopted across major public cloud platforms. However, its pay-as-use model necessitates frequent dynamic instance creation and destruction, making cold start latency a critical system bottleneck. While the industry predominantly employs caching techniques to mitigate cold start overheads, these approaches exhibit inherent limitations. Existing research prototypes often optimize container runtimes manually, resulting in systems incompatible with production environments that cannot leverage automated management capabilities or integrate optimizations effectively. Moreover, runtime optimizations fail to fit system features like asynchronous cold start. This paper presents Sprig, the first serverless system that seamlessly integrates cutting-edge academic research with production-grade platforms. We introduce a novel computational abstraction called Template Pod to bridge the gap between research-oriented flexible designs and minimal computational units in production systems. Sprig further incorporates Asynccfork, a low-latency startup technique optimized for large-scale serverless platforms, significantly reducing function cold start latency. Additionally, we design a multiplexed proxy queue with enhanced scheduling elasticity to resolve request accumulation issues in Knative's asynchronous startup mechanism. Evaluations show Sprig achieves a 70% average reduction in cold start latency compared with Knative, and decreases P99 end-to-end latency from 27s to 0.5s under high load while maintaining only 16.7% of baseline memory footprint for 20 instances. Dong Du 0003, Liang Zhang 0010, Yubin Xia |
JCC | 2 |
| 2025 | Fork in the Road: Reflections and Optimizations for Cold Start Latency in Production Serverless Systems
Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Anqi Shen, Dawei Shen, Qi Xing, Shun Song, Tongkai Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
OSDI | 14 |
| 2025 | How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS ServiceabstractIn modern systems, memory copy remains a critical performance bottleneck across various scenarios, playing a pervasive role in system-wide execution such as syscalls, IPC, and user-mode applications. Numerous efforts have aimed at optimizing copy performance, including zero-copy with page remapping and hardware-accelerated copy. However, they typically target specific use cases, such as Linux zero-copy send() for messages of ≥10KB. This paper argues for copy as a first-class OS service, offering three key benefits: (1) with the asynchronous copy abstraction provided by the service, applications can overlap their execution with copy; (2) the service can effectively utilize hardware capabilities to enhance copy performance; (3) the service's global view of copies further enables holistic optimization. To this end, we introduce Copier, a new OS service of coordinated asynchronous copy, to serve both user-mode applications and OS services. We build Copier-Linux to demonstrate Copier's ability to improve performance for diverse use cases, including Redis, Protobuf, network stack, proxy, etc. Evaluations show that Copier achieves up to a 1.8 × speedup for real-world applications like Redis and a 1.6 × improvement over zIO, the state-of-the-art in optimizing copy efficiency. To further facilitate adoption, we develop a toolchain to ease the use of Copier. We also integrate Copier into a commercial smartphone OS (HarmonyOS 5.0), achieving promising results. Jingkai He, Yunpeng Dong, Dong Du 0003, Mo Zou, Zhitai Yu, Yuxin Ren 0001, Ning Jia 0004, Yubin Xia, Haibo Chen 0001 |
SOSP | 3 |
| 2024 | sIOPMP: Scalable and Efficient I/O Protection for TEEsabstractTrusted Execution Environments (TEEs), like Intel SGX/TDX, AMD SEV-SNP, ARM TrustZone/CCA, have been widely adopted in prevailing architectures. However, these TEEs typically do not consider I/O isolation (e.g., defending against malicious DMA requests) as a first-class citizen, which may degrade the I/O performance. Traditional methods like using IOMMU or software I/O can degrade throughput by at least 20% for I/O intensive workloads. The main reason is that the isolation requirements for I/O devices differ from CPU ones. This paper proposes a novel I/O isolation mechanism for TEEs, named sIOPMP (scalable I/O Physical Memory Protection), with three key features. First, we design a Multi-stage-Tree-based checker, supporting more than 1,000 hardware regions. Second, we classify the devices into hot and cold, and support unlimited devices with the mountable entry. Third, we propose a remapping mechanism to switch devices between hot and cold status for dynamic I/O workloads. Evaluation results show that sIOPMP introduces only negligible performance overhead for both benchmarks and real-world workloads, and improves 20% ~ 38% network throughput compared with IOMMU-based mechanisms or software I/O adopted in TEEs. Erhu Feng, Dahu Feng, Dong Du 0003, Yubin Xia, Siqi Zhao, Haibo Chen 0001 |
ASPLOS (2) | 3 |
| 2024 | On-demand and Parallel Checkpoint/Restore for GPU ApplicationsabstractLeveraging serverless computing for cloud-based machine learning services is on the rise, promising cost-efficiency and flexibility are crucial for ML applications relying on high-performance GPUs and substantial memory. However, despite modern serverless platforms handling diverse devices like GPUs seamlessly on a pay-as-you-go basis, a longstanding challenge remains: startup latency, a well-studied issue when serverless is CPU-centric. For example, initializing GPU apps with minor GPU models, like MobileNet, demands several seconds. For more intricate models such as GPT-2, startup latency can escalate to around 10 seconds, vastly overshadowing the short computation time for GPU-based inference. Prior solutions tailored for CPU serverless setups, like fork() and Checkpoint/Restore, cannot be directly and effectively applied due to differences between CPUs and GPUs. Yanning Yang, Dong Du 0003, Haitao Song 0001, Yubin Xia |
SoCC | 2 |
| 2024 | sNPU: Trusted Execution Environments on Integrated NPUsabstractTrusted execution environment (TEE) promises strong security guarantee with hardware extensions for security-sensitive tasks. Due to its numerous benefits, TEE has gained widespread adoption, and extended from CPU-only TEEs to FPGA and GPU TEE systems. However, existing TEE systems exhibit inadequate and inefficient support for an emerging (and significant) processing unit, NPU. For instance, commercial TEE systems resort to coarse-grained and static protection approaches for NPUs, resulting in notable performance degradation (10%–20%), limited (or no) multitasking capabilities, and suboptimal resource utilization. In this paper, we present a secure NPU architecture, known as sNPU, which aims to mitigate vulnerabilities inherent to the design of NPU architectures. First, sNPU proposes NPU Guarder to enhance the NPU’s access control. Second, sNPU defines new attack surfaces leveraging in-NPU structures like scratchpad and NoC, and designs NPU Isolator to guarantee the isolation of scratchpad and NoC routing. Third, our system introduces a trusted software module called NPU Monitor to minimize the software TCB. Our prototype, evaluated on FPGA, demonstrates that sNPU significantly mitigates the runtime costs associated with security checking (from upto 20% to 0%) while incurring less than 1% resource costs. Erhu Feng, Dahu Feng, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
ISCA | 3 |
| 2024 | Using Dynamically Layered Definite Releases for Verifying the RefFS File System
Mo Zou, Dong Du 0003, Mingkai Dong 0002, Haibo Chen 0001 |
OSDI | 2 |
| 2024 | Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with Jiagu
Yanning Yang, Dong Du 0003, Yubin Xia, James R. Larus, Haibo Chen 0001 |
USENIX ATC | 3 |
| 2023 | The Gap Between Serverless Research and Real-world SystemsabstractWith the emergence of the serverless computing paradigm in the cloud, researchers have explored many challenges of serverless systems and proposed solutions such as snapshot-based booting. However, we have noticed that some of these optimizations are based on oversimplified assumptions that lead to infeasibility and hide real-world issues. This paper aims to analyze the gap between current serverless research and real-world systems from a perspective of industry, and present new observations, challenges, opportunities, and insights that may address the discrepancies. Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
SoCC | 2 |
| 2023 | Efficient Distributed Secure Memory with Migratable Merkle TreeabstractHardware-assisted enclaves with memory encryption have been widely adopted in the prevailing architectures, e.g., Intel SGX/TDX, AMD SEV, ARM CCA, etc. However, existing enclave designs fall short in supporting efficient cooperation among cross-node enclaves (i.e., multi-machines) because the range of hardware memory protection is within a single node. A naive approach is to leverage cryptography at the application level and transfer data between nodes through secure channels (e.g., SSL). However, it incurs orders of magnitude costs due to expensive encryption/decryption, especially for distributed applications with large data transfer, e.g., MapReduce and graph computing. A secure and efficient mechanism of distributed secure memory is necessary but still missing.This paper proposes Migratable Merkle Tree (MMT), a design enabling efficient distributed secure memory to support distributed confidential computing. MMT sets up an integrity forest for distributed memory on multiple nodes. It allows an enclave to securely delegate an MMT closure, which contains both data and metadata of a subtree, to a remote enclave. By reusing the memory encryption mechanisms of existing enclaves, our design achieves secure data transfer without software re-encryption. We have implemented a prototype of MMT and a trusted firmware for management, and further applied MMT to real-world distributed applications. The evaluation results show that compared with existing systems using the AES-NI instruction, MMT can achieve up to 13x speedup on data transferring, and gain 12%~58% improvement on the end-to-end performance of MapReduce and PageRank. Erhu Feng, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
HPCA | 2 |
| 2023 | Secure and Efficient Runtime Environment for Smart Contracts on JointCloudabstractMany cloud providers, including Amazon, Google, Microsoft, and Alibaba Cloud, offer support for blockchain cloud services that rely on a runtime environment, such as the Ethereum Virtual Machine (EVM), to execute smart contracts and ensure consistency between participants. However, existing runtime systems suffer from two main limitations. Firstly, traditional runtime systems like EVM cannot guarantee privacy protection as all the data uploaded to the blockchain is visible to all participants. This restricts the use of blockchain in limited scenarios. Secondly, each computation on the runtime system must be synchronized to all nodes in the network, resulting in a significant increase in computational overhead, which can be challenging to implement for more complex applications. One approach to address these limitations is to utilize Trusted Execution Environments (TEE) for blockchain runtime, which can provide privacy protection and mitigate redundant synchronization operations. However, using TEE for blockchain may significantly increase cloud costs. To overcome these challenges, this paper proposes PL-EVM, a new runtime environment for smart contracts that utilizes jointcloud. PL-EVM achieves high-security guarantees by using TEE to protect privacy-sensitive data and incorporates dynamic migration and splitting mechanisms to achieve high efficiency and low costs. Our evaluation results show that PL-EVM can improve performance and reduce costs by 4% to 32.22%. Yuhao Xue, Dong Du 0003, Yubin Xia |
JCC | 2 |
| 2023 | Accelerating Extra Dimensional Page Walks for Confidential ComputingabstractTo support highly scalable and fine-grained computing paradigms such as microservices and serverless computing better, modern hardware-assisted confidential computing systems, such as Intel TDX and ARM CCA, introduce permission table to achieve fine-grained and scalable memory isolation among different domains. However, it also adds an extra dimension to page walks besides page tables, leading to significantly more memory references (e.g., 4 → 12 for RISC-V Sv39)1. We observe that most costs (about 75%) caused by the extra dimension of page walks are used to validate page table pages. Based on this observation, this paper proposes HPMP (Hybrid Physical Memory Protection), a hardware-software co-design (on RISC-V) that protects page table pages using segment registers and normal pages using permission tables to balance scalability and performance. We have implemented HPMP and Penglai-HPMP (a TEE system based on HPMP) on FPGA with two RISC-V cores (both in-order and out-of-order). Evaluation results show that HPMP can reduce costs by 23.1%–73.1% on BOOM and significantly improve performance on real-world applications, including serverless computing (FunctionBench) and Redis. Dong Du 0003, Bicheng Yang, Yubin Xia, Haibo Chen 0001 |
MICRO | 1 |
| 2022 | Serverless computing on heterogeneous computersabstractExisting serverless computing platforms are built upon homogeneous computers, limiting the function density and restricting serverless computing to limited scenarios. We introduce Molecule, the first serverless computing system utilizing heterogeneous computers. Molecule enables both general-purpose devices (e.g., Nvidia DPU) and domain-specific accelerators (e.g., FPGA and GPU) for serverless applications that significantly improve function density (50% higher) and application performance (up to 34.6x). To achieve these results, we first propose XPU-Shim, a distributed shim to bridge the gap between underlying multi-OS systems (when using general-purpose devices) and our serverless runtime (i.e., Molecule). We further introduce vectorized sandbox, a sandbox abstraction to abstract hardware heterogeneity (when using domain-specific accelerators). Moreover, we also review state-of-the-art serverless optimizations on startup and communication latency and overcome the challenges to implement them on heterogeneous computers. We have implemented Molecule on real platforms with Nvidia DPUs and Xilinx FPGAs and evaluate it using benchmarks and real-world applications. Dong Du 0003, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
ASPLOS | 1 |
| 2021 | Third-Eye: Practical and Context-Aware Inference of Causal Relationship Violations in Commodity Kernels
Chuhong Yuan, Dong Du 0003, Haibo Chen 0001 |
DIMVA | 2 |
| 2021 | Scalable Memory Protection in the PENGLAI Enclave
Erhu Feng, Dong Du 0003, Bicheng Yang, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
OSDI | 3 |
| 2021 | Boosting Inter-process Communication with Architectural SupportabstractIPC (inter-process communication) is a critical mechanism for modern OSes, including not only microkernels such as seL4, QNX, and Fuchsia where system functionalities are deployed in user-level processes, but also monolithic kernels like Android where apps frequently communicate with plenty of user-level services. However, existing IPC mechanisms still suffer from long latency. Previous software optimizations of IPC usually cannot bypass the kernel that is responsible for domain switching and message copying/remapping across different address spaces; hardware solutions such as tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt to the new hardware primitives. In this article, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for efficient and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel and supports secure message passing across multiple processes without copying. We have implemented a prototype of XPC based on the ARM AArch64 with Gem5 simulator and RISC-V architecture with FPGA boards. The evaluation shows that XPC can reduce IPC call latency from 664 to 21 cycles, 14×–123× improvement on Android Binder (ARM), and improve the performance of real-world applications on microkernels by 1.6× on Sqlite3. Yubin Xia, Dong Du 0003, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001, Haibing Guan |
ACM Trans. Comput. Syst. | 2 |
| 2020 | Catalyzer: Sub-millisecond Startup for Serverless Computing with Initialization-less BootingabstractServerless computing promises cost-efficiency and elasticity for high-productive software development. To achieve this, the serverless sandbox system must address two challenges: strong isolation between function instances, and low startup latency to ensure user experience. While strong isolation can be provided by virtualization-based sandboxes, the initialization of sandbox and application causes non-negligible startup overhead. Conventional sandbox systems fall short in low-latency startup due to their application-agnostic nature: they can only reduce the latency of sandbox initialization through hypervisor and guest kernel customization, which is inadequate and does not mitigate the majority of startup overhead. Dong Du 0003, Yubin Xia, Binyu Zang, Guanglu Yan, Chenggang Qin, Qixuan Wu, Haibo Chen 0001 |
ASPLOS | 1 |
| 2020 | Characterizing serverless platforms with serverlessbenchabstractServerless computing promises auto-scalability and cost-efficiency (in "pay-as-you-go" manner) for high-productive software development. Because of its virtue, serverless computing has motivated increasingly new applications and services in the cloud. This, however, also presents new challenges including how to efficiently design high-performance serverless platforms and how to efficiently program on the platforms. Dong Du 0003, Yubin Xia, Binyu Zang, Ziqian Lu, Pingchao Yang, Chenggang Qin, Haibo Chen 0001 |
SoCC | 3 |
| 2020 | Iso-UniK: lightweight multi-process unikernel through memory protection keysabstractAbstract Unikernel, specializing a minimalistic libOS with an application, is an attractive design for cloud computing. However, the Achilles’ heel of unikernel is the lack of multi-process support, which makes it less flexible and applicable. Many applications rely on the process abstraction to isolate different components. For example, Apache with the multi-processing module isolates a request handler in a process to guarantee security. Prior art tackles the problem by simulating multi-process with multiple unikernels, which is incompatible with existing cloud providers and also introduces high overhead. This paper proposes Iso-UniK, a new unikernel design enabling multi-task applications with the support of both functionality and isolation. Iso-UniK leverages a recent hardware feature, named Intel Memory Protection Key (Intel MPK), to provide lightweight and efficient isolation for multi-process in unikernel. Our design has three benefits compared with previous approaches. First, Iso-UniK does not need hypervisor support and is thus compatible with existing cloud computing platforms; second, Iso-UniK promises fast system calls with only 45 cycles; last, a process can be isolated with a flexible configuration. We have implemented a prototype based on OSv, a unikernel system supporting unmodified applications. Iso-UniK can achieve fast fork operation with only 66 μs for multi-process applications. Our evaluation shows that the isolation and multi-process support in Iso-UniK will not damage the applications’ performance. Dong Du 0003, Yubin Xia |
Cybersecur. | 2 |
| 2019 | XPC: architectural support for secure and efficient cross process callabstractMicrokernel has many intriguing features like security, fault-tolerance, modularity and customizability, which recently stimulate a resurgent interest in both academia and industry (including seL4, QNX and Google's Fuchsia OS). However, IPC (inter-process communication), which is known as the Achilles' Heel of microkernels, is still the major factor for the overall (poor) OS performance. Besides, IPC also plays a vital role in monolithic kernels like Android Linux, as mobile applications frequently communicate with plenty of user-level services through IPC. Previous software optimizations of IPC usually cannot bypass the kernel which is responsible for domain switching and message copying/remapping; hardware solutions like tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt the new hardware primitives. In this paper, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for fast and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel, and supports message passing across multiple processes through the invocation chain without copying. The primitive is compatible with the traditional address space based isolation mechanism and can be easily integrated into existing microkernels and monolithic kernels. We have implemented a prototype of XPC based on a Rocket RISC-V core with FPGA boards and ported two microkernel implementations, seL4 and Zircon, and one monolithic kernel implementation, Android Binder, for evaluation. We also implement XPC on GEM5 simulator to validate the generality. The result shows that XPC can reduce IPC call latency from 664 to 21 cycles, up to 54.2x improvement on Android Binder, and improve the performance of real-world applications on microkernels by 1.6x on Sqlite3 and 10x on an HTTP server with minimal hardware resource cost. Dong Du 0003, Zhichao Hua 0001, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
ISCA | 1 |
| 2019 | Using concurrent relational logic with helpers for verifying the AtomFS file systemabstractConcurrent file systems are pervasive but hard to correctly implement and formally verify due to nondeterministic interleavings. This paper presents AtomFS, the first formally-verified, fine-grained, concurrent file system, which provides linearizable interfaces to applications. The standard way to prove linearizability requires modeling linearization point of each operation---the moment when its effect becomes visible atomically to other threads. We observe that path inter-dependency, where one operation (like rename) breaks the path integrity of other operations, makes the linearization point external and thus poses a significant challenge to prove linearizability. Mo Zou, Dong Du 0003, Ming Fu, Ronghui Gu, Haibo Chen 0001 |
SOSP | 3 |
| 2018 | EPTI: Efficient Defence against Meltdown Attack for Unpatched VMs
Zhichao Hua 0001, Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 2 |
| 2018 | SplitPass: A Mutually Distrusting Two-Party Password Manager
Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang, Zhenkai Liang |
J. Comput. Sci. Technol. | 2 |