Yubin Xia

dblp:01/615 · DBLP profile ↗
← Back
89ranked-venue papers
7as first author
55since 2021 · last 2026
0000-0001-6558-5298ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 59 · 7 first-author · 38 since 2021Software engineering, systems software and programming languages · 22 · 18 since 2021Security and privacy · 9 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Computer networks · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 History Doesn't Repeat Itself but Rollouts Rhyme: Accelerating Reinforcement Learning with RhymeRL
abstract
With the rapid advancement of large language models (LLMs), reinforcement learning (RL) has emerged as a pivotal methodology for enhancing the reasoning capabilities of LLMs. Unlike traditional pre-training approaches, RL encompasses multiple stages: rollout, reward, and training, which necessitates collaboration among various worker types. However, current RL systems continue to grapple with substantial GPU underutilization, due to two primary factors: (1) The rollout stage dominates the overall RL process due to test-time scaling; (2) Imbalances in rollout lengths (within the same batch) result in GPU bubbles. While prior solutions like asynchronous execution and truncation offer partial relief, they may compromise training accuracy for efficiency. Our key insight stems from a previously overlooked observation: rollout responses exhibit remarkable similarity across adjacent training epochs. Based on the insight, we introduce RhymeRL, an LLM RL system designed to accelerate RL training with two key innovations. First, to enhance rollout generation, we present HistoSpec, a speculative decoding inference engine that utilizes the similarity of historical rollout token sequences to obtain accurate drafts. Second, to tackle rollout bubbles, we introduce HistoPipe, a two-tier scheduling strategy that leverages the similarity of historical rollout distributions to balance workload among rollout workers. Experimental results demonstrate that RhymeRL achieves up to a 2.6x performance improvement over existing methods, without compromising accuracy or modifying the RL paradigm.
Jingkai He, Tianjian Li, Erhu Feng, Dong Du 0003, Qian Liu 0033, Yubin Xia, Haibo Chen 0001
ASPLOS (2)7
2026 DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
abstract
Large multimodal models (LMMs) have demonstrated excellent capabilities in both understanding and generation tasks with various modalities. While these models can accept flexible combinations of input data, their training efficiency suffers from two major issues: pipeline stage imbalance caused by heterogeneous model architectures, and training data dynamicity stemming from the diversity of multimodal data.
Zhenliang Xue, Hanpeng Hu, Xing Chen 0009, Yixin Song 0003, Zeyu Mi, Yibo Zhu 0001, Daxin Jiang, Yubin Xia, Haibo Chen 0001
ASPLOS (2)9
2026 SKernel: An Elastic and Efficient Secure Container System at Scale with a Split-Kernel Architecture
abstract
Secure containers leverage hardware virtualization to isolate container sandboxes, enabling dedicated guest kernels to mitigate shared kernel attacks prevalent in traditional systems. However, existing approaches struggle with a fundamental trade-off: VM-based solutions (e.g., Kata) prioritize performance but lack elasticity and on-demand usage for volatile and bursty workloads, while lightweight methods (e.g., gVisor) rely on the host kernel for dynamic resource management at the cost of significant performance degradation due to guest-host dependencies.
Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Guotao Tan, Anqi Shen, Dawei Shen, Xinyao Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004
EuroSys15
2026 Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
Mo Zou, Hengbin Zhang, Dong Du 0003, Yubin Xia, Haibo Chen 0001
FAST5
2026 CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Yanning Yang, Dong Du 0003, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen 0001
ISCA8
2026 OneSidedMW: Managing Disaggregated Memory Efficiently, Flexibly, and Securely with RNIC Offloading
Jinyu Gu 0001, Xingda Wei, Yubin Xia
NSDI4
2026 Experiences with the ChCore Experimental Operating System Kernel
abstract
ChCore is an experimental operating system kernel built from the ground up for both education and research. Compared to other similar kernels, ChCore is unique in three aspects: 1) ARM-centric in light of a recent shift of computing from x86 to ARM; 2) microkernelbased given a resurgent interest in microkernel from industry; 3) bimodal for both education and research such that students can easily switch from learning to researching. This paper will describe our five-year experiences in teaching and experimenting operating systems with ChCore.
Haibo Chen 0001, Yubin Xia, Jinyu Gu 0001
SIGCSE (1)2
2026 LayerTEE: Decoupled Memory Protection for Scalable Multilayer Communication on RISC-V
abstract
The Trusted Execution Environment (TEE) has been widely implemented by modern hardware vendors to protect security and privacy-sensitive applications and data, such as Intel SGX/TDX, ARM TrustZone, AMD SEV, and RISC-V Penglai. However, existing TEE systems face challenges in balancing memory isolation among security, performance, and scalability requirements. Segment-based memory isolation mechanisms, like RISC-V PMP, struggle to scale effectively to the large number of segments needed for confidential cloud and data center environments. On the other hand, table-based isolation methods, such as page tables, combine address translation with memory protection, leading to inefficient cross-enclave communication and potential security vulnerabilities like Rowhammer attacks. This paper introduces a novel TEE system, which decouples memory protection (to segments) from address translation (to page tables). This design improves communication performance by dynamically adjusting memory protection capabilities, without sacrificing application compatibility. LayerTEE enhances enclave security and scalability by designing a multi-layer segment-based isolation mechanism. We have built a prototype of based on FPGA, incorporating hardware extensions and software support. The evaluation demonstrates that significantly surpasses existing TEE solutions, achieving three orders of magnitude lower communication latency and 10x greater scalability while maintaining robust security guarantees.
Shangjie Pan, Yinghao Yang 0001, Xuanyao Peng, Xiquan Zhao, Dong Du 0003, Yubin Xia, Xiaowei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2026 EIDS: A Cloud Intrusion Detection System with High Performance and Maintainability
abstract
Intrusion Detection Systems (IDSes) are widely employed to identify potential attacks in guest virtual machines (VMs). Nonetheless, traditional IDSes fall short of the demands of high-performance clouds. First, monitoring VM events increases the tail latency of guest services. Second, the throughput of traditional IDSes cannot meet high-performance cloud requirements, leading to event loss and reduced detection accuracy. Finally, cloud providers typically run complex IDS tools within the VM. Updating IDS functionality requires modifying guest VMs, which hurts maintainability. To overcome these challenges, this article presents EIDS, a cloud IDS framework with high performance and good maintainability. We observe that the main bottleneck is collecting VM status, and the collected status can be divided into fundamental and supplementary status. EIDS then splits the status collection procedure spatially and temporally. First, we provide a status monitor with a separate architecture that isolates the status collection logic in a microVM, thus minimizing the code in guest VMs and improving maintainability. Second, EIDS introduces a two-phase status collection method to handle multiple events in batches, asynchronously, for high IDS throughput. A tiny tracer, implemented with eBPF, operates inside the user VM to collect fundamental status. The complex status collector runs in an isolated microVM. It utilizes Virtual Machine Introspection (VMI) to gather supplementary status, using the fundamental status to bridge the semantic gap. The status collector batches the collection for multiple events to amortize the fixed overhead of microVM switching and improve event tracing throughput. Finally, to minimize tail latency overhead, a fine-grained and workload-aware scheduler executes IDS logic with small time slices during user VM idle periods. We implemented a prototype of EIDS in Linux-KVM and conducted a comprehensive evaluation. We compared EIDS’s performance with Falco, an open-source IDS widely used by Kubernetes and AWS for runtime security monitoring. The results demonstrate that, compared to Falco, EIDS reduces the 99 th -percentile latency overhead by 97% and achieves a 13.8X improvement in IDS event handling throughput.
Xiaokang Hu, Zhichao Hua 0001, Naixuan Guan, Yibin Shen, Yang Yu 0002, Zeyu Mi, Yubin Xia, Jiesheng Wu
ACM Trans. Comput. Syst.8
2026 Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and Practicability
abstract
Current serverless platforms struggle to optimize resource utilization for both CPU and GPU functions due to their dynamic and fine-grained nature. Conventional techniques like overcommitment and autoscaling fall short, often sacrificing utilization for practicability or incurring performance tradeoffs. Overcommitment requires predicting performance to prevent QoS violation, introducing tradeoff between prediction accuracy and overheads. Autoscaling requires scaling instances in response to load fluctuations quickly to reduce resource wastage, but more frequent scaling also leads to more cold start overheads. The rich concurrency of GPU resources further complicates GPU instance orchestration, such as setting right batch sizes. This article introduces Jiagu to harmonize efficiency with practicability through the following novel techniques. First, pre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and scheduling. Second, dual-staged scaling achieves frequent adjustment of instances with minimum overhead. Third, Jiagu conducts an in-depth analysis about the complexity of the relationship between GPU function configuration and execution. It then proposes batch-aware scaling that achieves optimal configurations for both batch size setting and autoscaling, addressing all the challenges according to the analysis. We have implemented a prototype and evaluated it using real-world applications and traces from the public cloud platform. Our evaluation shows an improvement in deployment density over commercial clouds (with Kubernetes) while maintaining QoS for both CPU and GPU functions (54.8% and 18% respectively), and 81.0%–93.7% lower scheduling costs and a 57.4%–69.3% reduction in cold start latency compared to existing QoS-aware schedulers.
Yanning Yang, Dong Du 0003, Yubin Xia, Haibo Chen 0001
ACM Trans. Comput. Syst.4
2025 Dep-TEE: Decoupled Memory Protection for Secure and Scalable Inter-enclave Communication on RISC-V
abstract
Trusted Execution Environment (TEE) has been widely implemented by modern hardware vendors to protect security and privacy-sensitive applications and data, such as Intel SGX/TDX, ARM TrustZone, AMD SEV, and RISC-V Penglai. However, existing TEE systems face challenges in balancing memory isolation among security, performance, and scalability requirements. This paper introduces a novel TEE system, Dep-TEE, which decouples memory protection (to segments) from address translation (to page tables). This design improves communication performance by dynamically adjusting memory protection capabilities, without sacrificing application compatibility, and enhances security by safeguarding against attacks on page tables. We have built a prototype of Dep-TEE based on FPGA, incorporating hardware extensions and software support. The evaluation demonstrates that Dep-TEE significantly surpasses existing TEE solutions, achieving three orders of magnitude lower communication latency and 10x greater scalability while maintaining robust security guarantees.
Shangjie Pan, Xuanyao Peng, Zeyuan Man, Xiquan Zhao, Dongrong Zhang, Bicheng Yang, Dong Du 0003, Yubin Xia, Xiaowei Li 0001
ASP-DAC9
2025 D-VSync: Decoupled Rendering and Displaying for Smartphone Graphics
abstract
Rendering service, which typically orchestrates screen display and UI through Vertical Synchronization (VSync), is an indispensable system service for user experiences of smartphone OSes (e.g., Android, OpenHarmony, and iOS). The recent trend of large high-frame-rate screens, stunning visual effects, and physics-based animations has placed unprecedented pressure on the VSync-based rendering architecture, leading to higher frame drops and longer rendering latency.
Yuanpei Wu, Dong Du 0003, Yubin Xia, Ming Fu, Binyu Zang, Haibo Chen 0001
ASPLOS (1)4
2025 Offloading Cloud-Native Infrastructure with XpuPod
abstract
Cloud-native systems increasingly rely on infrastructure services (e.g., service meshes, monitoring agents), which compete for resources with user applications, thereby degrading performance and scalability. We propose XpuPod, a new abstraction that offloads these services to Data Processing Units (DPUs) to enforce strict isolation while reducing host resource contention and operational costs. XpuPod enables a cross-processing-unit (XPU) network system. This system features two key components: (1) transparent XPU networking, which provides a unified network abstraction for processes spanning both the CPU and DPU, and (2) elastic and efficient XPU communication, a mechanism that achieves shared-memory performance without the costs of pinned resources. By leveraging the XPU network system and the compositional nature of cloud-native workloads, XpuPod can optimally offload infrastructure containers to DPUs. We implement XpuPod on Linux and a corresponding cloud-native system, XpuK8s, on Kubernetes. We evaluate our system using NVIDIA BlueField-2 DPUs and a simulator backed by a CXL memory device. The results show that XpuK8s effectively supports complex, unmodified, commodity cloud-native applications. Compared to a kernel-bypass design, XpuK8s provides up to 31.9x lower latency and consumes 64x fewer resources. Furthermore, compared to state-of-the-art systems, XpuK8s can achieve 60% lower end-to-end latency and 55% higher scalability.
Bicheng Yang, Jingkai He, Dong Du 0003, Yubin Xia, Haibo Chen 0001
SoCC4
2025 A Hardware-Software Co-Design for Efficient Secure Containers
abstract
VM-level containers provide strong isolation by running each container with its own kernel in a VM. However, they rely on virtualization hardware designed for general-purpose VMs, causing non-negligible performance overhead compared to OS-level containers. This performance gap widens dramatically in nested virtualization scenarios, where secure containers run inside a VM.
Jiacheng Shi 0002, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia
EuroSys4
2025 Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
abstract
With the rapid development of artificial intelligence (AI) applications, an emerging class of AI accelerators, termed Inter-core Connected Neural Processing Units (NPU), has been adopted in both cloud and edge computing environments, like Graphcore IPU, Tenstorrent, etc.Despite their innovative design, these NPUs often demand substantial hardware resources, leading to suboptimal resource utilization due to the imbalance of hardware requirements across various tasks.To address this issue, prior research has explored virtualization techniques for monolithic NPUs, but has neglected inter-core connected NPUs with the hardware topology.This paper introduces vNPU, the first comprehensive virtualization design for inter-core connected NPUs, integrating three novel techniques: (1) NPU route virtualization, which redirects instruction and data flow from virtual NPU cores to physical ones, creating a virtual topology; (2) NPU memory virtualization, designed to minimize translation stalls for SRAM-centric and NoC-equipped NPU cores, thereby maximizing the memory bandwidth; and (3) Besteffort topology mapping, which determines the optimal mapping from all candidate virtual topologies, balancing resource utilization with end-to-end performance.We have developed a prototype of vNPU on both an FPGA platform (Chipyard+FireSim) and a simulator (DCRA).Evaluation results demonstrate that when executing multiple NPU workloads on virtual NPUs, vNPU achieves performance improvements of up to 1.92x and 1.28x for the Transformer and ResNet models, respectively, in comparison to the MIG-based virtualization method.Furthermore, the hardware performance * Both authors contributed equally to this research.
Dahu Feng, Erhu Feng, Dong Du 0003, Pinjie Xu, Yubin Xia, Haibo Chen 0001
ISCA5
2025 Sprig: Low-Latency Startup for Knative Serverless Platforms with Async-Cfork
abstract
Serverless computing has emerged as a transformative paradigm in the cloud-native era, widely adopted across major public cloud platforms. However, its pay-as-use model necessitates frequent dynamic instance creation and destruction, making cold start latency a critical system bottleneck. While the industry predominantly employs caching techniques to mitigate cold start overheads, these approaches exhibit inherent limitations. Existing research prototypes often optimize container runtimes manually, resulting in systems incompatible with production environments that cannot leverage automated management capabilities or integrate optimizations effectively. Moreover, runtime optimizations fail to fit system features like asynchronous cold start. This paper presents Sprig, the first serverless system that seamlessly integrates cutting-edge academic research with production-grade platforms. We introduce a novel computational abstraction called Template Pod to bridge the gap between research-oriented flexible designs and minimal computational units in production systems. Sprig further incorporates Asynccfork, a low-latency startup technique optimized for large-scale serverless platforms, significantly reducing function cold start latency. Additionally, we design a multiplexed proxy queue with enhanced scheduling elasticity to resolve request accumulation issues in Knative's asynchronous startup mechanism. Evaluations show Sprig achieves a 70% average reduction in cold start latency compared with Knative, and decreases P99 end-to-end latency from 27s to 0.5s under high load while maintaining only 16.7% of baseline memory footprint for 20 instances.
Dong Du 0003, Liang Zhang 0010, Yubin Xia
JCC4
2025 Fork in the Road: Reflections and Optimizations for Cold Start Latency in Production Serverless Systems
Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Anqi Shen, Dawei Shen, Qi Xing, Shun Song, Tongkai Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004
OSDI15
2025 OS Rendering Service Made Parallel with Out-of-Order Execution and In-Order Commit
Yuanpei Wu, Yubin Xia, Yang Yu 0002, Ming Fu, Binyu Zang, Haibo Chen 0001
OSDI3
2025 Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
abstract
With the rapid advancement of artificial intelligence technologies such as ChatGPT, AI agents, and video generation, contemporary mobile systems have begun integrating these AI capabilities on local devices to enhance privacy and reduce response latency. To meet the computational demands of AI tasks, current mobile SoCs are equipped with diverse AI accelerators, including GPUs and Neural Processing Units (NPUs). However, there has not been a comprehensive characterization of these heterogeneous processors, and existing designs typically only leverage a single AI accelerator for LLM inference, leading to suboptimal use of computational resources and memory bandwidth.
Dahu Feng, Erhu Feng, Yingrui Wang, Yubin Xia, Pinjie Xu, Haibo Chen 0001
SOSP6
2025 How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS Service
abstract
In modern systems, memory copy remains a critical performance bottleneck across various scenarios, playing a pervasive role in system-wide execution such as syscalls, IPC, and user-mode applications. Numerous efforts have aimed at optimizing copy performance, including zero-copy with page remapping and hardware-accelerated copy. However, they typically target specific use cases, such as Linux zero-copy send() for messages of ≥10KB. This paper argues for copy as a first-class OS service, offering three key benefits: (1) with the asynchronous copy abstraction provided by the service, applications can overlap their execution with copy; (2) the service can effectively utilize hardware capabilities to enhance copy performance; (3) the service's global view of copies further enables holistic optimization. To this end, we introduce Copier, a new OS service of coordinated asynchronous copy, to serve both user-mode applications and OS services. We build Copier-Linux to demonstrate Copier's ability to improve performance for diverse use cases, including Redis, Protobuf, network stack, proxy, etc. Evaluations show that Copier achieves up to a 1.8 × speedup for real-world applications like Redis and a 1.6 × improvement over zIO, the state-of-the-art in optimizing copy efficiency. To further facilitate adoption, we develop a toolchain to ease the use of Copier. We also integrate Copier into a commercial smartphone OS (HarmonyOS 5.0), achieving promising results.
Jingkai He, Yunpeng Dong, Dong Du 0003, Mo Zou, Zhitai Yu, Yuxin Ren 0001, Ning Jia 0004, Yubin Xia, Haibo Chen 0001
SOSP8
2025 μEFI: A Microkernel-Style UEFI with Isolation and Transparency
Yiyang Wu, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001
USENIX ATC4
2025 Serverless Functions Made Confidential and Efficient with Split Containers
Jiacheng Shi 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001
USENIX Security Symposium3
2025 Enhancing Knowledge Tracing through Decoupling Cognitive Pattern from Error-Prone Data
abstract
Knowledge tracing (KT) aims to predict students' future performance based on their past learning activities. However, no one is perfect. Factors such as carelessness, fatigue, and stress often cause students to make mistakes on problems they have already mastered, leading to anomalies in their historical learning data. These anomalies disrupt inherent patterns in the data, misleading the KT model. Extracting cognitive patterns that accurately reflect students' knowledge mastery from such error-prone data remains a significant challenge. Against this background, this paper proposes a novel KT method named RoubstKT, inspired by educational measurement theory and frequency-based decomposition. A cognitive decoupling analyzer is proposed to decouple the student's cognitive pattern and random factors from the data through smoothing and subtraction operations, then recombine them using a gating mechanism or adaptive parameter fusion strategy. To more effectively diagnose students' knowledge mastery, we employ a decay-based attention mechanism that focuses on random behaviors at adjacent time steps. We conducted comprehensive experiments based on real-world datasets and targeted datasets with added random noise. The experimental results demonstrated the effectiveness of the proposed method.
Teng Guo 0002, Yubin Xia, Mingliang Hou, Zitao Liu 0001, Feng Xia 0001, Weiqi Luo 0002
WWW3
2025 Harmonizing Security and Performance in Microkernel File Servers
Wentai Li, Jinyu Gu 0001, Yubin Xia, Binyu Zang
J. Comput. Sci. Technol.4
2025 XpuTEE: A High-Performance and Practical Heterogeneous Trusted Execution Environment for GPUs
abstract
AI applications are employed in diverse scenarios, including data centers, personal computers, smart cars, and so on. Their privacy is threatened by the intricate software stacks and the potential malfeasance of system maintainers. The Trusted Execution Environment (TEE) has become popular for safeguarding applications from untrusted system software. However, AI applications are always speeded up with heterogeneous accelerators, e.g., GPU, which requires the TEE to be heterogeneous. A heterogeneous TEE should satisfy three requirements: (1) the joint heterogeneous abstraction that covers the CPU and GPUs and minimizes cooperation overhead among enclaves on them; (2) the high performance for supporting high-speed GPUs and introducing limited performance overhead; and (3) the compatibility with existing CPUs and GPUs so that existing machines can directly benefit from it. To meet the above requirements, this article introduces XpuTEE, a practical and high-performance heterogeneous TEE system. XpuTEE provides a new abstraction called XpuEnclave, comprising the CEnclave to protect CPU-side logic and numerous XEnclaves to guard GPU tasks. XpuEnclave is a joint TEE crossing the CPU and connected GPUs, and it removes all cryptographic operations and extra memory copies for CPU-GPU communication, which allows XpuTEE to achieve high performance. The results demonstrate that XpuTEE has an average performance overhead of 2.48% for common AI applications.
Shulin Fan, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001
ACM Trans. Comput. Syst.3
2024 sIOPMP: Scalable and Efficient I/O Protection for TEEs
abstract
Trusted Execution Environments (TEEs), like Intel SGX/TDX, AMD SEV-SNP, ARM TrustZone/CCA, have been widely adopted in prevailing architectures. However, these TEEs typically do not consider I/O isolation (e.g., defending against malicious DMA requests) as a first-class citizen, which may degrade the I/O performance. Traditional methods like using IOMMU or software I/O can degrade throughput by at least 20% for I/O intensive workloads. The main reason is that the isolation requirements for I/O devices differ from CPU ones. This paper proposes a novel I/O isolation mechanism for TEEs, named sIOPMP (scalable I/O Physical Memory Protection), with three key features. First, we design a Multi-stage-Tree-based checker, supporting more than 1,000 hardware regions. Second, we classify the devices into hot and cold, and support unlimited devices with the mountable entry. Third, we propose a remapping mechanism to switch devices between hot and cold status for dynamic I/O workloads. Evaluation results show that sIOPMP introduces only negligible performance overhead for both benchmarks and real-world workloads, and improves 20% ~ 38% network throughput compared with IOMMU-based mechanisms or software I/O adopted in TEEs.
Erhu Feng, Dahu Feng, Dong Du 0003, Yubin Xia, Siqi Zhao, Haibo Chen 0001
ASPLOS (2)4
2024 On-demand and Parallel Checkpoint/Restore for GPU Applications
abstract
Leveraging serverless computing for cloud-based machine learning services is on the rise, promising cost-efficiency and flexibility are crucial for ML applications relying on high-performance GPUs and substantial memory. However, despite modern serverless platforms handling diverse devices like GPUs seamlessly on a pay-as-you-go basis, a longstanding challenge remains: startup latency, a well-studied issue when serverless is CPU-centric. For example, initializing GPU apps with minor GPU models, like MobileNet, demands several seconds. For more intricate models such as GPT-2, startup latency can escalate to around 10 seconds, vastly overshadowing the short computation time for GPU-based inference. Prior solutions tailored for CPU serverless setups, like fork() and Checkpoint/Restore, cannot be directly and effectively applied due to differences between CPUs and GPUs.
Yanning Yang, Dong Du 0003, Haitao Song 0001, Yubin Xia
SoCC4
2024 sNPU: Trusted Execution Environments on Integrated NPUs
abstract
Trusted execution environment (TEE) promises strong security guarantee with hardware extensions for security-sensitive tasks. Due to its numerous benefits, TEE has gained widespread adoption, and extended from CPU-only TEEs to FPGA and GPU TEE systems. However, existing TEE systems exhibit inadequate and inefficient support for an emerging (and significant) processing unit, NPU. For instance, commercial TEE systems resort to coarse-grained and static protection approaches for NPUs, resulting in notable performance degradation (10%–20%), limited (or no) multitasking capabilities, and suboptimal resource utilization. In this paper, we present a secure NPU architecture, known as sNPU, which aims to mitigate vulnerabilities inherent to the design of NPU architectures. First, sNPU proposes NPU Guarder to enhance the NPU’s access control. Second, sNPU defines new attack surfaces leveraging in-NPU structures like scratchpad and NoC, and designs NPU Isolator to guarantee the isolation of scratchpad and NoC routing. Third, our system introduces a trusted software module called NPU Monitor to minimize the software TCB. Our prototype, evaluated on FPGA, demonstrates that sNPU significantly mitigates the runtime costs associated with security checking (from upto 20% to 0%) while incurring less than 1% resource costs.
Erhu Feng, Dahu Feng, Dong Du 0003, Yubin Xia, Haibo Chen 0001
ISCA4
2024 The Design and Optimization of Memory Ballooning in SEV Confidential Virtual Machines
abstract
With the popularity of confidential computing, confidential virtual machines (CVMs) have been widely adopted and they guarantee strong security by hardware. However, there still exist some problems in memory management in CVMs. Since private memory pages of CVMs are encrypted and cannot be accessed by hypervisors, existing CVMs employ static page management to avoid crashes due to the relocation of encrypted memory pages, leading cloud platforms managing CVMs to face more severe memory management pressures than before. Memory ballooning, as an efficient, flexible, and highly compatible memory management mechanism in virtualization, is not available in CVMs based on SEV (Secure Encrypted Virtualization). In this paper, we analyze the design of SEV CVMs and memory ballooning, and enable memory ballooning on SEV CVMs by substituting static page management with dynamic page management and addressing communication issues between the guest frontend and the host backend. Besides, we propose three performance optimization strategies for memory ballooning on SEV CVMs, including asynchronous reclaiming on the host side, an additional shadow vCPU on the guest side, and accelerating cache flushing operations in the host kernel. Experiments show that the time cost of reclaiming memory from SEV CVMs by memory ballooning can be reduced by up to 38 times and up to 55% overhead caused by delaying reclamation in real-world applications like MYSQL can be eliminated.
Chang Deng, Zheyun Shen, Dingji Li, Zeyu Mi, Yubin Xia
JCC5
2024 PrometheusMigrate: Efficient Live Migration of Confidential Virtual Machine with Software Abstraction
abstract
With the rise of cloud computing, the use of virtual machines in data centers has become genuinely common. To achieve load balance and disaster recovery between different hosts when a virtual machine is providing services, the concept of live migration has been proposed and has now become one of the essential capabilities in data centers. On the other hand, users’ concerns about data privacy when applications are deployed on cloud vendors’ servers are also rising, and for this reason, confidential virtual machines, a new type of virtualization hardware, have been proposed. However, since the memory of confidential VMs is usually encrypted, existing live migration solutions cannot be applied to confidential VMs directly, so hardware vendors have come up with various solutions to support this critical feature. For example, the mainstream confidential virtualization hardware platform – AMD-SEV, adds a series of interfaces in their secure firmware to assist the hypervisor to achieve live migration capability. However, this solution can cause the live migration time to be too long due to the limited number of secure processors and computing power. In this paper, we propose a new software design: PROMETHEUSMIGRATE. PROMETHEUSMIGRATE presents a novel software abstraction that bypasses SEV’s secure firmware with guaranteed security, thereby dramatically accelerating the total live migration time as well as downtime of the AMD-SEV platform.
Chenhui Ji, Dingji Li, Zeyu Mi, Yubin Xia
JCC4
2024 CPC: Flexible, Secure, and Efficient CVM Maintenance with Confidential Procedure Calls
Zeyu Mi, Yubin Xia, Haibing Guan, Haibo Chen 0001
USENIX ATC3
2024 Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with Jiagu
Yanning Yang, Dong Du 0003, Yubin Xia, James R. Larus, Haibo Chen 0001
USENIX ATC4
2023 The Gap Between Serverless Research and Real-world Systems
abstract
With the emergence of the serverless computing paradigm in the cloud, researchers have explored many challenges of serverless systems and proposed solutions such as snapshot-based booting. However, we have noticed that some of these optimizations are based on oversimplified assumptions that lead to infeasibility and hide real-world issues. This paper aims to analyze the gap between current serverless research and real-world systems from a perspective of industry, and present new observations, challenges, opportunities, and insights that may address the discrepancies.
Dong Du 0003, Yubin Xia, Haibo Chen 0001
SoCC3
2023 Efficient Distributed Secure Memory with Migratable Merkle Tree
abstract
Hardware-assisted enclaves with memory encryption have been widely adopted in the prevailing architectures, e.g., Intel SGX/TDX, AMD SEV, ARM CCA, etc. However, existing enclave designs fall short in supporting efficient cooperation among cross-node enclaves (i.e., multi-machines) because the range of hardware memory protection is within a single node. A naive approach is to leverage cryptography at the application level and transfer data between nodes through secure channels (e.g., SSL). However, it incurs orders of magnitude costs due to expensive encryption/decryption, especially for distributed applications with large data transfer, e.g., MapReduce and graph computing. A secure and efficient mechanism of distributed secure memory is necessary but still missing.This paper proposes Migratable Merkle Tree (MMT), a design enabling efficient distributed secure memory to support distributed confidential computing. MMT sets up an integrity forest for distributed memory on multiple nodes. It allows an enclave to securely delegate an MMT closure, which contains both data and metadata of a subtree, to a remote enclave. By reusing the memory encryption mechanisms of existing enclaves, our design achieves secure data transfer without software re-encryption. We have implemented a prototype of MMT and a trusted firmware for management, and further applied MMT to real-world distributed applications. The evaluation results show that compared with existing systems using the AES-NI instruction, MMT can achieve up to 13x speedup on data transferring, and gain 12%~58% improvement on the end-to-end performance of MapReduce and PageRank.
Erhu Feng, Dong Du 0003, Yubin Xia, Haibo Chen 0001
HPCA3
2023 SecDFS: A Secure and Decentralized File System
abstract
Consortium networks are usually constructed among enterprises and organizations. A consortium network is a specific type of decentralized network which only has a limited number of untrusted nodes. Consortium systems are designed to achieve consensus in such a network and provide stable and trusted services. Meanwhile, secure data sharing plays a crucial role in the consortium network, introducing new requirements to the underlying file system. Firstly, a consortium file system should be decentralized since any party of the consortium network is untrusted. Secondly, it should enforce data confidentiality and integrity and provide access control. Finally, the system should offer high performance. Existing file systems are designed for public decentralized or centralized networks and cannot meet all the above requirements. To address this problem, we present SecDFS, a decentralized, secure, high-performance file system for consortium networks. SecDFS first combines trusted execution environment and consistent hashing to design a consistent consensus protocol, through which SecDFS distributes files across consortium nodes and achieves data consensus. After that, SecDFS introduces a secure storage constructing method to protect the confidentiality and integrity of both the data and metadata, as well as a two-level cache to speed up the file system performance. We have implemented a prototype of SecDFS and performed a detailed evaluation with it. The results show that SecDFS has 74.83X latency speedup compared with the InterPlanetary File System on average.
Shenglong Zhao, Zhichao Hua 0001, Yubin Xia
ICPADS3
2023 ISA-Grid: Architecture of Fine-grained Privilege Control for Instructions and Registers
abstract
Isolation is a critical mechanism for enhancing the security of computer systems. By controlling the access privileges of software and hardware resources, isolation mechanisms can decouple software into multiple isolated components and enforce the principle of least privilege. While existing isolation systems primarily focus on memory isolation, they overlook the isolation of instruction and register resources, which we refer to as ISA (Instruction Set Architecture) resources. However, previous works have shown that exploiting ISA resources can lead to serious security problems, such as breaking the system's memory isolation property by abusing x86's CR3 register. Furthermore, existing hardware only provides privilege-level-based access control for ISA resources, which is too coarse-grained for software decoupling. For example, ARM Cortex A53 has several hundred system instructions/registers, but only four exception levels (EL0 to EL3) are provided. Additionally, more than 100 instructions/registers for system control are available in only EL1 (the kernel mode). To address this problem, this paper proposes ISA-Grid, an architecture of fine-grained privilege control for instructions and registers. ISA-Grid is a hardware extension that enables the creation of multiple ISA domains, with each domain having different privileges to access instructions and registers. The ISA domain can provide bit-level fine-grained privilege control for registers. We implemented prototypes of ISA-Grid based on two different CPU cores: 1) a RISC-V CPU core on an FPGA board and 2) an x86 CPU core on a simulator. We applied ISA-Grid to different cases, including Linux kernel decomposition and enhancing existing security systems, to demonstrate how ISA-Grid can isolate ISA resources and mitigate attacks based on abusing them. The performance evaluation results on both x86 and RISC-V platforms with real-world applications showed that ISA-Grid has negligible runtime overhead (less than 1%).
Shulin Fan, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang
ISCA3
2023 Secure and Efficient Runtime Environment for Smart Contracts on JointCloud
abstract
Many cloud providers, including Amazon, Google, Microsoft, and Alibaba Cloud, offer support for blockchain cloud services that rely on a runtime environment, such as the Ethereum Virtual Machine (EVM), to execute smart contracts and ensure consistency between participants. However, existing runtime systems suffer from two main limitations. Firstly, traditional runtime systems like EVM cannot guarantee privacy protection as all the data uploaded to the blockchain is visible to all participants. This restricts the use of blockchain in limited scenarios. Secondly, each computation on the runtime system must be synchronized to all nodes in the network, resulting in a significant increase in computational overhead, which can be challenging to implement for more complex applications. One approach to address these limitations is to utilize Trusted Execution Environments (TEE) for blockchain runtime, which can provide privacy protection and mitigate redundant synchronization operations. However, using TEE for blockchain may significantly increase cloud costs. To overcome these challenges, this paper proposes PL-EVM, a new runtime environment for smart contracts that utilizes jointcloud. PL-EVM achieves high-security guarantees by using TEE to protect privacy-sensitive data and incorporates dynamic migration and splitting mechanisms to achieve high efficiency and low costs. Our evaluation results show that PL-EVM can improve performance and reduce costs by 4% to 32.22%.
Yuhao Xue, Dong Du 0003, Yubin Xia
JCC4
2023 Accelerating Extra Dimensional Page Walks for Confidential Computing
abstract
To support highly scalable and fine-grained computing paradigms such as microservices and serverless computing better, modern hardware-assisted confidential computing systems, such as Intel TDX and ARM CCA, introduce permission table to achieve fine-grained and scalable memory isolation among different domains. However, it also adds an extra dimension to page walks besides page tables, leading to significantly more memory references (e.g., 4 → 12 for RISC-V Sv39)1. We observe that most costs (about 75%) caused by the extra dimension of page walks are used to validate page table pages. Based on this observation, this paper proposes HPMP (Hybrid Physical Memory Protection), a hardware-software co-design (on RISC-V) that protects page table pages using segment registers and normal pages using permission tables to balance scalability and performance. We have implemented HPMP and Penglai-HPMP (a TEE system based on HPMP) on FPGA with two RISC-V cores (both in-order and out-of-order). Evaluation results show that HPMP can reduce costs by 23.1%–73.1% on BOOM and significantly improve performance on real-world applications, including serverless computing (FunctionBench) and Redis.
Dong Du 0003, Bicheng Yang, Yubin Xia, Haibo Chen 0001
MICRO3
2023 Encrypted Databases Made Secure Yet Maintainable
Cheng Tan 0005, Huorong Li, Sheng Wang 0011, Zeyu Mi, Yubin Xia, Feifei Li 0001, Haibo Chen 0001
OSDI8
2023 Temporal Graph Cube
abstract
Data warehouse and OLAP (Online Analytical Processing) are effective tools for decision support on traditional relational data and static multidimensional network data. However, many real-world multidimensional networks are often modeled as temporal multidimensional networks, where the edges in the network are associated with temporal information. Such temporal multidimensional networks typically cannot be handled by traditional data warehouse and OLAP techniques. To fill this gap, we propose a novel data warehouse model, named$\mathsf {Temporal{ }\; Graph{ }\; Cube}$, to support OLAP queries on temporal multidimensional networks. Through supporting OLAP queries in any time range, users can obtain summarized information of the network in the time range of interest, which cannot be derived by using traditional static graph OLAP techniques. We propose a segment-tree based indexing technique to speed up the OLAP queries, and also develop an index-updating technique to maintain the index when the temporal multidimensional network evolves over time. In addition, we also propose a novel concept called$\mathsf {similarity{ }\; of{ }\; snapshots}$which shows a strong correlation with the efficiency of indexing technique and can provide a good reference on the necessity of building the index. The results of extensive experiments on two large real-world datasets demonstrate the effectiveness and efficiency of the proposed method.
Guoren Wang, Yue Zeng 0004, Rong-Hua Li 0001, Hongchao Qin, Xuanhua Shi, Yubin Xia, Xuequn Shang 0001, Liang Hong 0001
IEEE Trans. Knowl. Data Eng.6
2022 Serverless computing on heterogeneous computers
abstract
Existing serverless computing platforms are built upon homogeneous computers, limiting the function density and restricting serverless computing to limited scenarios. We introduce Molecule, the first serverless computing system utilizing heterogeneous computers. Molecule enables both general-purpose devices (e.g., Nvidia DPU) and domain-specific accelerators (e.g., FPGA and GPU) for serverless applications that significantly improve function density (50% higher) and application performance (up to 34.6x). To achieve these results, we first propose XPU-Shim, a distributed shim to bridge the gap between underlying multi-OS systems (when using general-purpose devices) and our serverless runtime (i.e., Molecule). We further introduce vectorized sandbox, a sandbox abstraction to abstract hardware heterogeneity (when using domain-specific accelerators). Moreover, we also review state-of-the-art serverless optimizations on startup and communication latency and overcome the challenges to implement them on heterogeneous computers. We have implemented Molecule on real platforms with Nvidia DPUs and Xilinx FPGAs and evaluate it using benchmarks and real-world applications.
Dong Du 0003, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001
ASPLOS4
2022 Towards A Secure Joint Cloud With Confidential Computing
abstract
As data security in public clouds attracts more attention and concerns, researchers and practitioners have proposed techniques to secure cloud computing. Confidential computing (CC) is a compelling approach that guarantees both privacy and integrity of data and code in public clouds. In this paper, we first survey the status of CC in today’s commercialized public clouds, including the cloud CC abstractions, infrastructures, metrics, third-party service vendors, and real-world cloud use cases. We also discover the limitations such as re-programming efforts, extra cost, limited availability, etc. We further take a step forward to prospect CC in the joint cloud scenario. We finally showcase the challenges of realizing a secure joint cloud and propose possible solutions.
Erhu Feng, Yubin Xia
JCC4
2022 EPK: Scalable and Efficient Memory Protection Keys
Jinyu Gu 0001, Wentai Li, Yubin Xia, Haibo Chen 0001
USENIX ATC4
2022 A Hardware-Software Co-design for Efficient Intra-Enclave Isolation
Jinyu Gu 0001, Bojun Zhu, Wentai Li, Yubin Xia, Haibo Chen 0001
USENIX Security Symposium5
2022 Unified Enclave Abstraction and Secure Enclave Migration on Heterogeneous Security Architectures
Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Chenggang Qin, Zhengyu He
J. Comput. Sci. Technol.3
2022 Colony: A Privileged Trusted Execution Environment With Extensibility
abstract
The code base of system software is growing fast, which results in a large number of vulnerabilities: for example, 296 CVEs have been found in Xen hypervisor and 2195 CVEs in Linux kernel. To reduce the reliance on the trust of system software, many researchers try to provide trusted execution environments (TEEs), which can be categorized into two types: non-privileged TEEs and privileged TEEs. Non-privileged TEEs (e.g., Intel SGX) are extensible, but cannot protect security services like virtual machine introspection (VMI) due to the lack of system-level semantics. On the contrary, privileged TEEs (e.g., the secure world of ARM TrustZone) have system-level semantics, but any additional service implemented in the privileged TEE directly increases the TCB of the entire system. In this article, we propose a new design of TEE to support system-level security services and achieve better extensibility with a small TCB. Each TEE instance of the proposed design is named aColony. Specifically, we introduce asecure monitorfor isolation and capability management. EachColonyis assigned capabilities to access only necessary system-level semantics. We use the new TEE to build four security services, including secure device accessing, VMI tools, a system call tracer, and a much more complex service to virtualize ARM TrustZone with multipleColonies. We have implemented the system on ARMv7 and ARMv8 platforms, in Xen hypervisor and Linux kernel, and perform a detailed evaluation to show its efficiency.11.This paper is an extended version of the conference paper published in USENIX Security’17: vTZ: Virtualizing ARM TrustZone[29]. A brief summary of differences is in Section8.
Yubin Xia, Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan
IEEE Trans. Computers1
2021 Fast and Accurate Optimizer for Query Processing over Knowledge Graphs
abstract
This paper presents Gpl, a fast and accurate optimizer for query processing over knowledge graphs. Gpl is novel in three ways. First, Gpl proposes a type-centric approach to enhance the accuracy of cardinality estimation prominently, which naturally embeds the correlation of multiple query conditions into the existing type system of knowledge graphs. Second, to predict execution time accurately, Gpl constructs a specialized cost model for graph exploration scheme and tunes the coefficients with target hardware platform and graph data. Third, Gpl further uses a budget-aware strategy for plan enumeration with a greedy heuristic to boost the overall performance (i.e., optimization time and execution time) for various workloads. Evaluations with representative knowledge graphs and query benchmarks show that Gpl can select optimal plans for 33 of 39 queries and only incurs less than 5% slowdown on average compared to optimal results. In contrast, the state-of-the-art optimizer and manually tuned results will cause 100% and 36% slowdown, respectively.
Jingqi Wu, Rong Chen 0001, Yubin Xia
SoCC3
2021 Confidential Serverless Made Efficient with Plug-In Enclaves
abstract
Serverless computing has become a fact of life on modern clouds. A serverless function may process sensitive data from clients. Protecting such a function against untrusted clouds using hardware enclave is attractive for user privacy. In this work, we run existing serverless applications in SGX enclave, and observe that the performance degradation can be as high as 5.6× to even 422.6×. Our investigation identifies these slowdowns are related to architectural features, mainly from page-wise enclave initialization. Leveraging insights from our overhead analysis, we revisit SGX hardware design and make minimal modification to its enclave model. We extend SGX with a new primitive—region-wise plugin enclaves that can be mapped into existing enclaves to reuse attested common states amongst functions. By remapping plugin enclaves, an enclave allows in-situ processing to avoid expensive data movement in a function chain. Experiments show that our design reduces the enclave function latency by 94.74-99.57%, and boosts the autoscaling throughput by 19-179×.
Yubin Xia, Haibo Chen 0001
ISCA2
2021 Tora: A Trusted Blockchain Oracle Based on a Decentralized TEE Network
abstract
Smart Contracts cannot get external data directly due to the closure and determinacy of blockchain itself, limiting the extensibility of blockchain applications. Oracle is proposed to serve as a data feed to offer authenticated and deterministic external data to smart contracts. However, existing centralized oracles are efficient but vulnerable to targeted attacks and suffering from a single point of failure, while existing decentralized oracles are inefficient. The paper presents a trusted blockchain oracle based on a decentralized Trusted Execution Environment (TEE) network called Tora, which embraces both efficiency and availability. The key of Tora is a hybrid consensus mechanism with the confidential available group selection based on Proof-of-Availability (PoA). We have implemented a prototype of Tora on Ethereum with Intel Software Guard eXtensions (SGX) and evaluated it with a data-fetch use case on different measurements, including gas costs, off-chain execution time, group selection results and throughput. The results show that Tora provides considerable efficiency and scalability.
Yubin Xia
JCC3
2021 Scalable Memory Protection in the PENGLAI Enclave
Erhu Feng, Dong Du 0003, Bicheng Yang, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001
OSDI6
2021 Bringing Decentralized Search to Decentralized Services
Jinhao Zhu, Tianxu Zhang, Cheng Tan 0005, Yubin Xia, Sebastian Angel, Haibo Chen 0001
OSDI5
2021 TwinVisor: Hardware-isolated Confidential Virtual Machines for ARM
abstract
Confidential VM, which offers an isolated execution environment for cloud tenants with limited trust in the cloud provider, has recently been deployed in major clouds such as AWS and Azure. However, while ARM has become increasingly popular in cloud data centers, existing confidential VM designs mainly leverage specialized x86 hardware extensions (e.g., AMD SEV and Intel TDX) to isolate VMs upon a shared hypervisor.
Dingji Li, Zeyu Mi, Yubin Xia, Binyu Zang, Haibo Chen 0001, Haibing Guan
SOSP3
2021 TZ-Container: protecting container from untrusted OS with ARM TrustZone
Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang
Sci. China Inf. Sci.4
2021 Enclavisor: A Hardware-Software Co-Design for Enclaves on Untrusted Cloud
abstract
The releases of Intel SGX and AMD SEV mark the transition of hardware-based enclaves from research prototypes to mainstream products. These two paradigms of secure enclaves are attractive to both the cloud providers and tenants, since security is one of the key pillars of cloud computing. However, it is found that current hardware-defined enclaves are not flexible and efficient enough for the cloud. For example, although SGX can provide strong memory protection with both confidentiality and integrity, the size of secure memory is tightly restricted. On the contrary, SEV enables enclaves to use more memory but has critical security flaws due to no memory integrity protection. Meanwhile, both types of enclaves have relatively long booting latency, which makes them not suitable for short-term tasks like serverless workloads. After an in-depth analysis, we find that there are some intrinsic tradeoffs between security and performance due to the limitation of architectural designs. In this article, we investigate a novel hardware-software co-design of enclaves to meet the requirements of cloud by placing a part of the logic of the enclave mechanism into a lightweight software layer, named Enclavisor, to achieve a balance between security, performance, and flexibility. Specifically, our implementation is based on AMD's SEV and, Enclavisor is placed in the guest kernel mode of SEV's secure virtual machines. Enclavisor inherently supports memory encryption with no memory limitation and also achieves efficient booting, multiple enclave granularities, and post-launch remote attestation. Meanwhile, we also propose hardware/software solutions to mitigate the security flaws caused by the lack of memory integrity. We implement a prototype of Enclavisor on an AMD SEV server. The experiments on both micro-benchmarks and application benchmarks show that enclaves on Enclavisor can have close-to-native performance.
Jinyu Gu 0001, Xinyue Wu, Bojun Zhu, Yubin Xia, Binyu Zang, Haibing Guan, Haibo Chen 0001
IEEE Trans. Computers4
2021 Boosting Inter-process Communication with Architectural Support
abstract
IPC (inter-process communication) is a critical mechanism for modern OSes, including not only microkernels such as seL4, QNX, and Fuchsia where system functionalities are deployed in user-level processes, but also monolithic kernels like Android where apps frequently communicate with plenty of user-level services. However, existing IPC mechanisms still suffer from long latency. Previous software optimizations of IPC usually cannot bypass the kernel that is responsible for domain switching and message copying/remapping across different address spaces; hardware solutions such as tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt to the new hardware primitives. In this article, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for efficient and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel and supports secure message passing across multiple processes without copying. We have implemented a prototype of XPC based on the ARM AArch64 with Gem5 simulator and RISC-V architecture with FPGA boards. The evaluation shows that XPC can reduce IPC call latency from 664 to 21 cycles, 14×–123× improvement on Android Binder (ARM), and improve the performance of real-world applications on microkernels by 1.6× on Sqlite3.
Yubin Xia, Dong Du 0003, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001, Haibing Guan
ACM Trans. Comput. Syst.1
2020 Catalyzer: Sub-millisecond Startup for Serverless Computing with Initialization-less Booting
abstract
Serverless computing promises cost-efficiency and elasticity for high-productive software development. To achieve this, the serverless sandbox system must address two challenges: strong isolation between function instances, and low startup latency to ensure user experience. While strong isolation can be provided by virtualization-based sandboxes, the initialization of sandbox and application causes non-negligible startup overhead. Conventional sandbox systems fall short in low-latency startup due to their application-agnostic nature: they can only reduce the latency of sandbox initialization through hypervisor and guest kernel customization, which is inadequate and does not mitigate the majority of startup overhead.
Dong Du 0003, Yubin Xia, Binyu Zang, Guanglu Yan, Chenggang Qin, Qixuan Wu, Haibo Chen 0001
ASPLOS3
2020 Occlum: Secure and Efficient Multitasking Inside a Single Enclave of Intel SGX
abstract
Intel Software Guard Extensions (SGX) enables user-level code to create private memory regions called enclaves, whose code and data are protected by the CPU from software and hardware attacks outside the enclaves. Recent work introduces library operating systems (LibOSes) to SGX so that legacy applications can run inside enclaves with few or even no modifications. As virtually any non-trivial application demands multiple processes, it is essential for LibOSes to support multitasking. However, none of the existing SGX LibOSes support multitasking both securely and efficiently.
Youren Shen, Hongliang Tian, Yu Chen 0004, Kang Chen 0001, Runji Wang, Yubin Xia, Shoumeng Yan
ASPLOS7
2020 Characterizing serverless platforms with serverlessbench
abstract
Serverless computing promises auto-scalability and cost-efficiency (in "pay-as-you-go" manner) for high-productive software development. Because of its virtue, serverless computing has motivated increasingly new applications and services in the cloud. This, however, also presents new challenges including how to efficiently design high-performance serverless platforms and how to efficiently program on the platforms.
Dong Du 0003, Yubin Xia, Binyu Zang, Ziqian Lu, Pingchao Yang, Chenggang Qin, Haibo Chen 0001
SoCC4
2020 TEEp: Supporting Secure Parallel Processing in ARM TrustZone
abstract
Machine learning applications are getting prevelent on various computing platforms, including cloud servers, smart phones, IoT devices, etc. For these applications, security is one of the most emergent requirements. While trusted execution environment (TEE) like ARM TrustZone has been widely used to protect critical prodecures including fingerprint authentication and mobile payment, state-of-the-art implementations of TEE OS lack the support for multi-threading and are not suitable for computing-intensive workloads. This is because current TEE OSes are usually designed for hosting security critical tasks, which are typically small and non-computing-intensive. Thus, most of TEE OSes do not support multi-threading in order to minimize the size of the trusted computing base (TCB). In this paper, we propose TEEp, a system that enables multi-threading in TEE without weakening security, and supports existing multi-threaded applications to run directly in TEE. Our design includes a novel multithreading mechanism based on the cooperation between the TEE OS and the host OS, without trusting the host OS. We implement our system based on OP-TEE and port it to two platforms: a HiKey 970 development board as mobile platform, and a Huawei Hi1610 ARM server as server platform. We run TensorFlow Lite on the development board and TensorFlow on the server for performance evaluation in TEE. The result shows that our system can improve the throughput of TensorFlow Lite on 5 models to 3.2x when 4 cores are available, with 13.5% overhead compared with Linux on average.
Zinan Li, Wenhao Li 0009, Yubin Xia, Binyu Zang
ICPADS3
2020 HCloud: A Serverless Platform for JointCloud Computing
abstract
Most of today's cloud providers, including Amazon, Google and Microsoft, usually offer computing services with the abstraction of virtual machines (VMs), which are also known as instances. Even though these public clouds share similar infrastructures, they provide different service quality standards and price models for users, which fluctuate according to the dynamic loads. The JointCloud computing model empowers the cooperation among multiple public clouds, which can provide the users better service quality with reasonable prices. However, it is difficult to dynamically cooperate among clouds based on the traditional VM abstraction. Any computation migration among clouds will trigger costly VM states transfer, which makes it extremely hard for users to adjust their deployment models swiftly. Fortunately, serverless computing model is getting popular recently and has been supported by major cloud vendors. This model allows users to upload their computing tasks to the cloud in the unit of function. Users do not need to consider all of the tedious works like virtual server maintenance; instead, the cloud automatically instantiates function workers to handle incoming requests on the fly. In this paper, we propose a new JointCloud platform called HCloud, which efficiently manages computing resources of multiple clouds while offering the server-less model to the users. The fine-grained granularity of function enables HCloud to flexibly migrate workloads among clouds, thus brings better service qualities and lower prices at the same time. For example, HCloud can reduce the cost by requesting resource from a cheaper cloud for latency-insensitive jobs while routing the requests of latency-sensitive ones to the nearest and performant clouds. We have implemented a prototype of HCloud and evaluated it by simulating multiple cloud providers. The evaluation results show that HCloud can greatly improve the performance of serverless workloads with small costs.
Zeyu Mi, Zhichao Hua 0001, Yubin Xia
JCC5
2020 A Survey on Serverless Computing and Its Implications for JointCloud Computing
abstract
Serverless computing is known as an appealing alternative cloud computing paradigm with its auto-scaling nature and pay-as-you-go charging model. Mainstream cloud vendors have proposed their own serverless platforms, while various kinds of applications have been refactored in a serverless manner for execution. However, the serverless computing model still entails refinement as it introduces performance and security issues. In this paper, we conduct a comprehensive survey on the serverless computing, mainly in three aspects: the type of applications suitable for serverless, the performance issues, and the security issues. We specifically elaborate previous efforts on resolving issues in serverless and shed light on the unresolved issues. We also discuss the opportunities and challenges in integrating serverless computing in the jointcloud infrastructure.
Mingyu Wu 0001, Zeyu Mi, Yubin Xia
JCC3
2020 Harmonizing Performance and Isolation in Microkernels with Efficient Intra-kernel Isolation and Communication
Jinyu Gu 0001, Xinyue Wu, Wentai Li, Zeyu Mi, Yubin Xia, Haibo Chen 0001
USENIX ATC6
2020 Iso-UniK: lightweight multi-process unikernel through memory protection keys
abstract
Abstract Unikernel, specializing a minimalistic libOS with an application, is an attractive design for cloud computing. However, the Achilles’ heel of unikernel is the lack of multi-process support, which makes it less flexible and applicable. Many applications rely on the process abstraction to isolate different components. For example, Apache with the multi-processing module isolates a request handler in a process to guarantee security. Prior art tackles the problem by simulating multi-process with multiple unikernels, which is incompatible with existing cloud providers and also introduces high overhead. This paper proposes Iso-UniK, a new unikernel design enabling multi-task applications with the support of both functionality and isolation. Iso-UniK leverages a recent hardware feature, named Intel Memory Protection Key (Intel MPK), to provide lightweight and efficient isolation for multi-process in unikernel. Our design has three benefits compared with previous approaches. First, Iso-UniK does not need hypervisor support and is thus compatible with existing cloud computing platforms; second, Iso-UniK promises fast system calls with only 45 cycles; last, a process can be isolated with a flexible configuration. We have implemented a prototype based on OSv, a unikernel system supporting unmodified applications. Iso-UniK can achieve fast fork operation with only 66 μs for multi-process applications. Our evaluation shows that the isolation and multi-process support in Iso-UniK will not damage the applications’ performance.
Dong Du 0003, Yubin Xia
Cybersecur.3
2019 XPC: architectural support for secure and efficient cross process call
abstract
Microkernel has many intriguing features like security, fault-tolerance, modularity and customizability, which recently stimulate a resurgent interest in both academia and industry (including seL4, QNX and Google's Fuchsia OS). However, IPC (inter-process communication), which is known as the Achilles' Heel of microkernels, is still the major factor for the overall (poor) OS performance. Besides, IPC also plays a vital role in monolithic kernels like Android Linux, as mobile applications frequently communicate with plenty of user-level services through IPC. Previous software optimizations of IPC usually cannot bypass the kernel which is responsible for domain switching and message copying/remapping; hardware solutions like tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt the new hardware primitives. In this paper, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for fast and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel, and supports message passing across multiple processes through the invocation chain without copying. The primitive is compatible with the traditional address space based isolation mechanism and can be easily integrated into existing microkernels and monolithic kernels. We have implemented a prototype of XPC based on a Rocket RISC-V core with FPGA boards and ported two microkernel implementations, seL4 and Zircon, and one monolithic kernel implementation, Android Binder, for evaluation. We also implement XPC on GEM5 simulator to validate the generality. The result shows that XPC can reduce IPC call latency from 664 to 21 cycles, up to 54.2x improvement on Android Binder, and improve the performance of real-world applications on microkernels by 1.6x on Sqlite3 and 10x on an HTTP server with minimal hardware resource cost.
Dong Du 0003, Zhichao Hua 0001, Yubin Xia, Binyu Zang, Haibo Chen 0001
ISCA3
2019 TEEv: virtualizing trusted execution environments on mobile platforms
abstract
Trusted Execution Environments (TEE) are widely deployed, especially on smartphones. A recent trend in TEE development is the transition from vendor-controlled, single-purpose TEEs to open TEEs that host Trusted Applications (TAs) from multiple sources with independent tasks. This transition is expected to create a TA ecosystem needed for providing stronger and customized security to apps and OS running in the Rich Execution Environment (REE). However, the transition also poses two security challenges: enlarged attack surface resulted from the increased complexity of TAs and TEEs; the lack of trust (or isolation) among TAs and the TEE.
Wenhao Li 0009, Yubin Xia, Long Lu, Haibo Chen 0001, Binyu Zang
VEE2
2019 Protecting mobile devices from physical memory attacks with targeted encryption
abstract
Sensitive data in a process could be scattered over the memory of a computer system for a prolonged period of time. Unfortunately, DRAM chips were proven insecure in previous studies. The problem becomes worse in the mobile environment, in which users' smartphones are easily lost or stolen. The powered-on phones may contain sensitive data in the vulnerable DRAM chips. In this paper, we propose MemVault, a mechanism to protect sensitive data in Android devices against physical memory attacks. MemVault keeps track of the propagation of well-marked sensitive data sources, and selectively encrypts tainted sensitive memory contents in the DRAM chip. When a tainted object is accessed, MemVault redirects the access to the internal RAM (iRAM), where the cipher-text object is decrypted transparently. iRAM is a system-on-chip (SoC) component which is by nature immune to physical memory exploits. We have implemented a MemVault prototype system, and have evaluated it with extensive experiments. Our results validate that MemVault effectively eliminates the occurrences of clear-text sensitive objects in DRAM chips, and imposes acceptable overheads.
Le Guan, Chen Cao 0004, Sencun Zhu, Jingqiang Lin 0001, Peng Liu 0005, Yubin Xia, Bo Luo
WiSec6
2018 VButton: Practical Attestation of User-driven Operations in Mobile Apps
abstract
More and more malicious apps and mobile rootkits are found to perform sensitive operations on behalf of legitimate users without their awareness. Malware does so by either forging user inputs or tricking users into making unintended requests to online service providers. Such malware is hard to detect and generates large revenues for cybercriminals, which is often used for committing ad/click frauds, faking reviews/ratings, promoting people or business on social networks, etc.
Wenhao Li 0009, Shiyu Luo, Zhichuang Sun, Yubin Xia, Long Lu, Haibo Chen 0001, Binyu Zang, Haibing Guan
MobiSys4
2018 EPTI: Efficient Defence against Meltdown Attack for Unpatched VMs
Zhichao Hua 0001, Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang
USENIX ATC3
2018 SplitPass: A Mutually Distrusting Two-Party Password Manager
Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang, Zhenkai Liang
J. Comput. Sci. Technol.3
2018 ShadowEth: Private Smart Contract on Public Blockchain
Yubin Xia, Haibo Chen 0001, Binyu Zang, Jan Xie
J. Comput. Sci. Technol.2
2017 Secure Live Migration of SGX Enclaves on Untrusted Cloud
abstract
The recent commercial availability of Intel SGX (Software Guard eXtensions) provides a hardware-enabled building block for secure execution of software modules in an untrusted cloud. As an untrusted hypervisor/OS has no access to an enclave's running states, a VM (virtual machine) with enclaves running inside loses the capability of live migration, a key feature of VMs in the cloud. This paper presents the first study on the support for live migration of SGX-capable VMs. We identify the security properties that a secure enclave migration process should meet and propose a software-based solution. We leverage several techniques such as two-phase checkpointing and self-destroy to implement our design on a real SGX machine. Security analysis confirms the security of our proposed design and performance evaluation shows that it incurs negligible performance overhead. Besides, we give suggestions on the future hardware design for supporting transparent enclave migration.
Jinyu Gu 0001, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan
DSN3
2017 Deconstructing Xen
Le Shi, Yuming Wu, Yubin Xia, Nathan Dautenhahn, Haibo Chen 0001, Binyu Zang
NDSS3
2017 vTZ: Virtualizing ARM TrustZone
Zhichao Hua 0001, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan
USENIX Security Symposium3
2017 Secure Outsourcing of Virtual Appliance
abstract
Computation outsourcing using virtual appliance is getting prevalent in cloud computing. However, with both hardware and software being controlled by potentially curious or even malicious cloud operators, it is no surprise to see frequent reports of security accidents, like data leakages or abuses. This paper proposes Kite, a hardware-software framework that guards the security of tenant's virtual machine (VM), in which the outsourced computation is encapsulated. Kite only trusts the processor and makes no security assumption on external memory, devices, or hypervisor. Unlike prior hardware-based approaches, Kite retains transparency with existing VM and requires few changes to the (untrusted) hypervisor by introducing VM-Shim mechanism. Each VM-Shim instance runs in between its VM and the hypervisor, which only transfers necessary information designated by the VM to the hypervisor and external environments. Kite also considers the high-level semantic of interaction between VM and hypervisor to defend against attacks through legitimate operations or interfaces. We have implemented a prototype of Kite's secure processor in a QEMU-based full-system emulator and its software components on real machine. Evaluation shows that the performance overhead of Kite ranges from 0.5-14.0 percent on simulated platform and 0.4-7.3 percent on real hardware.
Yubin Xia, Haibing Guan, Yunji Chen, Tianshi Chen 0002, Binyu Zang, Haibo Chen 0001
IEEE Trans. Cloud Comput.1
2017 Efficient and Available In-Memory KV-Store with Hybrid Erasure Coding and Replication
abstract
In-memory key/value store (KV-store) is a key building block for many systems like databases and large websites. Two key requirements for such systems are efficiency and availability, which demand a KV-store to continuously handle millions of requests per second. A common approach to availability is using replication, such as primary-backup (PBR), which, however, requires M +1 times memory to tolerate M failures. This renders scarce memory unable to handle useful user jobs. This article makes the first case of building highly available in-memory KV-store by integrating erasure coding to achieve memory efficiency, while not notably degrading performance. A main challenge is that an in-memory KV-store has much scattered metadata. A single KV put may cause excessive coding operations and parity updates due to excessive small updates to metadata. Our approach, namely Cocytus, addresses this challenge by using a hybrid scheme that leverages PBR for small-sized and scattered data (e.g., metadata and key), while only applying erasure coding to relatively large data (e.g., value). To mitigate well-known issues like lengthy recovery of erasure coding, Cocytus uses an online recovery scheme by leveraging the replicated metadata information to continuously serve KV requests. To further demonstrate the usefulness of Cocytus, we have built a transaction layer by using Cocytus as a fast and reliable storage layer to store database records and transaction logs. We have integrated the design of Cocytus to Memcached and extend it to support in-memory transactions. Evaluation using YCSB with different KV configurations shows that Cocytus incurs low overhead for latency and throughput, can tolerate node failures with fast online recovery, while saving 33% to 46% memory compared to PBR when tolerating two failures. A further evaluation using the SmallBank OLTP benchmark shows that in-memory transactions can run atop Cocytus with high throughput, low latency, and low abort rate and recover fast from consecutive failures.
Haibo Chen 0001, Mingkai Dong 0002, Yubin Xia, Haibing Guan, Binyu Zang
ACM Trans. Storage5
2016 A Case for Virtualizing Persistent Memory
abstract
With the proliferation of software and hardware support for persistent memory (PM) like PCM and NV-DIMM, we envision that PM will soon become a standard component of commodity cloud, especially for those applications demanding high performance and low latency. Yet, current virtualization software lacks support to efficiently virtualize and manage PM to improve cost-effectiveness, performance, and endurance.
Rong Chen 0001, Haibo Chen 0001, Yubin Xia, KwanJong Park, Binyu Zang, Haibing Guan
SoCC4
2016 Mitigating Sync Amplification for Copy-on-write Virtual Disk
Qingshu Chen, Yubin Xia, Haibo Chen 0001
FAST3
2015 Thwarting Memory Disclosure with Efficient Hypervisor-enforced Intra-domain Isolation
abstract
Exploiting memory disclosure vulnerabilities like the HeartBleed bug may cause arbitrary reading of a victim's memory, leading to leakage of critical secrets such as crypto keys, personal identity and financial information. While isolating code that manipulates critical secrets into an isolated execution environment is a promising countermeasure, existing approaches are either too coarse-grained to prevent intra-domain attacks, or require excessive intervention from low-level software (e.g., hypervisor or OS), or both. Further, few of them are applicable to large-scale software with millions of lines of code. This paper describes a new approach, namely SeCage, which retrofits commodity hardware virtualization extensions to support efficient isolation of sensitive code manipulating critical secrets from the remaining code. SeCage is designed to work under a strong adversary model where a victim application or even the OS may be controlled by the adversary, while supporting large-scale software with small deployment cost. SeCage combines static and dynamic analysis to decompose monolithic software into several compart- ments, each of which may contain different secrets and their corresponding code. Following the idea of separating control and data plane, SeCage retrofits the VMFUNC mechanism and nested paging in Intel processors to transparently provide different memory views for different compartments, while allowing low-cost and transparent invocation across domains without hypervisor intervention.
Haibo Chen 0001, Yubin Xia
CCS5
2015 TinMan: eliminating confidential mobile data exposure with security oriented offloading
abstract
The wide adoption of smart devices has stimulated a fast shift of security-critical data from desktop to mobile devices. However, recurrent device theft and loss expose mobile devices to various security threats and even physical attacks. This paper presents TinMan, a system that protects confidential data such as web site password and credit card number (we use the term cor to represent these data, which is short for Confidential Record) from being leaked or abused even under device theft. TinMan separates accesses of cor from the rest of the functionalities of an app, by introducing a trusted node to store cor and offloading any code from a mobile device to the trusted node to access cor. This completely eliminates the exposure of cor on the mobile devices. The key challenges to TinMan include deciding when and how to efficiently and transparently offload execution; TinMan addresses these challenges with security-oriented offloading with a low-overhead tainting scheme called asymmetric tainting to track accesses to cor to trigger offloading, as well as transparent SSL session injection and TCP pay-load replacement to offload accesses to cor. We have implemented a prototype of TinMan based on Android and demonstrated how TinMan protects the information of user's bank account and credit card number without modifying the apps. Evaluation results also show that TinMan incurs only a small amount of performance and power overhead.
Yubin Xia, Cheng Tan 0005, Haibing Guan, Binyu Zang, Haibo Chen 0001
EuroSys1
2015 "Anti-Caching"-based elastic memory management for Big Data
abstract
The increase in the capacity of main memory coupled with the decrease in cost has fueled the development of in-memory database systems that manage data entirely in memory, thereby eliminating the disk I/O bottleneck. However, as we shall explain, in the Big Data era, maintaining all data in memory is impossible, and even unnecessary. Ideally we would like to have the high access speed of memory, with the large capacity and low price of disk. This hinges on the ability to effectively utilize both the main memory and disk. In this paper, we analyze state-of-the-art approaches to achieving this goal for in-memory databases, which is called as “Anti-Caching” to distinguish it from traditional caching mechanisms. We conduct extensive experiments to study the effect of each fine-grained component of the entire process of “Anti-Caching” on both performance and prediction accuracy. To avoid the interference from other unrelated components of specific systems, we implement these approaches on a uniform platform to ensure a fair comparison. We also study the usability of each approach, and how intrusive it is to the systems that intend to incorporate it. Based on our findings, we propose some guidelines on designing a good “Anti-Caching” approach, and sketch a general and efficient approach, which can be utilized in most in-memory database systems without much code modification.
Hao Zhang 0029, Gang Chen 0001, Beng Chin Ooi, Weng-Fai Wong, Shensen Wu, Yubin Xia
ICDE6
2015 Reducing world switches in virtualized environment with flexible cross-world calls
abstract
Modern computers are built with increasingly complex software stack crossing multiple layers (i.e., worlds), where cross-world call has been a necessity for various important purposes like security, reliability, and reduced complexity. Unfortunately, there is currently limited cross-world call support (e.g., syscall, vmcall), and thus other calls need to be emulated by detouring multiple times to the privileged software layer (i.e., OS kernel and hypervisor). This causes not only significant performance degradation, but also unnecessary implementation complexity.
Wenhao Li 0009, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan
ISCA2
2015 AdAttester: Secure Online Mobile Advertisement Attestation Using TrustZone
abstract
Mobile advertisement (ad for short) is a major financial pillar for developers to provide free mobile apps. However, it is frequently thwarted by ad fraud, where rogue code tricks ad providers by forging ad display or user clicks, or both. With the mobile ad market growing drastically (e.g., from $8.76 billion in 2012 to $17.96 billion in 2013), it is vitally important to provide a verifiable mobile ad framework to detect and prevent ad frauds. Unfortunately, this is notoriously hard as mobile ads usually run in an execution environment with a huge TCB.
Wenhao Li 0009, Haibo Chen 0001, Yubin Xia
MobiSys4
2015 Poster: TVisor - A Practical and Lightweight Mobile Red-Green Dual-OS Architecture
abstract
Mobile and embedded system software designer are often torn between choosing security and functionality. In particular, the security of out-of-band execution environment is sensitive to rich functionality. ARM TrustZone has been used to develop a Trusted Execution Environment (TEE), which runs in parallel with rich functionality commodity OS and provides an isolated and tamper-resistant execution context for trusted applications. ARM TrustZone splits access of the processor, memory and peripherals into two different worlds, namely normal world and secure world. The secure world is more privileged and the recommended context to implement TEE. However, despite the security of TrustZone TEE, the functionality is very limited.
Wenhao Li 0009, Yubin Xia, Haibo Chen 0001
MobiSys4
2014 Concurrent and consistent virtual machine introspection with hardware transactional memory
abstract
Virtual machine introspection, which provides tamperresistant, high-fidelity “out of the box” monitoring of virtual machines, has many prominent security applications including VM-based intrusion detection, malware analysis and memory forensic analysis. However, prior approaches are either intrusive in stopping the world to avoid race conditions between introspection tools and the guest VM, or providing no guarantee of getting a consistent state of the guest VM. Further, there is currently no effective means for timely examining the VM states in question. In this paper, we propose a novel approach, called TxIntro, which retrofits hardware transactional memory (HTM) for concurrent, timely and consistent introspection of guest VMs. Specifically, TxIntro leverages the strong atomicity of HTM to actively monitor updates to critical kernel data structures. Then TxIntro can mount introspection to timely detect malicious tampering. To avoid fetching inconsistent kernel states for introspection, TxIntro uses HTM to add related synchronization states into the read set of the monitoring core and thus can easily detect potential inflight concurrent kernel updates. We have implemented and evaluated TxIntro based on Xen VMM on a commodity Intel Haswell machine that provides restricted transactional memory (RTM) support. To demonstrate the effectiveness of TxIntro, we implemented a set of kernel rootkit detectors using TxIntro. Evaluation results show that TxIntro is effective in detecting these rootkits, and is efficient in adding negligible performance overhead.
Yubin Xia, Haibing Guan, Binyu Zang, Haibo Chen 0001
HPCA2
2013 Architecture support for guest-transparent VM protection from untrusted hypervisor and physical attacks
abstract
The privacy and integrity of tenant's data highly rely on the infrastructure of multi-tenant cloud being secure. However, with both hardware and software being controlled by potentially curious or even malicious cloud operators, it is no surprise to see frequent reports of data leakages or abuses in cloud. Unfortunately, most prior solutions require intrusive changes to the cloud platform and none can protect a VM against adversaries controlling the physical machine. This paper analyzes the challenges of transparent VM protection against sophisticated adversaries controlling the whole software and hardware stack. Based on the analysis, this paper proposes HyperCoffer, a hardware-software framework that guards the privacy and integrity of tenant's VMs. HyperCoffer only trusts the processor chip and makes no security assumption on external memory and devices. Hyper-Coffer extends existing processor virtualization with memory encryption and integrity checking to secure data communication with off-chip memory. Unlike prior hardware-based approaches, HyperCoffer retains transparency with existing virtual machines (i.e., operating systems) and requires very few changes to the (untrusted) hypervisor. HyperCoffer introduces a mechanism called VM-Shim that runs in-between a guest VM and the hypervisor. Each VM-Shim instance for a VM runs in a separate protected context and only declassifies necessary information designated by the VM to the hypervisor and external environments (e.g., through NICs). We have implemented a prototype of HyperCoffer in a QEMU-based full-system emulator and the VM-Shim mechanism in a real machine. Performance measurement using trace-based simulation and on a real hardware platform shows that the performance overhead is small (ranging from 0.6% to 13.9% on simulated platform and 0.3% to 6.8% on real hardware for the VM-Shim mechanism).
Yubin Xia, Haibo Chen 0001
HPCA1
2012 CFIMon: Detecting violation of control flow integrity using performance counters
abstract
Many classic and emerging security attacks usually introduce illegal control flow to victim programs. This paper proposes an approach to detecting violation of control flow integrity based on hardware support for performance monitoring in modern processors. The key observation is that the abnormal control flow in security breaches can be precisely captured by performance monitoring units. Based on this observation, we design and implement a system called CFIMon, which is the first non-intrusive system that can detect and reason about a variety of attacks violating control flow integrity without any changes to applications (either source or binary code) or requiring special-purpose hardware. CFIMon combines static analysis and runtime training to collect legal control flow transfers, and leverages the branch tracing store mechanism in commodity processors to collect and analyze runtime traces on-the-fly to detect violation of control flow integrity. Security evaluation shows that CFIMon has low false positives or false negatives when detecting several realistic security attacks. Performance results show that CFIMon incurs only 6.1% performance overhead on average for a set of typical server applications.
Yubin Xia, Haibo Chen 0001, Binyu Zang
DSN1
2009 PaS: A Preemption-aware Scheduling Interface for Improving Interactive Performance in Consolidated Virtual Machine Environment
abstract
As virtualization technology is used widely in cloud computing, there are more and more interactive workloads being deployed on virtual machine (VM) environment. Although improving interactive performance has been heavily studied in operating system area, in consolidated VM environment, the improvements of guest OS are usually offset by the more coarse-grained VM scheduler, which may cause poor interactive performance. The guest OS scheduler and VM scheduler are totally independent with each other, which leads to the so called 'semantic gap'. To reduce this semantic gap, this paper presents PaS (Preemption-aware Scheduling) as an extension of VM scheduling interface. PaS introduces only two interfaces: one to register VM preemption conditions, the other to check if a VM is preempting. Thanks to the sophisticated techniques of interactive-process identification and optimization in traditional OS, it is trivial for guest OS to use the new interfaces: only 10 lines of code are added into Linux 2.6.18.8. The evaluation results show that PaS can significantly improve the interactive performance of consolidated VMs while keeping the fairness and performance isolation.
Yubin Xia, Xu Cheng 0001
ICPADS1
2007 A Fast Lossless Codec of Continuous-Tone Images for Thin Client Computing
abstract
Summary form only given. We propose a fast and efficient lossless codec of continuous-tone images, SPEDIC (simple predictor and edge detector based image codec), which is uniquely suitable for coding screen updates generated by multimedia applications in thin client computing systems. A codec for thin client computing should take account of the tradeoff between compression ratio and coding complexity, because screen update images in thin client computing should be sent from server to client in a timely fashion. We argue that by combining similar but simpler building blocks of the state-of-the-art methods, a little inferior compression ratio of JPEG-LS (the standard for lossless image compression today) can be attained with much lower coding complexity.
Yan Niu, Yubin Xia, Xu Cheng 0001
DCC3
2007 A Fast and Efficient Codec for Multimedia Applications in Wireless Thin-Client Computing
abstract
Thin-client computing is uniquely suitable for mobile environments, where resource-poor devices may need to access critical applications over wireless networks. In thin-client computing, applications run on a powerful server, which sends screen updates to the client in real time. However, fast and efficient coding methods for screen updates of multimedia applications are very challenging in wireless environments, for the limited bandwidth and the real-time constraint. In this paper, we present SPEDIC (Simple Predictor and Edge Detector based Image Codec), a fast and efficient codec of continuous-tone images, which is appropriate for multimedia applications in wireless thin-client computing. SPEDIC adopts predictive coding, edge coding and run coding, depending on the characteristics of image blocks, and trades off between compression ratio and coding complexity. The experimental results show that SPEDIC provides good compression ratio with very low coding complexity, especially for continuous-tone images, and achieves the best end-to-end latency over wireless networks. Compared with JPEG-LS, the standard of lossless image compression, SPEDIC compresses 2.5 times faster and decompresses 3.2 times faster than it does for a series of continuous-tone images, and is only 6.2% inferior to it in compression ratio.
Yan Niu, Yubin Xia, Xu Cheng 0001
WOWMOM3