Hao Fan 0006

dblp:23/2535-6 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-6741-0448ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 6 first-author · 17 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 AEP: Achieving Hierarchical Fault Tolerance in DSM Through Atomic Execution Protection
abstract
Recent advances in CXL 3.0 have revitalized Distributed Shared Memory (DSM), enabling zero-copy sharing across machine nodes at near-NUMA latency and making shared-memory communication practical for distributed applications. However, DSM memory allocators remain vulnerable to partial failures: client crashes during critical, non-atomic memory operations can corrupt metadata and stall the entire system. Existing solutions either block all clients during recovery or impose heavy runtime overhead due to distributed fault-tolerance protocols. We propose Atomic Execution Protection (AEP), a kernel-assisted fault tolerance mechanism that guarantees atomic completion of user-space critical operations by deferring termination until protected regions finish. AEP tolerates intra-node partial failures without costly distributed coordination. To extend beyond a single node, we further design AEP-DSM, a hierarchical DSM allocator that combines AEP-protected node-level allocators with cross-node coordination. Our Linux AEP prototype integrated with a cross-node DSM consistency protocol (CXL-SHM) achieves near-native local performance and improves throughput by up to 13.3x over state-of-the-art DSM allocators.
Zixuan Wang 0030, Hang Huang, Jia Rao, Hui Lu 0001, Hao Fan 0006, Song Wu 0001, Hai Jin 0001
EuroSys6
2026 Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
abstract
Serverless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage.
Hao Wu 0032, Yaochen Liu, Minchen Yu, Qizhen Weng 0001, Junxiao Deng, Hao Fan 0006, Song Wu 0001, Wei Wang 0030, Hai Jin 0001
EuroSys7
2026 MCOP: A Multiple Containers in One Pod Placement Strategy towards Application Completion Time Minimization
Ziyou Si, Lin Gu 0002, Deze Zeng, Hao Fan 0006, Quan Chen 0002
INFOCOM4
2026 KPT-Fork: Enhancing User Data Isolation via Kernel Page Table Fork
abstract
Modern operating systems (OS) utilize a shared kernel address design, enabling the OS kernel to access physical memory addresses conveniently and efficiently. However, this design introduces vulnerabilities that attackers can exploit to arbitrarily read/write content from any kernel-space address, known as data-oriented attacks. Existing MMU-based isolation approaches face inherent challenges in scalability, compatibility, and context switch overhead. Inspired by the reference kernel page table, we propose to use user-private reference kernel page tables (UP-KPTs) as the foundation for data isolation in a shared kernel address space. A UP-KPT is a customized kernel page table tailored for a specific container, encompassing address mappings of private user data that are not visible to other containers. To protect sensitive system-wide data, an administrator can unmap such data from UP-KPTs and make it inaccessible to potentially malicious containers. UP-KPT provides strong data isolation while maintaining full compatibility with legacy applications and existing kernel components. We develop KPT-fork, a kernel isolation framework that enables the creation and management of UP-KPTs. KPT-fork entails two important designs: 1) a UP-KPT structure that enables data isolation while preserving the simplicity, convenience, and efficiency of a shared kernel page table; 2) a private memory allocator that efficiently transforms static and direct address mappings of private data in shared kernel addresses into dynamic user-private mappings. We demonstrate through three case studies that KPT-fork can be extended to protect various types of data. Our evaluation shows that KPT fork can defend against a wide range of data-oriented attacks while incurring negligible overhead.
Zixuan Wang 0030, Lisong Pan, Jia Rao, Hao Fan 0006, Song Wu 0001
IEEE Trans. Cloud Comput.4
2026 When Server Joins INA: The Resource-Aware Repair Acceleration for Erasure-Coded Storage Systems
Geyao Cheng, Junxu Xia, Hao Fan 0006, Fengzeng Liu, Haibo Mi, Deke Guo
IEEE Trans. Parallel Distributed Syst.3
2026 Puffer: A Serverless Platform Based on Vertical Memory Scaling
abstract
This paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. Based on Faascale, we realize a serverless platform, named Puffer. Experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies, Puffer reduces time for cold-starting MicroVMs by 89.01%, improves memory utilization by 17.66%, and decreases functions execution time by 23.93% on average.
Hao Fan 0006, Kun Wang 0059, Haibo Mi, Song Wu 0001, Chen Yu 0003
IEEE Trans. Parallel Distributed Syst.1
2026 HyDuo: A Low-Latency and Cost-Efficient Storage System for Stateful Serverless Applications
Hao Fan 0006, Hanxiang Huang, Song Wu 0001
IEEE Trans. Serv. Comput.4
2025 WAF: An Efficient WebAssembly-Based Execution Environment for User-Defined Functions
abstract
User-Defined Functions (UDFs) have long served as the standard method for extending the capabilities of data management systems. With the advent of WebAssembly (WASM), UDFs' dependencies, such as language runtimes and libraries, can be compiled into a WASM module, which is then instantiated to execute the UDF. This approach offers several key advantages: 1) it allows developers to write UDFs in their preferred programming language, rather than being limited to those natively supported by the database engine; 2) it isolates UDFs' dependencies within the WASM module, mitigating the risk of errors caused by conflicting dependencies on the same host; and 3) it promotes cross-platform compatibility, enabling seamless execution of UDFs across different engines, operating systems, and architectures. However, our analysis reveals that executing a WASM-based UDF incurs overhead due to data transfer between the database engine and the WASM runtime. This process involves data copying and data layout adjustments, which can significantly impact performance. To address these challenges, we present WAF, a WASM-based UDF execution environment. WAF leverages shared memory to eliminate data copying and shifts data layout adjustments from the execution phase to the compilation phase. Experimental results show that WAF reduces the execution overhead of WASM-based UDFs by 3.1x and achieves an 18.1x speedup compared to the container-based approach, eliminating nearly all data transfer delays.
Hao Fan 0006, Junhui Peng, Song Wu 0001, Chen Yu 0003, Hai Jin 0001, Wei Yang 0013
ICDE2
2025 EDDE: Container Deployment Framework Beyond the Cloud
abstract
Containers, renowned for their lightweight nature and flexibility, have seen growing adoption for deploying edge services such as web applications. However, existing cloud-oriented container deployment frameworks fail to address the unique challenges of edge environments, including geographical distribution, device heterogeneity, and resource constraints. This oversight leads to suboptimal performance for latency-sensitive edge services like HPC/AI-powered autonomous driving and edge gaming, which demand rapid startup and immediate responsiveness.
Hao Fan 0006, Shadi Ibrahim, Lin Gu 0002, Song Wu 0001
SC1
2025 High-performance BFT consensus for Metaverse through block linking and shortcut loop
Chaozheng Ding, Xiaohai Dai, Hao Fan 0006, Jianwen Xiang
Comput. Commun.4
2025 System log isolation for containers
abstract
Abstract Container-based virtualization is increasingly popular in cloud computing due to its efficiency and flexibility. Isolation is a fundamental property of containers and weak isolation could cause significant performance degradation and security vulnerability. However, existing works have almost not discussed the isolation problems of system log which is critical for monitoring and maintenance of containerized applications. In this paper, we present a detailed isolation analysis of system log in current container environment. First, we find several system log isolation problems which can cause significant impacts on system usability, security, and efficiency. For example, system log accidentally exposes information of host and co-resident containers to one container, causing information leakage. Second, we reveal that the root cause of these isolation problems is that containers share the global log configuration, the same log storage, and the global log view. To address these problems, we design and implement a system named private logs (POGs). POGs provides each container with its own log configuration and stores logs individually for each container, avoiding log configuration and storage sharing, respectively. In addition, POGs enables private log view to help distinguish which container the logs belong to. The experimental results show that POGs can effectively enhance system log isolation for containers with negligible performance overhead.
Kun Wang 0005, Song Wu 0001, Yanxiang Cui, Hao Fan 0006, Hai Jin 0001
Frontiers Comput. Sci.5
2025 CBuild: Cluster-Oriented Collaborative Image Building for Containers
abstract
Starting a container needs to build a container image layer-by-layer if the required image is not available. However, the image building involves downloading a large amount of data, which significantly delays the development and deployment of containerized services. To reduce data downloads and accelerate image building, current methods typically focus on improving data sharing through reconstructing images. Unfortunately, these approaches show limited performance improvement in clusters as they only improve data sharing on a single node. In this paper, we find that there are significant duplicated remote file downloads between nodes in a cluster. Accordingly, we propose cBuild, a distributed file cache to minimize costly image data downloads in cluster environments. Specifically, to enable inter-node image data sharing, cBuild designs a non-intrusive interception mechanism based on network namespace, instead of directly detecting building instructions that dirty images. Based on the distribution characteristics of duplicated files in layers, cBuild places image files among nodes in a balanced manner to prevent transfer bottlenecks caused by hotspot nodes and employs a layer-aware searching strategy to quickly locate the desired files. We implement cBuild on the basis of Docker. Experiments show that cBuild improves building speed by up to 15.3 × and reduces the data downloading by 80%.
Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers2
2025 KubeSPT: Stateful Pod Teleportation for Service Resilience With Live Migration
abstract
Container orchestration systems, such as Kubernetes, streamline containerized application deployment. As more and more applications are being deployed in Kubernetes, there is an increasing need for rescheduling - relocating a running pod to different nodes - due to system upgrades, node failures, and load-balancing optimizations. Live migration, which transfers services from source nodes to target nodes with minimal downtime, is the ideal support for rescheduling. However, implementing live migration for pods that run stateful services is challenging, because Kubernetes manages pods as stateless. First, the current pod's network namespace initialization process causes a mismatch in the network state between the migrated pod and internal containers. Second, migrating the memory state results in extended downtime. Third, Kubernetes operations on pods do not consider preserving the state of the pods. Therefore, we propose KubeSPT to achieve live migration of stateful pods in rescheduling scenarios. Firstly, we synchronize the network state of pods and internal containers by controlling packet flow and implement fast service redirection. Secondly, we introduce a Hot Data and Lazy-Restore method for memory restoration to reduce migration downtime. Finally, we decouple pod migration operations from other Kubernetes operations to ensure compatibility with live migration. Experimental results show that KubeSPT reduces downtime by 86%-93% compared to current rescheduling methods.
Hansheng Zhang, Song Wu 0001, Hao Fan 0006, Weibin Xue, Chen Yu 0003, Shadi Ibrahim, Hai Jin 0001
IEEE Trans. Serv. Comput.3
2024 Faascale: Scaling MicroVM Vertically for Serverless Computing with Memory Elasticity
abstract
This paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. We implement Faascale on Amazon Firecracker to evaluate its gains for the serverless platform. The results of experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies based the state-of-the-art snapshots technique, Faascale reduces time for cold-starting MicroVMs by 89.01% and functions execution time by 23.93% on average.
Qiang He 0001, Hao Fan 0006, Song Wu 0001
SoCC3
2024 StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow
Hao Wu 0010, Junxiao Deng, Shadi Ibrahim, Song Wu 0001, Hao Fan 0006, Ziyue Cheng, Hai Jin 0001
USENIX ATC6
2024 Precise control of page cache for containers
Kun Wang 0005, Song Wu 0001, Shengbang Li, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001
Frontiers Comput. Sci.5
2024 QoS-pro: A QoS-enhanced Transaction Processing Framework for Shared SSDs
abstract
Solid State Drives (SSDs) are widely used in data-intensive scenarios due to their high performance and decreasing cost. However, in shared environments, concurrent workloads can interfere with each other, leading to a violation of Quality of Service (QoS). While QoS mechanisms like fairness guarantees and latency constraints have been integrated into SSDs, existing transaction processing frameworks offer limited QoS guarantees and can significantly degrade overall performance in a shared environment. The reason is that the internal components of an SSD, originally designed to exploit parallelism, struggle to coordinate effectively when QoS mechanisms are applied to them. This article proposes a novel QoS -enhanced transaction pro cessing framework, called QoS-pro, which enhances QoS guarantees for concurrent workloads while maintaining high parallelism for SSDs. QoS-pro achieves this by redesigning transaction processing procedures to fully exploit the parallelism of shared SSDs and enhancing QoS-oriented transaction translation and scheduling with parallelism features in mind. In terms of fairness guarantees, QoS-pro outperforms state-of-the-art methods by achieving 96% fairness improvement and 64% maximum latency reduction. QoS-pro also shows almost no loss in throughput when compared with parallelism-oriented methods. Additionally, QoS-pro triggers the fewest Garbage Collection (GC) operations and minimally affects concurrently running workloads during GC operations.
Hao Fan 0006, Yiliang Ye, Shadi Ibrahim, Xingru Li, Weibin Xue, Song Wu 0001, Chen Yu 0003, Xuanhua Shi, Hai Jin 0001
ACM Trans. Archit. Code Optim.1
2024 vKernel: Enhancing Container Isolation via Private Code and Data
abstract
Container technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead.
Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan
IEEE Trans. Computers5
2024 Multi-Grained Trace Collection, Analysis, and Management of Diverse Container Images
abstract
Container technology is getting popular in cloud environments due to its lightweight feature and convenient deployment. Container Registry plays a critical role in container-based clouds, as many container startups involve downloading layer-structured container images from Container Registry. However, Container Registry is struggling to efficiently manage images (i.e., transfer and store) with the emergence of diverse services and new image formats. The reason is that Container Registry manages images uniformly at layer granularity. On the one hand, such uniform layer-level management probably cannot fit the various requirements of different kinds of containerized services well. On the other hand, new image formats organizing data in blocks or files cannot benefit from such uniform layer-level image management. In this paper, we perform the first analysis of image traces at multiple granularities (i.e., image-, layer-, and file-level) for various services and provide an in-depth comparison of different image formats. The traces were collected from a production-level Container Registry, amounting to 24 million requests and involving more than 184 TB of transferred data. We provide a number of valuable insights, including request patterns of services, file-level access patterns, and bottlenecks associated with different image formats. Based on these insights, we propose two optimizations to improve image transfer. Both the traces and toolkit for trace collection will be open-sourced.
Qi Zhang 0009, Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers3
2023 Duo: Improving Data Sharing of Stateful Serverless Applications by Efficiently Caching Multi-Read Data
abstract
A growing number of applications are moving to serverless architectures for high elasticity and fine-grained billing. For stateful applications, however, the use of serverless architectures is likely to lead to significant performance degradation, as frequent data sharing between different execution stages involves time-consuming remote storage access. Current platforms leverage memory cache to speed up remote access. However, conventional caching strategies show limited performance improvement. We experimentally find that the reason is that current strategies overlook the stage-dependent access patterns of stateful serverless applications, i.e., data that are read multiple times across stages (denoted as multi-read data) are wrongly evicted by data that are read only once (denoted as read-once data), causing a high cache miss ratio.Accordingly, we propose a new caching strategy, Duo, whose design principle is to cache multi-read data as long as possible. Specifically, Duo contains a large cache list and a small cache list, which act as Leader list and Wingman list, respectively. Leader list ignores the data that is read for the first time to prevent itself from being polluted by massive read-once data at each stage. Wingman list inspects the data that are ignored or evicted by Leader list, and pre-fetches the data that will probably be read again based on the observation that multi-read data usually appear periodically in groups. Compared to the state-of-the-art works, Duo improves hit ratio by 1.1×-2.1× and reduces the data sharing overhead by 25%-62%.
Hao Fan 0006, Chaoyi Cheng, Song Wu 0001, Hai Jin 0001
IPDPS2
2023 QoS-Aware and Cost-Efficient Dynamic Resource Allocation for Serverless ML Workflows
abstract
Machine Learning (ML) workflows are increasingly deployed on serverless computing platforms to benefit from their elasticity and fine-grain pricing. Proper resource allocation is crucial to achieve fast and cost-efficient execution of serverless ML workflows (specially for hyperparameter tuning and model training). Unfortunately, existing resource allocation methods are static, treat functions equally, and rely on offline prediction, which limit their efficiency. In this paper, we introduce CE-scaling – a Cost-Efficient autoscaling framework for serverless ML work-flows. During the hyperparameter tuning, CE-scaling partitions resources across stages according to their exact usage to minimize resource waste. Moreover, it incorporates an online prediction method to dynamically adjust resources during model training. We implement and evaluate CE-scaling on AWS Lambda using various ML models. Evaluation results show that compared to state-of-the-art static resource allocation methods, CE-scaling can reduce the job completion time and the monetary cost by up to 63% and 41% for hyperparameter tuning, respectively; and by up to 58% and 38% for model training.
Hao Wu 0010, Junxiao Deng, Hao Fan 0006, Shadi Ibrahim, Song Wu 0001, Hai Jin 0001
IPDPS3
2022 Container-aware I/O stack: bridging the gap between container storage drivers and solid state devices
abstract
Solid State Devices (SSDs) have been widely adopted in containerized cloud platforms as they provide parallel and high-speed data accesses for critical data-intensive applications. Unfortunately, the I/O stack of the physical host overlooks the layered and independent nature of containers, thus I/O operations require expensive file redirect (between the storage driver, Overlay2/EXT4, and the virtual file system, VFS) and are scheduled sequentially. Moreover, containers suffer from significant I/O contention as resources at the native file system are shared between them. This paper presents a Container-aware I/O stack (CAST). CAST is made up of Layer-aware VFS (LaVFS) and Container-aware Native File System (CaFS). LaVFS locates files based on layer information and enables simultaneous Copy-on-Write (CoW) operations and thus avoids the overhead of searching and modifying files. CaFS, on the other hand, provides contention-free access by designing fine-grain resource allocation at the native file system. Experimental results using a NVMe SSD with micro-benchmarks and real-world applications show that CAST achieves 216%-219% (38%-98%, respectively) improvement over the original I/O stack.
Song Wu 0001, Hao Fan 0006, Shadi Ibrahim, Hai Jin 0001
VEE4
2022 Container lifecycle-aware scheduling for serverless computing
abstract
Abstract Elastic scaling in response to changes on demand is a main benefit of serverless computing. When bursty workloads arrive, a serverless platform launches many new containers and initializes function environments (known as cold starts), which incurs significant startup latency. To reduce cold starts, platforms usually pause a container after it serves a request, and reuse this container for subsequent requests. However, this reuse strategy cannot efficiently reduce cold starts because the schedulers are agnostic of container lifecycle. For example, it may ignore soon available containers or evict soon needed containers. We propose a container lifecycle‐aware scheduling strategy for serverless computing, CAS. The key idea is to control distribution of requests and determine creation or eviction of containers according to different lifecycle phases of containers. We implement a prototype of CAS on OpenWhisk. Our evaluation shows that CAS reduces 81% cold starts and therefore brings a 63% reduction at 95th percentile latency compared with native scheduling strategy in OpenWhisk when there is worker contention between workloads, and does not add significant performance overhead.
Song Wu 0001, Zhiheng Tao, Hao Fan 0006, Hai Jin 0001, Chen Yu 0003, Chun Cao
Softw. Pract. Exp.3
2021 Gear: Enable Efficient Container Storage and Deployment with a New Image Format
abstract
Containers have been widely used in various cloud platforms as they enable agile and elastic application deployment through their process-based virtualization and layered image system. However, different layers of a container image may contain substantial duplicate and unnecessary data, which slows down its deployment due to long image downloading time and increased burden on the image registry. To accelerate the deployment and reduce the size of the registry, we propose a new image format, named Gear image, that consists of two parts: a Gear index describing the structure of the image's file system and a set of files that are required when running an application. The Gear index is represented as a single-layer image compatible with the existing deployment framework. Containers can be launched by pulling a Gear index and on demand retrieving files pointed to by the index. Furthermore, the Gear image enables a file-level sharing mechanism, which helps remove duplicate data in the registry and avoid repeated downloading of identical files by a client. We implement a prototype of the container framework, named Gear, supporting the new image format. Evaluation shows that Gear saves 54 % storage capacity in the registry, speeds up container startup by up to${5\times}$, and reduces 84 % bandwidth demands.
Hao Fan 0006, Shengwei Bian, Song Wu 0001, Song Jiang 0001, Shadi Ibrahim, Hai Jin 0001
ICDCS1
2021 Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice Control
abstract
Cloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. We evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75× performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications.
Hao Fan 0006, Song Wu 0001, Zhenjiang Xie, Sheng Di, Jiang Xiao 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers1
2020 BED: A Block-Level Deduplication-Based Container Deployment Framework
Shiqiang Zhang, Song Wu 0001, Hao Fan 0006, Deqing Zou, Hai Jin 0001
GPC3
2019 NCQ-Aware I/O Scheduling for Conventional Solid State Drives
abstract
While current fairness-driven I/O schedulers are successful in allocating equal time/resource share to concurrent workloads, they ignore the I/O request queueing or reordering in storage device layer, such as Native Command Queueing (NCQ). As a result, requests of different workloads cannot have an equal chance to enter NCQ (NCQ conflict) and fairness is violated. We address this issue by providing the first systematic empirical analysis on how NCQ affects I/O fairness and SSD utilization and accordingly proposing a NCQ-aware I/O scheduling scheme, NASS. The basic idea of NASS is to elaborately control the request dispatch of workloads to relieve NCQ conflict and improve NCQ utilization. NASS builds on two core components: an evaluation model to quantify important features of the workload, and a dispatch control algorithm to set the appropriate request dispatch of running workloads. We integrate NASS into four state-of-the-art I/O schedulers and evaluate its effectiveness using widely used benchmarks and real world applications. Results show that with NASS, I/O schedulers can achieve 11-23% better fairness and at the same time improve device utilization by 9-29%.
Hao Fan 0006, Song Wu 0001, Shadi Ibrahim, Hai Jin 0001, Jiang Xiao 0001, Haibing Guan
IPDPS1
2016 iShare: Balancing I/O performance isolation and disk I/O efficiency in virtualized environments
abstract
Summary Performance isolation has long been a challenging problem for disk resource allocation in virtualized environments. While there have been many researches working on I/O performance isolation and disk utilization, none of them addresses the I/O performance isolation and disk utilization as a whole. To this end, we investigate the impact of current disk I/O performance isolation schemes on disk I/O utilization. Interestingly, our studies report that current isolation schemes bring unnecessary disk idle and reduce the overall disk I/O performance because of ignoring the disk states and characteristics of requests. Accordingly, we propose an adaptive proportional‐share I/O scheduling framework, namediShare, in virtualized environments.iSharenot only ensures I/O performance isolation through proportionally allocating time slices according to the weights of virtual machines but also preserves high disk efficiency by detecting disk states and adaptively adjusting the time slice size based on characteristics of requests. We implement a prototype ofiShareon the Xen platform. The experimental results show thatiShareensures I/O performance isolation while improving disk I/O efficiency, compared withBlkio(i.e., the default I/O performance isolation method in Xen),iShareincreases disk I/O bandwidth by 58% and slightly improves the I/O performance isolation for the sequential write applications. Copyright © 2015 John Wiley & Sons, Ltd.
Song Wu 0001, Songqiao Tao, Hao Fan 0006, Hai Jin 0001, Shadi Ibrahim
Concurr. Comput. Pract. Exp.4