Huiba Li

dblp:48/9032 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
9since 2021 · last 2026
0009-0000-1344-6552ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 5 first-author · 8 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 ParaSync: Exploiting Fine-Grained Parallelism for Efficient File Synchronization
Lu Tang 0004, Huiba Li, Yue Yu 0001, Guangtao Xue, Jiwu Shu, Yiming Zhang 0003
FAST3
2026 SkySync: Accelerating File Synchronization with Collaborative Delta Generation
Huiba Li, Lu Tang 0004, Guangtao Xue, Jiwu Shu, Yiming Zhang 0003
FAST2
2026 PlanB: Efficient Software IPv6 Lookup with Linearized B+-Tree
Lanzheng Liu, Huiba Li, Jiwu Shu, Windsor W. Hsu, Yiming Zhang 0003
NSDI4
2026 zBuffer: Zero-Copy and Metadata-Free Serialization for Fast RPC with Scatter-Gather Reflection
abstract
This paper presents zBuffer, a zero-copy and metadata-free serialization library for high-performance and low-cost RPCs. At the core of zBuffer is scatter-gather reflection, a novel technique that collaboratively (i) leverages the NIC scatter-gather hardware feature to offload the costly data coalescing, and (ii) utilizes the static reflection mechanism of modern programming languages to enable type queries on complex data objects without requiring explicit metadata construction. We leverage C++ language features, mainly including template meta-programming and macros, to realize static reflection at compile time. Based on zBuffer, we design a fast RPC system (called zRPC) which eliminates all RPC memory copy overheads not only in (de)serialization but also in network transmission. Extensive evaluation shows that zBuffer/zRPC significantly outperforms state-of-the-art serialization/RPC mechanisms: zBuffer is approximately 7× faster than Cornflakes in serialization for complex data objects; and zRPC reduces 99th percentile latency by 21% and achieves 62% higher throughput than eRPC on the Masstree key-value (KV) store with the YCSB benchmark.
Huiba Li, Shun Gai, Youmin Chen, Yiming Zhang 0003
PPoPP2
2024 xMeta: SSD-HDD-hybrid Optimization for Metadata Maintenance of Cloud-scale Object Storage
abstract
Object storage has been widely used in the cloud. Traditionally, the size of object metadata is much smaller than that of object data, and thus existing object storage systems (such as Ceph and Oasis) can place object data and metadata, respectively, on hard disk drives (HDDs) and solid-state drives (SSDs) to achieve high I/O performance at a low monetary cost. Currently, however, a wide range of cloud applications organize their data as large numbers of small objects of which the data size is close to (or even smaller than) the metadata size, thus greatly increasing the cost if placing all metadata on expensive SSDs. This article presents x Meta , an SSD-HDD-hybrid optimization for metadata maintenance of cloud-scale object storage. We observed that a substantial portion of the metadata of small objects is rarely accessed and thus can be stored on HDDs with little performance penalty. Therefore, x Meta first classifies the hot and cold metadata based on the frequency of metadata accesses of upper-layer applications and then adaptively stores the hot metadata on SSDs and the cold metadata on HDDs. We also propose a merging mechanism for hot metadata to further improve the efficiency of SSD storage and optimize range key query and insertion for hot metadata by designing composite keys. We have integrated the x Meta metadata service with Ceph to realize a high-performance, low-cost object store (called xCeph). The extensive evaluation shows that xCeph outperforms the original Ceph by an order of magnitude in the space requirement of SSD storage, while improving the throughput by up to 2.7×.
Qiwen Ke, Huiba Li, Yongwei Wu 0001, Yiming Zhang 0003
ACM Trans. Archit. Code Optim.3
2024 Efficient Block Storage in the Cloud
abstract
This paper presents URSAL, an HDD-only block storage system that achieves ultra-efficiency, reliability, scalability and availability at low cost. Compared to existing block stores such as URSA, Ceph, and Sheepdog, URSAL has the following distinctions. First, since parallelism is harmful to the random I/O performance on HDDs, we restrict URSAL storage servers to conservatively perform parallel I/O on HDDs for avoiding I/O contention and reducing tail latency. Second, URSAL designs a proxy-based storage architecture to separate the high-level and low-level I/O logic, where for each virtual machine (VM) there is one URSAL proxy process running at the client VM side to control (at a high level) the procedure of server-side low-level I/O. Third, to alleviate the problem of low random write performance of HDDs, URSAL selectively performs direct block writes on raw HDDs or indirect log appends to HDD journals (which are then asynchronously replayed to raw HDDs), depending on the characteristics of the workloads. Fourth, software failures are nontrivial in large-scale block storage systems of which the availability is vital to client VMs, and thus for high availability we design an efficient fault-tolerance mechanism by isolating the connection management module of URSAL proxy. We have implemented URSAL and deployed it at scale. Extensive evaluation results demonstrate that URSAL achieves much higher performance than the state-of-the-art solutions for underloaded scenarios.
Yiming Zhang 0003, Huiba Li, Ping Zhong 0002, Shengyun Liu, Dongsheng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Block-level Image Service for the Cloud
abstract
Businesses increasingly need agile and elastic computing infrastructure to respond quickly to real-world situations. By offering efficient process-based virtualization and a layered image system, containers are designed to enable agile and elastic application deployment. However, creating or updating large container clusters is still slow due to the image downloading and unpacking process. In this article, we present DADI Image Service (DADI), a block-level image service for increased agility and elasticity in deploying applications. DADI replaces the waterfall model of starting containers (downloading image, unpacking image, starting container) with fine-grained on-demand transfer of remote images, realizing instant start of containers. To accelerate the cold start of containers, DADI designs a pull-based prefetching mechanism that allows a host to read necessary image data beforehand at the granularity of image layers. We design a peer-to-peer–based decentralized image sharing architecture to balance traffic among all the participating hosts and propose a pull-push collaborative prefetching mechanism to accelerate cold start. DADI efficiently supports various kinds of runtimes including cgroups, QEMU, and so on, further realizing “build once, run anywhere.” DADI has been deployed at scale in the production environment of Alibaba, serving one of the world’s largest ecommerce platforms. Performance results show that DADI can cold start 10,000 containers on 1,000 hosts within 4 s.
Huiba Li, Lanzheng Liu, Yiming Zhang 0003, Windsor W. Hsu
ACM Trans. Storage1
2023 Hybrid Block Storage for Efficient Cloud Volume Service
abstract
The migration of traditional desktop and server applications to the cloud brings challenge of high performance, high reliability, and low cost to the underlying cloud storage. To satisfy the requirement, this article proposes a hybrid cloud-scale block storage system called Ursa . Trace analysis shows that the I/O patterns served by block storage have only limited locality to exploit. Therefore, instead of using solid state drives (SSDs) as a cache layer, Ursa proposes hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on hard disk drives (HDDs) . At the core of Ursa ’s hybrid storage design is an adaptive journal that can bridge the performance gap between primary SSDs and backup HDDs for random writes by transforming small backup writes into journal appends, which are then asynchronously replayed and merged to backup HDDs. To efficiently index the journal, we design a novel range-optimized merge-tree structure that combines a continuous range of keys into a single composite key {offset,length} . Ursa integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that Ursa in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (IOPS and throughput per core).
Yiming Zhang 0003, Huiba Li, Shengyun Liu, Peng Huang 0005
ACM Trans. Storage2
2021 FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function Compute
Shuai Chang, Huangshi Tian, Huiba Li, Yue Cheng 0001
USENIX ATC6
2020 URSAL: Ultra-Efficient, Reliable, Scalable, and Available Block Storage at Low Cost
abstract
Large-scale cloud block storage provides virtual disks for various applications and services like online booking, gaming, and offline data analytics. The state-of-the-art URSA [1] block store adopted a hybrid storage structure which placed primary data on solid-state drives (SSDs) and stored backup data on hard-disk drives (HDDs). URSA used small SSD journals to bridge the performance gap between SSDs and HDDs. Although URSA's SSD-HDD-hybrid storage structure achieves SSD-like I/O performance while using only one third of the SSDs required by the SSD-only storage pattern (storing both primary data and backup data on SSDs), we argue that the traditional HDDonly storage structure is still preferable for a large variety of relatively low-end customers and underloaded applications that are sensitive to the per-bit storage cost.To lower the storage cost, in this paper we design URSAL, an HDD-only block store which provides ultra efficiency, reliability, scalability and availability at low cost. Compared to existing block stores such as URSA, Ceph, and Sheepdog, URSAL has the following distinctions. First, URSAL designs the proxy-based storage architecture, where a proxy server process runs together with each virtual machine (VM) client mounting virtual disks and controls the procedure of all block-level I/O. Second, URSAL selectively performs direct block writes on raw HDDs or indirect log appends to HDD journals (which are then asynchronously replayed to raw HDDs), depending on the characteristics of the workloads. Third, URSAL runs one storage server process for each physical HDD, which conservatively has at most one active thread reading/writing the HDD to avoid I/O contention. We have implemented URSAL. Evaluation results show that URSAL significantly outperforms state-of-the-art HDD-only block stores (Ceph and Sheepdog) when providing virtual disks for underloaded applications.
Huiba Li, Yiming Zhang 0003, Ping Zhong 0002
INFOCOM1
2020 DADI: Block-Level Image Service for Agile and Elastic Application Deployment
Huiba Li, Lanzheng Liu, Windsor W. Hsu
USENIX ATC1
2020 PBS: An Efficient Erasure-Coded Block Storage System Based on Speculative Partial Writes
abstract
Block storage provides virtual disks that can be mounted by virtual machines (VMs). Although erasure coding (EC) has been widely used in many cloud storage systems for its high efficiency and durability, current EC schemes cannot provide high-performance block storage for the cloud. This is because they introduce significant overhead to small write operations (which perform partial write to an entire EC group), whereas cloud-oblivious applications running on VMs are often small-write-intensive. We identify the root cause for the poor performance of partial writes in state-of-the-art EC schemes: for each partial write, they have to perform a time-consuming write-after-read operation that reads the current value of the data and then computes and writes the parity delta, which will be used to “patch” the parity in journal replay. In this article, we present a speculative partial write scheme (called P ARI X) that supports fast small writes in erasure-coded storage systems. We transform the original formula of parity calculation to use the data deltas (between the current/original data values), instead of the parity deltas, to calculate the parities in journal replay. For each partial write, this allows P ARI X to speculatively log only the new value of the data without reading its original value. For a series of n partial writes to the same data, P ARI X performs pure write (instead of write-after-read) for the last n -1 ones while only introducing a small penalty of an extra network round-trip time to the first one. Based on P ARI X, we design and implement P ARI X Block Storage (PBS), an efficient block storage system that provides high-performance virtual disk service for VMs running cloud-oblivious applications. PBS not only supports fast partial writes but also realizes efficient full writes, background journal replay, and fast failure recovery with strong consistency guarantees. Both microbenchmarks and trace-driven evaluation show that PBS provides efficient block storage and outperforms state-of-the-art EC-based systems by orders of magnitude.
Yiming Zhang 0003, Huiba Li, Shengyun Liu, Guangtao Xue
ACM Trans. Storage2
2019 URSA: Hybrid Block Storage for Cloud-Scale Virtual Disks
abstract
This paper presents URSA, a hybrid block store that provides virtual disks for various applications to run efficiently on cloud VMs. Trace analysis shows that the I/O patterns served by block storage have limited locality to exploit. Therefore, instead of using SSDs as a cache layer, URSA proposes an SSD-HDD-hybrid storage structure that directly stores primary replicas on SSDs and replicates backup replicas on HDDs, using journals to bridge the performance gap between SSDs and HDDs. URSA integrates the hybrid structure with designs for high reliability, scalability, and availability. Experiments show that URSA in its hybrid mode achieves almost the same performance as in its SSD-only mode (storing all replicas on SSDs), and outperforms other block stores (Ceph and Sheepdog) even in their SSD-only mode while achieving much higher CPU efficiency (performance per core). We also discuss some practical issues in our deployment.
Huiba Li, Yiming Zhang 0003, Dongsheng Li 0001, Shengyun Liu, Peng Huang 0005, Zheng Qin 0002, Kai Chen 0005, Yongqiang Xiong
EuroSys1
2018 KylinX: A Dynamic Library Operating System for Simplified and Efficient Cloud Virtualization
Yiming Zhang 0003, Jon Crowcroft, Dongsheng Li 0001, Chengfen Zhang, Huiba Li, Yaozheng Wang, Yongqiang Xiong, Guihai Chen
USENIX ATC5
2017 PARIX: Speculative Partial Writes in Erasure-Coded Systems
Huiba Li, Yiming Zhang 0003, Shengyun Liu, Dongsheng Li 0001, Yuxing Peng 0001
USENIX ATC1
2017 Delay-bounded skyline computing for large-scale real-time online data analytics
abstract
Summary The proliferation of Internet applications, cloud systems, and mobile social networks results in unprecedented data set scale and high data generation rate. For us to be able to extract any meaningful information, it is important to achieve real‐time online data analytics. Skyline queries are important in many online data applications such as real‐time Web mining, multipreference analysis, and decision making. Most existing studies mainly focus on centralized systems, and distributed skyline query processing is still an emerging and challenging topic. In this paper, we propose SkyStorm, a delay‐bounded parallel skyline computing approach for large‐scale real‐time data analytics by dividing the search into multiple rounds and limiting the search in each round within a budget‐restricted range. The effectiveness of our proposals is demonstrated through analysis and simulations.
Yiming Zhang 0003, Huiba Li, Ping Zhong 0002
Concurr. Comput. Pract. Exp.4
2014 RAFlow: Read Ahead Accelerated I/O Flow through Multiple Virtual Layers
abstract
Virtualization is the foundation for cloud computing, and the virtualization can not be achieved without software defined, elastic, flexible and scalable virtual layers. Unfortunately, if multiple virtual storage devices are chained together, the system may be subject to severe performance degradation. While the read-ahead (RA) mechanism in storage devices plays a very important role to improve I/O performance, RA may not be effective as expected for multiple virtualization layers, since it is originally designed for one layer only. When I/O requests are passed through a long I/O path, they may trigger a chain reaction and lead to unnecessary data transmission and thus bandwidth waste. In this paper, we study the dynamic behavior of RA through multiple I/O layers and demonstrate that if controlled well, RA can greatly accelerate I/O speed. We present RAFlow, a RA control mechanism, to effectively improve I/O performance by strategically expanding RA window at each layer. Our real-world experiments show that it can achieve 20% to 50% performance improvement in I/O paths with up to 8 virtualized storage devices.
Zhaoning Zhang 0001, Kui Wu 0001, Huiba Li, Jinghua Feng, Yuxing Peng 0001, Xicheng Lu
NAS3
2014 VMThunder: Fast Provisioning of Large-Scale Virtual Machine Clusters
abstract
Infrastructure as a service (IaaS) allows users to rent resources from the Cloud to meet their various computing requirements. The pay-as-you-use model, however, poses a nontrivial technical challenge to the IaaS cloud service providers: how to fast provision a large number of virtual machines (VMs) to meet users' dynamic computing requests? We address this challenge with VMThunder, a new VM provisioning tool, which downloads data blockson demandduring the VM booting process and speeds up VM image streaming by strategically integrating peer-to-peer (P2P) streaming techniques with enhanced optimization schemes such as transfer on demand, cache on read, snapshot on local, and relay on cache. In particular, VMThunder stores the original images in a share storage and in the meantime it adopts a tree-based P2P streaming scheme so that common image blocks are cached and reused across the nodes in the cluster. We implement VMThunder in CentOS Linux and thoroughly test its performance. Comprehensive experimental results show that VMThunder outperforms the state-of-the-art VM provisioning methods, with respect to scalability, latency, and VM runtime I/O performance.
Zhaoning Zhang 0001, Ziyang Li 0003, Kui Wu 0001, Dongsheng Li 0001, Huiba Li, Yuxing Peng 0001, Xicheng Lu
IEEE Trans. Parallel Distributed Syst.5
2013 OPTAS: Optimal Data Placement in MapReduce
abstract
The data placement strategy greatly affects the efficiency of MapReduce. The current strategy only takes the map phase into account to optimize the map time. But the ignored shuffle phase may increase the total running time significantly in many jobs. We propose a new data placement strategy, named OPTAS, which optimizes both the map and shuffle phases to reduce their total time. However, the huge search space makes it difficult to find out an optimal data placement instance (DPI) rapidly. To address this problem, an algorithm is proposed which can prune most of the search space and find out an optimal result quickly. The search space firstly is segmented in ascending order according to the potential map time. Within each segment, we propose an efficient method to construct a local optimal DPI with the minimal total time of both the map and shuffle phases. To find the global optimal DPI, we scan the local optimal DPIs in order. We have proven that the global optimal DPI can be found as the first local optimal DPI whose total time stops decreasing, thus further pruning the search space. In practice, we find that at most fourteen local optimal DPIs are scanned in tens of thousands of segments with the pruning strategy. Extensive experiments with real trace data verify not only the theoretic analysis of our pruning strategy and construction method but also the optimality of OPTAS. The best improvements obtained in our experiments can be over 40% compared with the existing strategy used by MapReduce.
Yongrui Qin, Zhen Huang 0006, Yuxing Peng 0001, Dongsheng Li 0001, Huiba Li
ICPADS6
2010 Nexus: Speculative Execution for Event-Driven Networking Programs
abstract
The efficiency of communication is a key factor to the performance of networking applications, and concurrent communication is an important approach to the efficiency of communication. However, many concurrency opportunities are very difficult to exploit because they depend on some undeterministic conditions. If these conditions are highly predictable, speculative execution can be a very effective approach to cope with the uncertainties. Existing researches on speculation seldom target at networking systems, and none of them can handle the event-driven model that is very popular in such systems. In this paper, we propose Nexus, a novel speculation scheme that supports event-driven networking applications. Nexus analyzes the dependence relationship of events, and performs speculation according to the duality of events and threads. Evaluation on a prototype implementation of nexus shows that this approach can significantly reduces the time needed to complete an event-driven program.
Huiba Li, Xicheng Lu, Yuxing Peng 0001
ICPADS1
2010 Automatic Concurrency Management for distributed applications
abstract
Building distributed applications is difficult mostly because of concurrency management. Existing approaches primarily include events and threads. Researchers and developers have been debating for decades to prove which is superior. Although the conclusion is far from obvious, this long debate clearly shows that neither of them is perfect. One of the problems is that they are both complex and error-prone. Both events and threads need the programmers to explicitly manage concurrency, and we believe it is just the source of difficulties. In this paper, we propose a novel approach—automatic concurrency management by the runtime system. It dynamically analyzes the programs to discover potential concurrency opportunities; and it dynamically schedules the communication and the computation tasks, resulting in automatic concurrent execution. This approach is inspired by the instruction scheduling technologies used in modern microprocessors, which dynamically exploits instruction-level parallelism. However, hardware scheduling algorithms do not fit software in many aspects, thus we have to design a new scheme completely from scratch. automatic concurrency management is a runtime technique with no modification to the language, compiler or byte code, so it is good at backward compatibility. It is essentially a dynamic optimization for networking programs.
Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001
ISCC1
2010 Superscalar communication: A runtime optimization for distributed applications
Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001, Hangjun Zhou, Xicheng Lu
Sci. China Inf. Sci.1