Ning Li 0010

dblp:14/5410-10 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-6510-1687ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 REEF: Energy-Efficient, Application-QoS-Aware Thread Processing in Oversubscribed Server Environments
abstract
Modern storage-intensive server applications such as key-value stores and relational databases often rely on thread oversubscription to sustain high throughput in cloud environments. While effective at hiding I/O stalls, this practice introduces serious challenges, including unpredictable query behaviors, instruction-per-query (IPQ) inflation, and instability in dynamic power management (DPM). Existing quality-of-service (QoS)-centric or energy-centric techniques, developed in isolation at one of the service layers, fail to holistically optimize resource and energy efficiency under connection-level QoS constraints. This paper presents REEF (Resource- and Energy-Efficient user-space scheduling Framework), a non-intrusive, cross-layer approach that coordinates query processing across the network stack, application layer, and OS resource manager. REEF transforms self-serving threads into on-call threads activated by near-optimal, proactively batched scheduling, enabling deep CPU C-state residency and mitigating CPU and I/O contention. Our extensive evaluation on real server applications (MongoDB and MySQL) demonstrates that REEF can substantially improve the energy efficiency of server applications under different connection-level QoS schemes, by up to 53.45% for throughput per power, up to 2.18× and 5.33× for coefficient of P99 and P99.9 tail-latency per power respectively, and significantly reduce resource consumption, by up to 73.65% in CPU frequency.
Ning Li 0010, Hong Jiang 0001, Hao Che, Zhijun Wang 0001
SoCC1
2023 User Disengagement-Oriented Target Enforcement for Multi-Tenant Database Systems
abstract
Unexpected long query latency of a database system can cause domino effects on all the upstream services and severely degrade end users' experience with unpredicted long waits, resulting in an increasing number of users disengaged with the services and thus leading to a high user disengagement ratio (UDR). A high UDR usually translates to reduced revenue for service providers. This paper proposes UTSLO, a UDR-oriented SLO guaranteed system, which enables a database system to support multi-tenant UDR targets in a cost-effective fashion through UDR-oriented capacity planning and dynamic UDR target enforcement. The former aims to estimate the feasibility of UDR targets while the latter dynamically tracks and regulates per-connection query latency distribution needed for accurate UDR target guarantee. In UTSLO, the database service capacity can be fully exploited to efficiently accommodate tenants while minimizing resources required for UDR target guarantee.
Ning Li 0010, Hong Jiang 0001, Hao Che, Zhijun Wang 0001, Minh Nguyen 0003, Todd Rosenkrantz
SoCC1
2022 Improving scalability of database systems by reshaping user parallel I/O
abstract
Modern database systems suffer from compromised throughput, persistent unfair I/O processing and unpredictable, high latency variability of user requests as a result of mismatches between highly scaled user parallel I/O and the I/O capacity afforded by the database and its underlying storage I/O stack. To address this problem, we introduce an efficient user-centric QoS-aware scheduling shim, called AppleS, for user-level fine-grained I/O regulation that delivers the right amount and pattern of user parallel I/O requests to the database system and supports user SLOs with high-level performance isolation and reduced I/O resource contention. It is designed to enable database systems to proactively regulate user request behaviors based on runtime conditions to reshape user access pattern to hide excessive user parallelism from the I/O stack that has a limited concurrent processing capability. This helps achieve scalable throughput for multi-user workloads in a fair and stable manner. AppleS is implemented as a user-space shim for transparent user-differentiated I/O scheduling, making it highly flexible and portable. Our extensive evaluation, run on real databases (MySQL and MongoDB), demonstrates that, by incorporating AppleS in the existing database systems, our solution can not only improve the throughput (up to 39.2%) in a fairer (3.2× to 40.6× fairness improvement) and more stable (up to 2× lower latency variability) manner, but also support user SLOs with less I/O provisioning.
Ning Li 0010, Hong Jiang 0001, Hao Che, Zhijun Wang 0001, Minh Nguyen 0003
EuroSys1
2021 An Incast-Coflow-Aware Minimum-Rate-Guaranteed Congestion Control Protocol for Datacenter Applications
abstract
Today s datacenters need to meet service level objectives (SLOs) for applications, which can be translated into deadlines for (co)flows running between job execution stages. As a result, meeting (co)flow deadlines with high probabilities is essential to attract and retain customers and hence, generate high revenue. To fill the lack of a transport protocol that can facilitate low (co)flow deadline miss rate, especially in the face of incast congestion, in this paper, we propose DCMRG, an incast-coflow-aware, ECN-based soft minimum-rate-guaranteed congestion control protocol for datacenter applications. DCMRG is composed of two major components, i.e., a congestion controller running on the send host and an incast congestion controller running on the receive host. DCMRG possesses three salient features. First, it is the first congestion control protocol that integrates congestion control with coflow-aware incast control while providing soft minimum flow rate guarantee. Second, DCMRG is readily deployable in datacenter networks. It only requires software upgrade in the hosts and minimum assistance (i.e., ECN) from in-network nodes. Third, DCMRG is backward compatible with and, by design, friendly to the widely deployed, standard-based transport protocols, such as DCTCP. The results from large-scale datacenter network simulation demonstrate that in the absence of incast congestion, DCMRG can reduce flow deadline miss rates by 3x and 1.6x compared to D2TCP and MRG, respectively. Moreover, DCMRG further reduces the coflow deadline miss rate by more than 40% and 60% and lowers the packet drop probability by 60% and 80%, in the face of incast congestion, compared to D2TCP with ICTCP and MRG with ICTCP, respectively.
Zhijun Wang 0001, Yunxiang Wu, Stoddard Rosenkrantz, Ning Li 0010, Minh Nguyen 0003, Hao Che
NAS4
2020 Optimal Encoding and Decoding Algorithms for the RAID-6 Liberation Codes
abstract
RAID-6 is gradually replacing RAID-5 as the dominant form of disk arrays due to its capability of tolerating concurrent failures of any two disks, as well as the case of encountering an uncorrectable read error during recovery. Implementing a RAID-6 system relies on some erasure coding schemes, and so far the most representative solutions are EVENODD codes [1], RDP codes [2] and Liberation codes [3], none of which has emerged as a clear "all-around" winner. In this paper, we are interested in revealing the undiscovered potential of the Liberation codes, since these codes have the following attractive features: (a) they have the best update performance, (b) they have better scalability, and (c) they are open-sourced and publicly available, as well as the following drawbacks: fair encoding performance and, more importantly, relatively poor decoding performance. Specificly, we present novel optimal encoding and decoding algorithms for the Liberation codes by introducing an alternative, geometric presentation of these codes. The proposed algorithms completely eliminate redundant computations during the encoding and decoding procedures by extracting and reusing common expressions between the two types of parity constraints, and do not involve any matrix operations on which the original algorithms are based. Our experiment results show that compared with the original solution, the proposed encoding and decoding algorithms reduce the number of XOR's by up to 16 percent and 15 ~20 percent respectively, and the encoding and decoding throughputs are increased by 22.3 percent and at most 155 percent respectively. Moreover, the encoding complexity reaches the theoretical lower bound, while the decoding complexity is also very close to the theoretical lower bound.
Hong Jiang 0001, Zhirong Shen, Hao Che, Nong Xiao 0001, Ning Li 0010
IPDPS6
2020 A Black-Box Fork-Join Latency Prediction Model for Data-Intensive Applications
abstract
The workflows of the predominant datacenter services are underlaid by various Fork-Join structures. Due to the lack of good understanding of the performance of Fork-Join structures in general, today's datacenters often operate under low resource utilization to meet stringent service level objectives (SLOs), e.g., in terms of tail and/or mean latency, for such services. Hence, to achieve high resource utilization, while meeting stringent SLOs, it is of paramount importance to be able to accurately predict the tail and/or mean latency for a broad range of Fork-Join structures of practical interests. In this article, we propose a black-box Fork-Join model that covers a wide range of Fork-Join structures for the prediction of tail and mean latency, called ForkTail and ForkMean, respectively. We derive highly computational effective, empirical expressions for tail and mean latency as functions of means and variances of task response times. Our extensive testing results based on model-based and trace-driven simulations, as well as a real-world case study in a cloud environment demonstrate that the models can consistently predict the tail and mean latency within 20 and 15 percent prediction errors at 80 and 90 percent load levels, respectively, for heavy-tailed workloads, and at any load levels for light-tailed workloads. Moreover, our sensitivity analysis demonstrates that such errors can be well compensated for with no more than 7 percent resource overprovisioning. Consequently, the proposed prediction model can be used as a powerful tool to aid the design of tail-and-mean-latency guaranteed job scheduling and resource provisioning, especially at high load, for datacenter applications.
Minh Nguyen 0003, Sami Alesawi, Ning Li 0010, Hao Che, Hong Jiang 0001
IEEE Trans. Parallel Distributed Syst.3
2019 Efficient MDS Array Codes for Correcting Multiple Column Erasures
abstract
The RΛ-Code is an efficient family of maximum distance separable (MDS) array codes of column distance 4, which involves two types of parity constraints: the row parity and the Λ parity formed by diagonal lines of slopes 1 and -1. Benefitting from the common expressions between the two parity constraints, the encoding and decoding complexities are distinctly lower than most (if not all) of other triple-erasure-correcting codes. It was left as an open problem generalizing the RΛ-Code to arbitrary column distances. In this paper, we present such a generalization, namely, we construct a family of MDS array codes being capable of correcting any prescribed number of erasures/errors by introducing multiple Λ parity constraints. Essentially, the generalized RΛ-Code is derived from a certain variant of the Blaum-Roth codes, and hence retains the error/erasure correcting capability of the latter. Compared with the Blaum-Roth codes, the generalized RΛ-Code has two advantages: a) by exploiting common expressions between row parity and different Λ parity constraints, and reusing the intermidate results during the syndrome calculations, it can encode and decode faster; and b) the memory footprint during encoding/decoding, and the I/O cost caused by degraded reads, are both reduced by 50%.
Hong Jiang 0001, Hao Che, Nong Xiao 0001, Ning Li 0010
ISIT5
2019 Storage Sharing Optimization Under Constraints of SLO Compliance and Performance Variability
abstract
SLO enforcement with the required strong SLO compliance and the desired low level of performance variability is necessary to ensure QoS for user applications with precisely differentiated service levels. However, for the cloud consolidating a large number of VMs rented by users, it is a great challenge to formulate an IO capacity allocation among consolidated VMs under the user-customized QoS constraints of SLO compliance and performance fluctuation for consolidated VMs. To address this challenge, we propose SASLO, an end-to-end VM-oriented control framework that supports users in customizing SLO targets and QoS constraints for each VM. SASLO can dynamically coordinate the throughput target and IO size limit for each VM adapting to the status of SLO enforcement so as to maximize the IO capacity allocation among consolidated VMs under QoS constraints. To accurately enforce time-varying throughput target, SASLO establishes a proportional-integral IO controller for each individual VM to converge the actual throughput to the target with an expected settling time. Our extensive evaluation driven by representative benchmarks demonstrates that SASLO is able to formulate a satisfactory IO capacity allocation plan for consolidated VMs under the constraints of SLO compliance and performance variability.
Ning Li 0010, Hong Jiang 0001, Dan Feng 0001, Zhan Shi 0001
IEEE Trans. Serv. Comput.1
2018 ForkTail: a black-box fork-join tail latency prediction model for user-facing datacenter workloads
abstract
The workflows of the predominant user-facing datacenter services, including web searching and social networking, are underlaid by various Fork-Join structures. Due to the lack of understanding the performance of Fork-Join structures in general, today's datacenters often resort to resource overprovisioning, operating under low resource utilization, to meet stringent tail-latency service level objectives (SLOs) for such services. Hence, to achieve high resource utilization, while meeting stringent tail-latency SLOs, it is of paramount importance to be able to accurately predict the tail latency for a broad range of Fork-Join structures of practical interests.
Minh Nguyen 0003, Sami Alesawi, Ning Li 0010, Hao Che, Hong Jiang 0001
HPDC3
2017 Customizable SLO and Its Near-Precise Enforcement for Storage Bandwidth
abstract
Cloud service is being adopted as a utility for large numbers of tenants by renting Virtual Machines (VMs). But for cloud storage, unpredictable IO characteristics make accurate Service-Level-Objective (SLO) enforcement challenging. As a result, it has been very difficult to support simple-to-use and technology-agnostic SLO specifying a particular value for a specific metric (e.g., storage bandwidth). This is because the quality of SLO enforcement depends on performance error and fluctuation that measure the precision of SLO enforcement . High precision of SLO enforcement is critical for user-oriented performance customization and user experiences. To address this challenge, this article presents V-Cup, a framework for VM-oriented customizable SLO and its near-precise enforcement. It consists of multiple auto-tuners, each of which exports an interface for a tenant to customize the desired storage bandwidth for a VM and enable the storage bandwidth of the VM to converge on the target value with a predictable precision. We design and implement V-Cup in the Xen hypervisor based on the fair sharing scheduler for VM-level resource management. Our V-Cup prototype evaluation shows that it achieves satisfying performance guarantees through near-precise SLO enforcement.
Ning Li 0010, Hong Jiang 0001, Dan Feng 0001, Zhan Shi 0001
ACM Trans. Storage1
2016 PSLO: enforcing the Xth percentile latency and throughput SLOs for consolidated VM storage
abstract
It is desirable but challenging to simultaneously support latency SLO at a pre-defined percentile, i.e., the Xth percentile latency SLO, and throughput SLO for consolidated VM storage. Ensuring the Xth percentile latency contributes to accurately differentiating service levels in the metric of the application-level latency SLO compliance, especially for the application built on multiple VMs. However, the Xth percentile latency SLO and throughput SLO enforcement are the opposite sides of the same coin due to the conflicting requirements for the level of IO concurrency. To address this challenge, this paper proposes PSLO, a framework supporting the Xth percentile latency and throughput SLOs under consolidated VM environment by precisely coordinating the level of IO concurrency and arrival rate for each VM issue queue. It is noted that PSLO can take full advantage of the available IO capacity allowed by SLO constraints to improve throughput or reduce latency with the best effort. We design and implement a PSLO prototype in the real VM consolidation environment created by Xen. Our extensive trace-driven prototype evaluation shows that our system is able to optimize the Xth percentile latency and throughput for consolidated VMs under SLO constraints.
Ning Li 0010, Hong Jiang 0001, Dan Feng 0001, Zhan Shi 0001
EuroSys1