Ailidani Ailijiang

dblp:185/7102 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0002-3234-4081ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Distributed systems · 76% Performance modeling and evaluation · 19% Processor architecture and microarchitecture · 6%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
consensus
2.552026
How to Evaluate Distributed Coordination Systems?-A Survey and Analysis · IEEE Trans. Parallel Distributed Syst. 2026
Scaling Replicated State Machines with Compartmentalization · Proc. VLDB Endow. 2021
PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus · SIGMOD Conference 2021
Distributed systems
distributed coordination
1.122026
How to Evaluate Distributed Coordination Systems?-A Survey and Analysis · IEEE Trans. Parallel Distributed Syst. 2026
Retroscope: Retrospective Monitoring of Distributed Systems · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
benchmarking
1.012026
How to Evaluate Distributed Coordination Systems?-A Survey and Analysis · IEEE Trans. Parallel Distributed Syst. 2026
Distributed systems › consensus
paxos
0.922021
PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus · SIGMOD Conference 2021
WPaxos: Wide Area Network Flexible Consensus · IEEE Trans. Parallel Distributed Syst. 2020
Distributed systems
replication
0.922021
PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus · SIGMOD Conference 2021
WPaxos: Wide Area Network Flexible Consensus · IEEE Trans. Parallel Distributed Syst. 2020
Processor architecture and microarchitecture
compartmentalization
0.512021
Scaling Replicated State Machines with Compartmentalization · Proc. VLDB Endow. 2021
Distributed systems › replication
state machine replication
0.512021
Scaling Replicated State Machines with Compartmentalization · Proc. VLDB Endow. 2021
Distributed systems › observability
distributed monitoring
0.412019
Retroscope: Retrospective Monitoring of Distributed Systems · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
queueing models
0.412019
Dissecting the Performance of Strongly-Consistent Replication Protocols · SIGMOD Conference 2019
Performance modeling and evaluation › benchmarking
distributed system benchmarking
0.312026
How to Evaluate Distributed Coordination Systems?-A Survey and Analysis · IEEE Trans. Parallel Distributed Syst. 2026
Distributed systems
fault tolerance
0.312026
How to Evaluate Distributed Coordination Systems?-A Survey and Analysis · IEEE Trans. Parallel Distributed Syst. 2026

Methods — techniques the papers use, named apart from their topics

survey · 1.0analysis · 1.0simulation · 0.8queueing theory · 0.8prototyping · 0.8piggybacking · 0.5compartmentalization · 0.5communication aggregation · 0.5object stealing · 0.4flexible quorums · 0.4
YearPublicationVenuePosition
2026 How to Evaluate Distributed Coordination Systems?-A Survey and Analysis
abstract
Coordination services and protocols are critical components of distributed systems and are essential for providing consistency, fault tolerance, and scalability. However, due to the lack of standard benchmarking and evaluation tools for distributed coordination services, coordination service developers/researchers either use a NoSQL standard benchmark and omit evaluating consistency, distribution, and fault tolerance; or create their own ad-hoc microbenchmarks and skip comparability with other services. In this study, we analyze and compare the evaluation mechanisms for known and widely used consensus algorithms, distributed coordination services, and distributed applications built on top of these services. We identify the most important requirements of distributed coordination service benchmarking, such as the metrics and parameters for the evaluation of the performance, scalability, availability, and consistency of these systems. Finally, we discuss why the existing benchmarks fail to address the complex requirements of distributed coordination system evaluation.
Bekir O. Turkkan, Elvis Rodrigues, Tevfik Kosar, Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
IEEE Trans. Parallel Distributed Syst.5
2021 PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus
abstract
Strongly consistent replication helps keep application logic simple and provides significant benefits for correctness and manageability. Unfortunately, the adoption of strongly-consistent replication protocols has been curbed due to their limited scalability and performance. To alleviate the leader bottleneck in strongly-consistent replication protocols, we introduce Pig, an in-protocol communication aggregation and piggybacking technique. Pig employs randomly selected nodes from follower subgroups to relay the leader's message to the rest of the followers in the subgroup, and to perform in-network aggregation of acknowledgments back from these followers. By randomly alternating the relay nodes across replication operations, Pig shields the relay nodes as well as the leader from becoming hotspots and improves throughput scalability. We showcase Pig in the context of classical Paxos protocols employed for strongly consistent replication by many cloud computing services and databases. We implement and evaluate PigPaxos, in comparison to Paxos and EPaxos protocols under various workloads over clusters of size 5 to 25 nodes. We show that the aggregation at the relay has little latency overhead, and PigPaxos can provide more than 3 folds improved throughput over Paxos and EPaxos with little latency deterioration. We support our experimental observations with the analytical modeling of the bottlenecks and show that the communication bottlenecks are minimized when employing only one randomly rotating relay node.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
SIGMOD Conference2
2021 Scaling Replicated State Machines with Compartmentalization
abstract
State machine replication protocols, like MultiPaxos and Raft, are a critical component of many distributed systems and databases. However, these protocols offer relatively low throughput due to several bottlenecked components. Numerous existing protocols fix different bottlenecks in isolation but fall short of a complete solution. When you fix one bottleneck, another arises. In this paper, we introduce compartmentalization, the first comprehensive technique to eliminate state machine replication bottlenecks. Compartmentalization involves decoupling individual bottlenecks into distinct components and scaling these components independently. Compartmentalization has two key strengths. First, compartmentalization leads to strong performance. In this paper, we demonstrate how to compartmentalize MultiPaxos to increase its throughput by 6× on a write-only workload and 16× on a mixed read-write workload. Unlike other approaches, we achieve this performance without the need for specialized hardware. Second, compartmentalization is a technique, not a protocol. Industry practitioners can apply compartmentalization to their protocols incrementally without having to adopt a completely new protocol.
Michael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Neil Giridharan, Joseph M. Hellerstein, Heidi Howard, Ion Stoica, Adriana Szekeres
Proc. VLDB Endow.2
2020 WPaxos: Wide Area Network Flexible Consensus
abstract
WPaxos is a multileader Paxos protocol that provides low-latency and high-throughput consensus across wide-area network (WAN) deployments. WPaxos uses multileaders, and partitions the object-space among these multileaders. Unlike statically partitioned multiple Paxos deployments, WPaxos is able to adapt to the changing access locality through object stealing. Multiple concurrent leaders coinciding in different zones steal ownership of objects from each other using phase-1 of Paxos, and then use phase-2 to commit update-requests on these objects locally until they are stolen by other leaders. To achieve fast phase-2 commits, WPaxos adopts the flexible quorums idea in a novel manner, and appoints phase-2 acceptors to be close to their respective leaders. We implemented WPaxos and evaluated it over WAN deployments across 5 AWS regions. The dynamic partitioning of the objectspace and emphasis on zone-local commits allow WPaxos to significantly outperform both partitioned Paxos deployments and leaderless Paxos approaches.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Tevfik Kosar
IEEE Trans. Parallel Distributed Syst.1
2019 Linearizable Quorum Reads in Paxos
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
HotStorage2
2019 Dissecting the Performance of Strongly-Consistent Replication Protocols
abstract
Many distributed databases employ consensus protocols to ensure that data is replicated in a strongly-consistent manner on multiple machines despite failures and concurrency. Unfortunately, these protocols show widely varying performance under different network, workload, and deployment conditions, and no previous study offers a comprehensive dissection and comparison of their performance. To fill this gap, we study single-leader, multi-leader, hierarchical multi-leader, and leaderless (opportunistic leader) consensus protocols, and present a comprehensive evaluation of their performance in local area networks (LANs) and wide area networks (WANs). We take a two-pronged systematic approach. We present an analytic modeling of the protocols using queuing theory and show simulations under varying controlled parameters. To cross-validate the analytic model, we also present empirical results from our prototyping and evaluation framework, Paxi. We distill our findings to simple throughput and latency formulas over the most significant parameters. These formulas enable the developers to decide which category of protocols would be most suitable under given deployment conditions.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas
SIGMOD Conference1
2019 Retroscope: Retrospective Monitoring of Distributed Systems
abstract
Retroscope is a comprehensive lightweight distributed monitoring tool that enables users to query and reconstruct past consistent global states of the system. Retroscope achieves this by augmenting the system with Hybrid Logical Clocks (HLC) and by streaming HLC-stamped event logs for storage and processing; these HLC timestamps are then used for constructing global (or nonlocal) snapshots upon request. Retroscope provides a rich querying language (RQL) to facilitate searching for global predicates across past consistent states. The search is performed by advancing through global states in small incremental steps, greatly reducing the amount of computation needed to construct consistent states. The Retroscope search algorithm is embarrassingly-parallel and can employ many worker processes (each processing up to 150,000 consistent snapshots per second) to handle a single query. We evaluate Retroscope's monitoring capabilities in two case studies: Chord and Apache ZooKeeper.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas, Sandeep S. Kulkarni
IEEE Trans. Parallel Distributed Syst.2
2018 Adapting to Access Locality via Live Data Migration in Globally Distributed Datastores
abstract
Storing data close to where it is used improves the performance of cloud applications. However, data access patterns change dynamically over time. Many datastores statically shard data making locality-adaptation difficult, and some provide limited capability for controlling the data-placement or migration. This leads to increased latency, reduced throughput, and expensive operations. To address this problem, we investigate the requirements for live data-migration and design four data-migration polices. Our policies use heuristics to determine the optimal data placement based on the access locality in the workload and load-balancing constraints. We show that even simple heuristics can be effective, and the topology-aware policies demonstrate overall better results with up to 70% latency improvement in medium locality workloads and nearly 95% improvement in workloads exhibiting very strong single-region access locality.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
IEEE BigData2
2017 Efficient Distributed Coordination at WAN-Scale
abstract
Traditional coordination services for distributed applications do not scale well over wide-area networks (WAN): centralized coordination fails to scale with respect to the increasing distances in the WAN, and distributed coordination fails to scale with respect to the number of nodes involved. We argue that it is possible to achieve scalability over WAN using a hierarchical coordination architecture and a smart token migration mechanism, and lay down the foundation of a novel design for a flexible-consistent coordination framework, called WanKeeper. We implemented WanKeeper based on the ZooKeeper API and deployed it over WAN as a proof of concept. Our experimental results based on the Yahoo! Cloud Serving Benchmark (YCSB), Apache BookKeeper replicated log service, and the Shared Cloud-backed File System (SCFS) show that WanKeeper provides multiple folds improvement in write/update performance in WAN compared to ZooKeeper, while keeping the same read performance.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Bekir O. Turkkan, Tevfik Kosar
ICDCS1
2017 Retrospective Lightweight Distributed Snapshots Using Loosely Synchronized Clocks
abstract
In order to take a consistent snapshot of a distributed system, it is necessary to collate and align local logs from each node to construct a pairwise concurrent cut. By leveraging NTP synchronized clocks, and augmenting them with logical clock causality information, Retroscope provides a lightweight solution for taking unplanned retrospective snapshots of past distributed system states. Instead of storing a multiversion copy of the entire system data, this is achieved efficiently by maintaining a configurable-size sliding window-log at each node to capture recent operations. In addition to retrospective snapshots, Retroscope also provides incremental and rolling snapshots that utilize an existing snapshot to reduce the cost of constructing a new snapshot in proximity. This capability is useful for performing stepwise debugging and root-cause analysis, and supporting data integrity monitoring and checkpoint-recovery. We implement Retroscope for the Voldemort distributed datastore and evaluate its performance under varying workloads.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas, Sandeep S. Kulkarni
ICDCS2
2016 Consensus in the Cloud: Paxos Systems Demystified
abstract
Coordination and consensus play an important role in datacenter and cloud computing, particularly in leader election, group membership, cluster management, service discovery, resource/access management, and consistent replication of the master nodes in services. Paxos protocols and systems provide a fault-tolerant solution to the distributed consensus problem and have attracted significant attention as well as generating substantial confusion. In order to elucidate the correct use of distributed coordination systems, we compare and contrast popular Paxos protocols and Paxos systems and present advantages and disadvantages for each. We also categorize the coordination use-patterns in cloud, and examine Google and Facebook infrastructures, as well as Apache top-level projects to investigate how they use Paxos protocols and systems. Finally, we analyze tradeoffs in the distributed coordination domain and identify promising future directions for achieving more scalable distributed coordination systems.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas
ICCCN1