Aleksey Charapko

dblp:142/2293 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-8072-0125ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 How to Evaluate Distributed Coordination Systems?-A Survey and Analysis
abstract
Coordination services and protocols are critical components of distributed systems and are essential for providing consistency, fault tolerance, and scalability. However, due to the lack of standard benchmarking and evaluation tools for distributed coordination services, coordination service developers/researchers either use a NoSQL standard benchmark and omit evaluating consistency, distribution, and fault tolerance; or create their own ad-hoc microbenchmarks and skip comparability with other services. In this study, we analyze and compare the evaluation mechanisms for known and widely used consensus algorithms, distributed coordination services, and distributed applications built on top of these services. We identify the most important requirements of distributed coordination service benchmarking, such as the metrics and parameters for the evaluation of the performance, scalability, availability, and consistency of these systems. Finally, we discuss why the existing benchmarks fail to address the complex requirements of distributed coordination system evaluation.
Bekir O. Turkkan, Elvis Rodrigues, Tevfik Kosar, Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
IEEE Trans. Parallel Distributed Syst.4
2025 Practical Considerations for Implementing State Machine Replication in the Cloud
abstract
State Machine Replication (SMR) protocols form the backbone of many distributed systems. Enterprises and startups increasingly build their distributed systems on the cloud due to its many advantages, such as scalability and cost-effectiveness. Due to their prevalence in systems, the practical aspects of SMR algorithms, such as efficiency and performance, become hugely important. These practical considerations can impact capacity planning, deployment strategies, and, ultimately, the cost of running stateful systems in the cloud. In this paper, we consider various practical choices that impact the performance and efficiency of state machine replication in the cloud. To that order, we design a language-agnostic Multi-Paxos-based state machine architecture and implement it in several languages popular for cloud deployment. In the process, we investigate the impact of threading, communication, and memory management models on replicated state machines’ performance and resource efficiency. We also examine the high-level implications of virtualization on the performance of SMR implementations in various languages. We present our findings as a collection of practical lessons backed by our experimental data and analysis.
Zhiying Liang, Vahab Jabrayilov, Abutalib Aghayev, Aleksey Charapko
ICDCS4
2025 HoliPaxos: Towards More Predictable Performance in State Machine Replication
abstract
State machine replication (SMR) algorithms ensure redundancy in critical systems and, as a result, underpin fault-tolerant distributed databases. Good SMR protocol performance is essential for capacity planning and meeting desired performance objectives. However, many implementations of popular SMR algorithms, such as MultiPaxos and Raft, have issues that make their performance unpredictable. This unpredictability often arises from certain "bolt-on" additions to core protocols, such as external failure detectors and replication log compaction. In this paper, we argue that tighter integration of such traditionally ad-hoc mechanisms with the core replication protocols can stabilize performance, making the solutions more reliable and more accessible to accurate capacity planning. Moreover, we show that these integrations can be non-disruptive for the underlying consensus algorithm, resulting in systems that preserve the simplicity and safety of traditional single-leader consensus-based SMR. To that order, we integrate the failure and slowdown detectors inside the SMR and achieve better performance and faster fail-over under various network partitions and node slowdown events. We also illustrate that tight integration of replication log management, pruning, and snapshotting can reduce memory and CPU usage while avoiding performance fluctuations associated with traditional log compaction and cleanup approaches.
Zhiying Liang, Vahab Jabrayilov, Abutalib Aghayev, Aleksey Charapko
Proc. VLDB Endow.4
2024 Cloudy Forecast: How Predictable is Communication Latency in the Cloud?
abstract
Many systems and services rely on timing assumptions for performance and availability to perform critical aspects of their operation, such as various timeouts for failure detectors or optimizations to concurrency control mechanisms. Many such assumptions rely on the ability of different components to communicate on time—a delay in communication may trigger the failure detector or cause the system to enter a less-optimized execution mode. Unfortunately, these timing assumptions are often set with little regard to actual communication guarantees of the underlying infrastructure – in particular, the variability of communication delays between processes in different nodes/servers. The higher communication variability holds especially true for systems deployed in the public cloud since the cloud is a utility shared by many users and organizations, making it prone to higher performance variance due to noisy neighbor syndrome. In this work, we present Cloud Latency Tester (CLT), a simple tool that can help measure the variability of communication delays between nodes to help engineers set proper values for their timing assumptions. We also provide our observational analysis of running CLT in three major cloud providers and share the lessons we learned.
Owen Hilyard, Bocheng Cui, Marielle Webster, Abishek Bangalore Muralikrishna, Aleksey Charapko
ICCCN5
2024 Simpler is Better: Revisiting Mencius State Machine Replication
abstract
State Machine Replication (SMR) algorithms provide durability and fault tolerance to many crucial applications, systems, and services. Most such systems rely on Viewstamped Replication and Paxos algorithms first proposed in the late 1980s or their more modern incarnation-Raft. However, many other SMR algorithms have been proposed in the past several decades, yet they do not boast successful industrial adoption. There are several reasons for such lack of adoption, ranging from increased implementation complexity to performance issues and unpredictability in various corner-case situations. In this paper, we examine these issues through the lens of Mencius, a multileader SMR algorithm popular in academia. We discuss its weaknesses that may prevent real-world use and propose Mencius Revisited to address them. In particular, our reimagining of the Mencius algorithm streamlines the handling of several corner cases, making it much simpler to implement while boasting multi-fold throughput improvement. Furthermore, Mencius Revisited has more predictable performance, especially under node failures-an ability crucial for production capacity planning. We achieve these improvements through algorithm simplification using opportunistic, low-latency batching and a revised failure-handling mechanism.
Bocheng Cui, Aleksey Charapko
SRDS2
2022 Metastable Failures in the Wild
Lexiang Huang, Matt Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko
OSDI8
2021 Metastable failures in distributed systems
abstract
We describe metastable failures---a failure pattern in distributed systems. Currently, metastable failures manifest themselves as black swan events; they are outliers because nothing in the past points to their possibility, have a severe impact, and are much easier to explain in hindsight than to predict. Although instances of metastable failures can look different at the surface, deeper analysis shows that they can be understood within the same framework.
Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, Timothy Zhu
HotOS3
2021 Scalable but wasteful: current state of replication in the cloud
abstract
Consensus protocols are at the core of strongly consistent replication deployed in cloud-based storage systems. There have been many proposals to optimize these protocols, most of which work by identifying and shifting load from bottlenecked nodes to underutilized nodes.
Venkata Swaroop Matte, Aleksey Charapko, Abutalib Aghayev
HotStorage2
2021 PigPaxos: Devouring the Communication Bottlenecks in Distributed Consensus
abstract
Strongly consistent replication helps keep application logic simple and provides significant benefits for correctness and manageability. Unfortunately, the adoption of strongly-consistent replication protocols has been curbed due to their limited scalability and performance. To alleviate the leader bottleneck in strongly-consistent replication protocols, we introduce Pig, an in-protocol communication aggregation and piggybacking technique. Pig employs randomly selected nodes from follower subgroups to relay the leader's message to the rest of the followers in the subgroup, and to perform in-network aggregation of acknowledgments back from these followers. By randomly alternating the relay nodes across replication operations, Pig shields the relay nodes as well as the leader from becoming hotspots and improves throughput scalability. We showcase Pig in the context of classical Paxos protocols employed for strongly consistent replication by many cloud computing services and databases. We implement and evaluate PigPaxos, in comparison to Paxos and EPaxos protocols under various workloads over clusters of size 5 to 25 nodes. We show that the aggregation at the relay has little latency overhead, and PigPaxos can provide more than 3 folds improved throughput over Paxos and EPaxos with little latency deterioration. We support our experimental observations with the analytical modeling of the bottlenecks and show that the communication bottlenecks are minimized when employing only one randomly rotating relay node.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
SIGMOD Conference1
2021 Scaling Replicated State Machines with Compartmentalization
abstract
State machine replication protocols, like MultiPaxos and Raft, are a critical component of many distributed systems and databases. However, these protocols offer relatively low throughput due to several bottlenecked components. Numerous existing protocols fix different bottlenecks in isolation but fall short of a complete solution. When you fix one bottleneck, another arises. In this paper, we introduce compartmentalization, the first comprehensive technique to eliminate state machine replication bottlenecks. Compartmentalization involves decoupling individual bottlenecks into distinct components and scaling these components independently. Compartmentalization has two key strengths. First, compartmentalization leads to strong performance. In this paper, we demonstrate how to compartmentalize MultiPaxos to increase its throughput by 6× on a write-only workload and 16× on a mixed read-write workload. Unlike other approaches, we achieve this performance without the need for specialized hardware. Second, compartmentalization is a technique, not a protocol. Industry practitioners can apply compartmentalization to their protocols incrementally without having to adopt a completely new protocol.
Michael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Neil Giridharan, Joseph M. Hellerstein, Heidi Howard, Ion Stoica, Adriana Szekeres
Proc. VLDB Endow.3
2020 WPaxos: Wide Area Network Flexible Consensus
abstract
WPaxos is a multileader Paxos protocol that provides low-latency and high-throughput consensus across wide-area network (WAN) deployments. WPaxos uses multileaders, and partitions the object-space among these multileaders. Unlike statically partitioned multiple Paxos deployments, WPaxos is able to adapt to the changing access locality through object stealing. Multiple concurrent leaders coinciding in different zones steal ownership of objects from each other using phase-1 of Paxos, and then use phase-2 to commit update-requests on these objects locally until they are stolen by other leaders. To achieve fast phase-2 commits, WPaxos adopts the flexible quorums idea in a novel manner, and appoints phase-2 acceptors to be close to their respective leaders. We implemented WPaxos and evaluated it over WAN deployments across 5 AWS regions. The dynamic partitioning of the objectspace and emphasis on zone-local commits allow WPaxos to significantly outperform both partitioned Paxos deployments and leaderless Paxos approaches.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Tevfik Kosar
IEEE Trans. Parallel Distributed Syst.2
2019 Linearizable Quorum Reads in Paxos
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
HotStorage1
2019 Dissecting the Performance of Strongly-Consistent Replication Protocols
abstract
Many distributed databases employ consensus protocols to ensure that data is replicated in a strongly-consistent manner on multiple machines despite failures and concurrency. Unfortunately, these protocols show widely varying performance under different network, workload, and deployment conditions, and no previous study offers a comprehensive dissection and comparison of their performance. To fill this gap, we study single-leader, multi-leader, hierarchical multi-leader, and leaderless (opportunistic leader) consensus protocols, and present a comprehensive evaluation of their performance in local area networks (LANs) and wide area networks (WANs). We take a two-pronged systematic approach. We present an analytic modeling of the protocols using queuing theory and show simulations under varying controlled parameters. To cross-validate the analytic model, we also present empirical results from our prototyping and evaluation framework, Paxi. We distill our findings to simple throughput and latency formulas over the most significant parameters. These formulas enable the developers to decide which category of protocols would be most suitable under given deployment conditions.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas
SIGMOD Conference2
2019 Retroscope: Retrospective Monitoring of Distributed Systems
abstract
Retroscope is a comprehensive lightweight distributed monitoring tool that enables users to query and reconstruct past consistent global states of the system. Retroscope achieves this by augmenting the system with Hybrid Logical Clocks (HLC) and by streaming HLC-stamped event logs for storage and processing; these HLC timestamps are then used for constructing global (or nonlocal) snapshots upon request. Retroscope provides a rich querying language (RQL) to facilitate searching for global predicates across past consistent states. The search is performed by advancing through global states in small incremental steps, greatly reducing the amount of computation needed to construct consistent states. The Retroscope search algorithm is embarrassingly-parallel and can employ many worker processes (each processing up to 150,000 consistent snapshots per second) to handle a single query. We evaluate Retroscope's monitoring capabilities in two case studies: Chord and Apache ZooKeeper.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas, Sandeep S. Kulkarni
IEEE Trans. Parallel Distributed Syst.1
2018 Adapting to Access Locality via Live Data Migration in Globally Distributed Datastores
abstract
Storing data close to where it is used improves the performance of cloud applications. However, data access patterns change dynamically over time. Many datastores statically shard data making locality-adaptation difficult, and some provide limited capability for controlling the data-placement or migration. This leads to increased latency, reduced throughput, and expensive operations. To address this problem, we investigate the requirements for live data-migration and design four data-migration polices. Our policies use heuristics to determine the optimal data placement based on the access locality in the workload and load-balancing constraints. We show that even simple heuristics can be effective, and the topology-aware policies demonstrate overall better results with up to 70% latency improvement in medium locality workloads and nearly 95% improvement in workloads exhibiting very strong single-region access locality.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas
IEEE BigData1
2017 Efficient Distributed Coordination at WAN-Scale
abstract
Traditional coordination services for distributed applications do not scale well over wide-area networks (WAN): centralized coordination fails to scale with respect to the increasing distances in the WAN, and distributed coordination fails to scale with respect to the number of nodes involved. We argue that it is possible to achieve scalability over WAN using a hierarchical coordination architecture and a smart token migration mechanism, and lay down the foundation of a novel design for a flexible-consistent coordination framework, called WanKeeper. We implemented WanKeeper based on the ZooKeeper API and deployed it over WAN as a proof of concept. Our experimental results based on the Yahoo! Cloud Serving Benchmark (YCSB), Apache BookKeeper replicated log service, and the Shared Cloud-backed File System (SCFS) show that WanKeeper provides multiple folds improvement in write/update performance in WAN compared to ZooKeeper, while keeping the same read performance.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas, Bekir O. Turkkan, Tevfik Kosar
ICDCS2
2017 Retrospective Lightweight Distributed Snapshots Using Loosely Synchronized Clocks
abstract
In order to take a consistent snapshot of a distributed system, it is necessary to collate and align local logs from each node to construct a pairwise concurrent cut. By leveraging NTP synchronized clocks, and augmenting them with logical clock causality information, Retroscope provides a lightweight solution for taking unplanned retrospective snapshots of past distributed system states. Instead of storing a multiversion copy of the entire system data, this is achieved efficiently by maintaining a configurable-size sliding window-log at each node to capture recent operations. In addition to retrospective snapshots, Retroscope also provides incremental and rolling snapshots that utilize an existing snapshot to reduce the cost of constructing a new snapshot in proximity. This capability is useful for performing stepwise debugging and root-cause analysis, and supporting data integrity monitoring and checkpoint-recovery. We implement Retroscope for the Voldemort distributed datastore and evaluate its performance under varying workloads.
Aleksey Charapko, Ailidani Ailijiang, Murat Demirbas, Sandeep S. Kulkarni
ICDCS1
2016 Consensus in the Cloud: Paxos Systems Demystified
abstract
Coordination and consensus play an important role in datacenter and cloud computing, particularly in leader election, group membership, cluster management, service discovery, resource/access management, and consistent replication of the master nodes in services. Paxos protocols and systems provide a fault-tolerant solution to the distributed consensus problem and have attracted significant attention as well as generating substantial confusion. In order to elucidate the correct use of distributed coordination systems, we compare and contrast popular Paxos protocols and Paxos systems and present advantages and disadvantages for each. We also categorize the coordination use-patterns in cloud, and examine Google and Facebook infrastructures, as well as Apache top-level projects to investigate how they use Paxos protocols and systems. Finally, we analyze tradeoffs in the distributed coordination domain and identify promising future directions for achieving more scalable distributed coordination systems.
Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas
ICCCN2
2014 Indexing and Retrieving Continuations in Musical Time Series Data Using Relational Databases
abstract
This paper proposed and tested a model that provides quick search and retrieval of continuations for time series, particularly musical data, using relational databases. The model extends an existing interactive music-generation system by focusing on large input sequences. Experiments using textural and musical data provided satisfactory performance results for the model.
Aleksey Charapko, Ching-Hua Chuan
ISM1
2013 Predicting Key Recognition Difficulty in Polyphonic Audio
abstract
In this paper, we present statistical models to predict the difficulty of recognizing musical keys from polyphonic audio signals. Automatic audio key finding has been studied for many years, and various approaches have been proposed and reported. Reports of these methods' performance are usually based on the proposers' own data sets. Without details on the data set, i.e., how challenging the data set is, directly comparing the effectiveness of these methods is not meaningful or even possible. Thus, in this study we focus on predicting the difficulty level of key recognition as perceived by human experts. Given an audio recording, represented as the extracted acoustic features, we apply multiple linear regression and proportional odds model to predict the difficulty level of the recording, annotated by experts as an integer on a 5-point Likert scale. We use four metrics to evaluate our prediction results: root mean square error, Pearson correlation coefficient, exact accuracy, and adjacent accuracy. We also examine the difference between experts' annotations and discuss their consistency.
Ching-Hua Chuan, Aleksey Charapko
ISM2