Alex Depoutovitch

dblp:47/2837 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0002-5648-2622ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 68% Hardware reliability and fault tolerance · 14% Cloud and datacenter computing · 9%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 50% Transaction processing and concurrency control · 50%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 8 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transaction processing and concurrency control › concurrency control
locking protocols
0.712023
Taurus MM: bringing multi-master to the cloud · Proc. VLDB Endow. 2023
Distributed systems
distributed coordination
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Distributed systems
fault tolerance
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Distributed systems › group communication
membership management
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Hardware reliability and fault tolerance › network fault tolerance
network partition tolerance
0.712023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Distributed systems › replication › replication and fault tolerance
replication and recovery
0.412020
Taurus Database: How to be Fast, Available, and Frugal in the Cloud · SIGMOD Conference 2020
Distributed systems › peer-to-peer systems
overlay networks
0.212023
Partial Network Partitioning · ACM Trans. Comput. Syst. 2023
Operating systems
fault tolerance
0.112010
Otherworld: giving applications a chance to survive OS kernel crashes · EuroSys 2010

Methods — techniques the papers use, named apart from their topics

overlay routing · 0.7constant-time snapshots · 0.4append-only storage · 0.4kernel error handling · 0.1
YearPublicationVenuePosition
2023 Taurus MM: bringing multi-master to the cloud
abstract
A single-master database has limited update capacity because a single node handles all updates. A multi-master database potentially has higher update capacity because the load is spread across multiple nodes. However, the need to coordinate updates and ensure durability can generate high network traffic. Reducing network load is particularly important in a cloud environment where the network infrastructure is shared among thousands of tenants. In this paper, we present Taurus MM, a shared-storage multi-master database optimized for cloud environments. It implements two novel algorithms aimed at reducing network traffic plus a number of additional optimizations. The first algorithm is a new type of distributed clock that combines the small size of Lamport clocks with the effective support of distributed snapshots of vector clocks. The second algorithm is a new hybrid page and row locking protocol that significantly reduces the number of lock requests sent over the network. Experimental results on a cluster with up to eight masters demonstrate superior performance compared to Aurora multi-master and CockroachDB.
Alex Depoutovitch, Per-Åke Larson, Jack Ng, Guanzhu Xiong, Paul Lee, Emad Boctor, Samiao Ren, Lengdong Wu, Calvin Sun
Proc. VLDB Endow.1
2023 Partial Network Partitioning
abstract
We present an extensive study focused on partial network partitioning. Partial network partitions disrupt the communication between some but not all nodes in a cluster. First, we conduct a comprehensive study of system failures caused by this fault in 13 popular systems. Our study reveals that the studied failures are catastrophic (e.g., lead to data loss), easily manifest, and are mainly due to design flaws. Our analysis identifies vulnerabilities in core systems mechanisms including scheduling, membership management, and ZooKeeper-based configuration management. Second, we dissect the design of nine popular systems and identify four principled approaches for tolerating partial partitions. Unfortunately, our analysis shows that implemented fault tolerance techniques are inadequate for modern systems; they either patch a particular mechanism or lead to a complete cluster shutdown, even when alternative network paths exist. Finally, our findings motivate us to build Nifty, a transparent communication layer that masks partial network partitions. Nifty builds an overlay between nodes to detour packets around partial partitions. Nifty provides an approach for applications to optimize their operation during a partial partition. We demonstrate the benefit of this approach through integrating Nifty with VoltDB, HDFS, and Kafka.
Basil Alkhatib, Sreeharsha Udayashankar, Sara Qunaibi, Ahmed Alquraan, Mohammed Alfatafta, Wael Al-Manasrah, Alex Depoutovitch, Samer Al-Kiswany
ACM Trans. Comput. Syst.7
2020 Taurus Database: How to be Fast, Available, and Frugal in the Cloud
abstract
Using cloud Database as a Service (DBaaS) offerings instead of on-premise deployments is increasingly common. Key advantages include improved availability and scalability at a lower cost than on-premise alternatives. In this paper, we describe the design of Taurus, a new multi-tenant cloud database system. Taurus separates the compute and storage layers in a similar manner to Amazon Aurora and Microsoft Socrates and provides similar benefits, such as read replica support, low network utilization, hardware sharing and scalability. However, the Taurus architecture has several unique advantages. Taurus offers novel replication and recovery algorithms providing better availability than existing approaches using the same or fewer replicas. Also, Taurus is highly optimized for performance, using no more than one network hop on critical paths and exclusively using append-only storage, delivering faster writes, reduced device wear, and constant-time snapshots. This paper describes Taurus and provides a detailed description and analysis of the storage node architecture, which has not been previously available from the published literature.
Alex Depoutovitch, Jin Chen 0006, Per-Åke Larson, Jack Ng, Wenlin Cui
SIGMOD Conference1
2010 Otherworld: giving applications a chance to survive OS kernel crashes
abstract
The default behavior of all commodity operating systems today is to restart the system when a critical error is encountered in the kernel. This terminates all running applications with an attendant loss of "work in progress" that is nonpersistent.
Alex Depoutovitch, Michael Stumm
EuroSys1