EDBT 2026 Demo / reviewers in the wild / expert
Michael Abebe 0001
dblp:216/4991-1
· DBLP profile ↗
11ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0002-9519-7883ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Caerus: Low-Latency Distributed Transactions for Geo-Replicated SystemsabstractDistributed deterministic database systems achieve high transaction throughput for geographically replicated data. Supporting transactions with ACID guarantees requires deterministic databases to order transactions globally to dictate execution order. In a geographically distributed environment, ordering transactions globally can take multiple wide-area network (WAN) round trips of messaging, which adds significant latency to transaction response times, leading to poor user experiences. To improve the response time of transactions in deterministic databases, we propose an ordering protocol that can include a transaction in the global order in a single WAN round trip to the primary regions of the data items within the transaction's read and write set. The protocol reduces the cost of determining the global order for all transactions by leveraging deterministic merging of partial sequences of transactions per geographic region. We implement the protocol in Caerus, our geo-replicated deterministic database system that serializably commits and replicates transactions after a delay of only a single WAN round trip of messaging. Using popular workload benchmarks over geographically replicated data in Azure, we show that Caerus outperforms state-of-the-art comparison systems to deliver low-latency transaction execution. Josh Hildred, Michael Abebe 0001, Khuzaima Daudjee |
Proc. VLDB Endow. | 2 |
| 2022 | Proteus: Autonomous Adaptive Storage for Mixed WorkloadsabstractEnterprises use distributed database systems to meet the demands of mixed or hybrid transaction/analytical processing (HTAP) workloads that contain both transactional (OLTP) and analytical (OLAP) requests. Distributed HTAP systems typically maintain a complete copy of data in row-oriented storage format that is well-suited for OLTP workloads and a second complete copy in column-oriented storage format optimized for OLAP workloads. Maintaining these data copies consumes significant storage space and system resources. Conversely, if a system stores data in a single format, OLTP or OLAP workload performance suffers. This paper presents Proteus, a distributed HTAP database system that adaptively and autonomously selects and changes its storage layout to optimize for mixed workloads. Proteus generates physical execution plans that utilize storage-aware operators for efficient transaction execution. Using comprehensive HTAP workloads and state-of-the-art comparison systems, we demonstrate that Proteus delivers superior HTAP performance while providing OLTP and OLAP performance on par with designs specialized for either type of workload. Michael Abebe 0001, Horatiu Lazu, Khuzaima Daudjee |
SIGMOD Conference | 1 |
| 2022 | Tiresias: Enabling Predictive Autonomous Storage and IndexingabstractTo efficiently store and query a DBMS, administrators must select storage and indexing configurations. For example, one must decide whether data should be stored in rows or columns, in-memory or on disk, and which columns to index. These choices can be challenging to make for workloads that are mixed requiring hybrid transactional and analytical processing (HTAP) support. There is growing interest in system designs that can adapt how data is stored and indexed to execute these workloads efficiently. We present Tiresias , a predictor that learns the cost of data accesses and predicts their latency and likelihood under different storage scenarios. Tiresias makes these predictions by collecting observed latencies and access histories to build predictive models in an online manner, enabling autonomous storage and index adaptation. Experimental evaluation shows the benefits of predictive adaptation and the trade-offs for different predictive techniques. Michael Abebe 0001, Horatiu Lazu, Khuzaima Daudjee |
Proc. VLDB Endow. | 1 |
| 2020 | DynaMast: Adaptive Dynamic Mastering for Replicated SystemsabstractSingle-master replicated database systems strive to be scalable by offloading reads to replica nodes. However, single-master systems suffer from the performance bottleneck of all updates executing at a single site. Multi-master replicated systems distribute updates among sites but incur costly coordination for multi-site transactions. We present DynaMast, a lazily replicated, multi-master database system that guarantees one-site transaction execution while effectively distributing both reads and updates among multiple sites. DynaMast benefits from these advantages by dynamically transferring the mastership of data, or remastering, among sites using a lightweight metadata-based protocol. DynaMast leverages remastering to adaptively place master copies to balance load and minimize future remastering. Using benchmark workloads, we demonstrate that DynaMast delivers superior performance over existing replicated database system architectures. Michael Abebe 0001, Brad Glasbergen, Khuzaima Daudjee |
ICDE | 1 |
| 2020 | Sentinel: Understanding Data SystemsabstractThe complexity of modern data systems and applications greatly increases the challenge in understanding system behaviour and diagnosing performance problems. When these problems arise, system administrators are left with the difficult task of remedying them by relying on large debug log files, vast numbers of metrics, and system-specific tooling. We demonstrate the Sentinel system, which enables administrators to analyze systems and applications by building models of system execution and comparing them to derive key differences in behaviour. The resulting analyses are then presented as system reports to administrators and developers in an intuitive fashion. Users of Sentinel can locate, identify and take steps to resolve the reported performance issues. As Sentinel's models are constructed online by intercepting debug logging library calls, Sentinel's functionality incurs little overhead and works with all systems that use standard debug logging libraries. Brad Glasbergen, Michael Abebe 0001, Khuzaima Daudjee, Daniel Vogel 0001, Jian Zhao 0010 |
SIGMOD Conference | 2 |
| 2020 | ChronoCache: Predictive and Adaptive Mid-Tier Query Result CachingabstractThe performance of data-driven, web-scale client applications is sensitive to access latency. To address this concern, enterprises strive to cache data on edge nodes that are closer to users, thereby avoiding expensive round-trips to remote data centers. However, these geo-distributed approaches are limited to caching static data. In this paper we present ChronoCache, a mid-tier caching system that exploits the presence of geo-distributed edge nodes to cache database query results closer to users. ChronoCache transparently learns and leverages client application access patterns to predictively combine query requests and cache their results ahead of time, thereby reducing costly round-trips to the remote database. We show that ChronoCache reduces query response times by up to 2/3 over prior approaches on multiple representative benchmark workloads. representative benchmark workloads. Brad Glasbergen, Kyle Langendoen, Michael Abebe 0001, Khuzaima Daudjee |
SIGMOD Conference | 3 |
| 2020 | MorphoSys: Automatic Physical Design Metamorphosis for Distributed Database SystemsabstractDistributed database systems are widely used to meet the demands of storing and managing computation-heavy workloads. To boost performance and minimize resource and data contention, these systems require selecting a distributed physical design that determines where to place data, and which data items to replicate and partition. Deciding on a physical design is difficult as each choice poses a trade-off in the design space, and a poor choice can significantly degrade performance. Current design decisions are typically static and cannot adapt to workload changes or are unable to combine multiple design choices such as data replication and data partitioning integrally. This paper presents MorphoSys , a distributed database system that dynamically chooses, and alters, its physical design based on the workload. MorphoSys makes integrated design decisions for all of the data partitioning, replication and placement decisions on-the-fly using a learned cost model. MorphoSys provides efficient transaction execution in the face of design changes via a novel concurrency control and update propagation scheme. Our experimental evaluation, using several benchmark workloads and state-of-the-art comparison systems, shows that MorphoSys delivers excellent system performance through effective and efficient physical designs. Michael Abebe 0001, Brad Glasbergen, Khuzaima Daudjee |
Proc. VLDB Endow. | 1 |
| 2020 | Sentinel: Universal Analysis and Insight for Data Systems
Brad Glasbergen, Michael Abebe 0001, Khuzaima Daudjee, Amit Levi 0001 |
Proc. VLDB Endow. | 2 |
| 2019 | WatDFS: A Project for Understanding Distributed Systems in the Undergraduate CurriculumabstractThe ubiquity of distributed computing systems has led to an increased focus on distributed systems in the undergraduate curriculum. In this paper, we describe the design of, and our experiences with, a new distributed systems project where students implement the core components of a distributed file system called WatDFS. The WatDFS project enables students to meet learning objectives from across the distributed systems curriculum and interact with real systems, while providing a high-quality testing environment that yields actionable feedback to students on their submissions. Michael Abebe 0001, Brad Glasbergen, Khuzaima Daudjee |
SIGCSE | 1 |
| 2018 | Apollo: Learning Query Correlations for Predictive Caching in Geo-Distributed Systems
Brad Glasbergen, Michael Abebe 0001, Khuzaima Daudjee, Scott Foggo, Anil Pacaci |
EDBT | 2 |
| 2018 | EC-Store: Bridging the Gap between Storage and Latency in Distributed Erasure Coded SystemsabstractCloud storage systems typically choose between replicating or erasure encoding data to provide fault tolerance. Replication ensures that data can be accessed from a single site but incurs a much higher storage overhead, which is a costly downside for large-scale storage systems. Erasure coding has a lower storage requirement but relies on encoding/decoding and distributed data retrieval, which can result in straggling requests that increase response times. We propose strategies for data access and data movement within erasure-coded storage systems that significantly reduce data retrieval times. We present EC-Store, a system that incorporates these dynamic strategies for data access and movement based on workload access patterns. Through detailed evaluation using two benchmark workloads, we show that EC-Store incurs significantly less storage overhead than replication while achieving better performance than both replicated and erasure-coded storage systems. Michael Abebe 0001, Khuzaima Daudjee, Brad Glasbergen, Yuanfeng Tian |
ICDCS | 1 |