Brendan Cully

dblp:33/6584 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0001-6395-0500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorComputer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 51% Distributed systems · 35% Cloud and datacenter computing · 7%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 57% Database system architecture and tuning · 43%
Software engineering, system software, and programming languages
2 papers
Program verification · 55% Software maintenance and evolution · 35% Debugging and program repair · 10%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
crash consistency
0.512021
Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3 · SOSP 2021
Storage systems
key-value storage
0.512021
Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3 · SOSP 2021
Storage systems
storage reliability
0.512021
Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3 · SOSP 2021
Distributed systems › replication
database replication
0.322013
RemusDB: transparent high availability for database systems · VLDB J. 2013
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Distributed systems
replication
0.322013
RemusDB: transparent high availability for database systems · VLDB J. 2013
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Distributed systems
fault tolerance
0.222013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Remus: High Availability via Asynchronous Virtual Machine Replication. (Best Paper) · NSDI 2008
Distributed systems › fault tolerance
high availability
0.222011
RemusDB: Transparent High Availability for Database Systems · Proc. VLDB Endow. 2011
Remus: High Availability via Asynchronous Virtual Machine Replication. (Best Paper) · NSDI 2008
Memory systems
non-volatile memory
0.212014
Strata: scalable high-performance storage on virtualized non-volatile memory · FAST 2014
Distributed and cloud data management
high availability
0.212013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Cloud and datacenter computing
virtualization
0.222008
Remus: High Availability via Asynchronous Virtual Machine Replication. (Best Paper) · NSDI 2008
Parallax: virtual disks for virtual machines · EuroSys 2008
Program verification
lightweight formal methods
0.112021
Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3 · SOSP 2021
Software maintenance and evolution
program comprehension
0.112009
Tralfamadore: unifying source code and execution experience · EuroSys 2009
Storage systems
distributed storage
0.112008
Parallax: virtual disks for virtual machines · EuroSys 2008
Distributed systems › replication
virtual machine replication
0.112008
Remus: High Availability via Asynchronous Virtual Machine Replication. (Best Paper) · NSDI 2008
Cloud and datacenter computing › virtualization
virtual machine storage
0.112008
Parallax: virtual disks for virtual machines · EuroSys 2008
Storage systems
file systems
0.112014
Strata: scalable high-performance storage on virtualized non-volatile memory · FAST 2014
Storage systems
scalable storage
0.112014
Strata: scalable high-performance storage on virtualized non-volatile memory · FAST 2014
Distributed systems
consensus
0.012013
RemusDB: transparent high availability for database systems · VLDB J. 2013
Debugging and program repair
fault localization
0.012009
Tralfamadore: unifying source code and execution experience · EuroSys 2009
Memory systems
cache
0.012008
Parallax: virtual disks for virtual machines · EuroSys 2008
Storage systems › file systems
snapshot
0.012008
Parallax: virtual disks for virtual machines · EuroSys 2008

Methods — techniques the papers use, named apart from their topics

model checking · 1.0lightweight formal methods · 1.0asynchronous replication · 0.1
YearPublicationVenuePosition
2025 Serverless Elasticsearch: the Architecture Transformation from Stateful to Stateless
abstract
Elasticsearch (ES) is a distributed search and analytics engine consisting of a cluster of nodes, each hosting a disjoint subset of data. ES has a shared-nothing architecture that relies on local disks to store cluster data such as index files, transaction logs, and cluster state metadata. This stateful architecture couples compute with storage, and leads to different data tiers (e.g., hot, warm, cold, frozen) of hardware and configurations that the administrator chooses from to balance cost, performance, and high availability. In this paper, we show a new serverless architecture that decouples compute from storage. Serverless ES offloads data to an affordable, highly available cloud object store, while supporting the same APIs and read-after-write semantics. We show why and how this stateless architecture simplifies the tiers to just two: indexing and search, allowing indexing and searching practically limitless data while scaling each tier independently. We describe how we wrap index data in a custom batch commit format to the object store to decrease upload costs by up to 100x, how we batch transaction log uploads to decrease upload costs by up to 30x, and how we delete files from the object store. We experimentally show that Serverless ES can get twice the indexing throughput of (stateful) ES on comparable compute hardware by using object storage for durability instead of replication, and can scale linearly to match ingestion workloads.
Iraklis Psaroudakis, Pooya Salehi, Jason Bryan, Francisco Fernández Castaño, Brendan Cully, Ankita Kumar, Henning Andersen, Thomas Repantis
SoCC5
2021 Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3
abstract
This paper reports our experience applying lightweight formal methods to validate the correctness of ShardStore, a new key-value storage node implementation for the Amazon S3 cloud object storage service. By "lightweight formal methods" we mean a pragmatic approach to verifying the correctness of a production storage node that is under ongoing feature development by a full-time engineering team. We do not aim to achieve full formal verification, but instead emphasize automation, usability, and the ability to continually ensure correctness as both software and its specification evolve over time. Our approach decomposes correctness into independent properties, each checked by the most appropriate tool, and develops executable reference models as specifications to be checked against the implementation. Our work has prevented 16 issues from reaching production, including subtle crash consistency and concurrency problems, and has been extended by non-formal-methods experts to check new features and properties as ShardStore has evolved.
James Bornholt, Rajeev Joshi, Vytautas Astrauskas, Brendan Cully, Bernhard Kragl, Seth Markle, Kyle Sauri, Drew Schleit, Grant Slatton, Serdar Tasiran, Jacob Van Geffen, Andy Warfield
SOSP4
2014 Strata: scalable high-performance storage on virtualized non-volatile memory
Brendan Cully, Jake Wires, Dutch T. Meyer, Kevin Jamieson 0002, Keir Fraser, Tim Deegan, Daniel Stodden, Geoffrey Lefebvre, Daniel Ferstay, Andy Warfield
FAST1
2013 RemusDB: transparent high availability for database systems
Umar Farooq Minhas, Shriram Rajagopalan, Brendan Cully, Ashraf Aboulnaga, Kenneth Salem, Andy Warfield
VLDB J.3
2012 Execution mining
abstract
Operating systems represent large pieces of complex software that are carefully tested and broadly deployed. Despite this, developers frequently have little more than their source code to understand how they behave. This static representation of a system results in limited insight into execution dynamics, such as what code is important, how data flows through a system, or how threads interact with one another. We describe Tralfamadore, a system that preserves complete traces of machine execution as an artifact that can be queried and analyzed with a library of simple, reusable operators, making it easy to develop and run new dynamic analyses. We demonstrate the benefits of this approach with several example applications, including a novel unified source and execution browser.
Geoffrey Lefebvre, Brendan Cully, Christopher Head, Mark Spear, Norman C. Hutchinson, Michael J. Feeley, Andy Warfield
VEE2
2012 SecondSite: disaster tolerance as a service
abstract
This paper describes the design and implementation of SecondSite, a cloud-based service for disaster tolerance. SecondSite extends the Remus virtualization-based high availability system by allowing groups of virtual machines to be replicated across data centers over wide-area Internet links. The goal of the system is to commodify the property of availability, exposing it as a simple tick box when configuring a new virtual machine. To achieve this in the wide area, we have had to tackle the related issues of replication traffic bandwidth, reliable failure detection across geographic regions and traffic redirection over a wide-area network without compromising on transparency and consistency.
Shriram Rajagopalan, Brendan Cully, Ryan O'Connor, Andy Warfield
VEE2
2011 RemusDB: Transparent High Availability for Database Systems
Umar Farooq Minhas, Shriram Rajagopalan, Brendan Cully, Ashraf Aboulnaga, Kenneth Salem, Andy Warfield
Proc. VLDB Endow.3
2009 Tralfamadore: unifying source code and execution experience
abstract
Program source is an intermediate representation of software; it lies between a developer's intention and the hardware's execution. Despite advances in languages and development tools, source itself and the applications we use to view it remain an essentially static representation of software, from which developers can spend considerable energy postulating actual behavior.
Geoffrey Lefebvre, Brendan Cully, Michael J. Feeley, Norman C. Hutchinson, Andy Warfield
EuroSys2
2008 Parallax: virtual disks for virtual machines
abstract
Parallax is a distributed storage system that uses virtualization to provide storage facilities specifically for virtual environments. The system employs a novel architecture in which storage features that have traditionally been implemented directly on high-end storage arrays and switches are relocated into a federation of storage VMs, sharing the same physical hosts as the VMs that they serve. This architecture retains the single administrative domain and OS agnosticism achieved by array- and switch-based approaches, while lowering the bar on hardware requirements and facilitating the development of new features. Parallax offers a comprehensive set of storage features including frequent, low-overhead snapshot of virtual disks, the 'gold-mastering' of template images, and the ability to use local disks as a persistent cache to dampen burst demand on networked storage.
Dutch T. Meyer, Gitika Aggarwal, Brendan Cully, Geoffrey Lefebvre, Michael J. Feeley, Norman C. Hutchinson, Andy Warfield
EuroSys3
2008 Remus: High Availability via Asynchronous Virtual Machine Replication. (Best Paper)
Brendan Cully, Geoffrey Lefebvre, Dutch T. Meyer, Michael J. Feeley, Norman C. Hutchinson, Andy Warfield
NSDI1
2008 Understanding Performance for Two 802.11 Competing Flows
Kan Cai, Michael J. Feeley, Brendan Cully, Sharath J. George
J. Comput. Sci. Technol.3
2007 Understanding Performance for Two 802.11 Competing Flows
abstract
It is well known that 802.11 suffers from both inefficiency and unfairness in the face of competition and interference. This paper provides a detailed analysis of the impact of topology and traffic type on network performance when two flows compete with each other for airspace. We consider both TCP and UDP flows and a comprehensive set of node topologies. We vary these topologies to consider all combinations of the following four node-to-node interactions: (1) nodes unable to read or sense each other, (2) nodes able to sense each other but not able to read each other's packets and nodes able to communicate with (3) weak and with (4) strong signal. We evaluate all possible cases through simulation and show that the cases can be reduced to 9 UDP and 10 TCP 802.1 Ig models with similar efficiency/fairness characteristics. We also validate our simulation results with extensive experiments conducted in a laboratory testbed. These more detailed models improve on previous work such as hidden-Zexposed-terminal categorization and are thus better suited as a basis for adaptive techniques to improve performance in 802.11 multi-hop WLAN or Mesh Networks.
Kan Cai, Michael J. Feeley, Brendan Cully, Sharath J. George
MASS3