Chris Douglas

dblp:20/11411 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0003-2628-354XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Cloud and datacenter computing · 76% Distributed systems · 19% Storage systems · 5%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 80% Database system architecture and tuning · 20%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cloud economics
cloud cost optimization
0.812024
SkyPIE: A Fast & Accurate Oracle for Object Placement · Proc. ACM Manag. Data 2024
Cloud and datacenter computing
cloud storage
0.812024
SkyPIE: A Fast & Accurate Oracle for Object Placement · Proc. ACM Manag. Data 2024
Distributed systems › distributed resource management
object placement
0.812024
SkyPIE: A Fast & Accurate Oracle for Object Placement · Proc. ACM Manag. Data 2024
Cloud and datacenter computing
cluster resource management and scheduling
0.732017
Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015
Blind men and an elephant coalescing open-source, academic, and industrial perspectives on BigData · ICDE 2015
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.622019
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
REEF: Retainable Evaluator Execution Framework · SIGMOD Conference 2015
Cloud and datacenter computing › cluster resource management and scheduling
resource scheduling
0.412019
Hydra: a federated resource manager for data-center scale analytics · NSDI 2019
Distributed and cloud data management › distributed data store
distributed file system
0.312017
Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics · SIGMOD Conference 2017
Mathematical optimization › continuous optimization
convex optimization
0.212024
SkyPIE: A Fast & Accurate Oracle for Object Placement · Proc. ACM Manag. Data 2024
Cloud and datacenter computing › multi-tenancy
multi-tenant scheduling
0.212015
Blind men and an elephant coalescing open-source, academic, and industrial perspectives on BigData · ICDE 2015
Distributed systems
fault tolerance
0.222017
Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017
REEF: Retainable Evaluator Execution Framework · SIGMOD Conference 2015
Cloud and datacenter computing › multi-tenancy
tenant isolation
0.112021
QuiCK: A Queuing System in CloudKit · SIGMOD Conference 2021
Storage systems › object storage
cloud object store
0.112012
Walnut: a unified cloud object store · SIGMOD Conference 2012
Storage systems
object storage
0.112012
Walnut: a unified cloud object store · SIGMOD Conference 2012
Distributed and cloud data management
big data systems
0.112015
Blind men and an elephant coalescing open-source, academic, and industrial perspectives on BigData · ICDE 2015
Distributed systems
distributed coordination
0.112015
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015
Distributed systems
distributed coordination and fault tolerance
0.112015
Blind men and an elephant coalescing open-source, academic, and industrial perspectives on BigData · ICDE 2015
Cloud and datacenter computing
resource allocation
0.112015
Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters · USENIX ATC 2015
Cloud and datacenter computing
cloud data management
0.012012
Walnut: a unified cloud object store · SIGMOD Conference 2012
Cloud and datacenter computing › multi-tenancy
multi-tenant storage
0.012012
Walnut: a unified cloud object store · SIGMOD Conference 2012

Methods — techniques the papers use, named apart from their topics

integer linear programming · 1.5convex optimization · 1.5geometric algorithms · 0.8geometric algorithm · 0.8two-level sharding · 0.5FoundationDB Record Layer · 0.5collaborative filtering · 0.3task scheduling · 0.2state management · 0.2data caching · 0.2elasticity · 0.1
YearPublicationVenuePosition
2024 SkyPIE: A Fast & Accurate Oracle for Object Placement
abstract
Cloud object stores offer vastly different price points for object storage as a function of workload and geography. Poor object placement can thus lead to significant cost overheads. Prior cost-saving techniques attempt to optimize placement policies on the fly, deciding object placements for each object individually. In practice, these techniques do not scale to the size of the modern cloud. In this work, we leverage the static nature and pay-per-use pricing model of cloud environments to explore a different approach. Rather than computing object placements on the fly, we precompute a SkyPIE oracle---a lookup structure representing all possible placement policies and the workloads for which they are optimal. Internally, SkyPIE represents placement policies as a matrix of cost-hyperplanes, which we effectively precompute through pruning and convex optimization. By leveraging a fast geometric algorithm, online queries then are 1 to 8 orders of magnitude faster but as accurate as Integer-Linear-Programming. This makes exact optimization tractable for real workloads and we show >10x cost savings compared to state-of-the-art heuristic approaches.
Tiemo Bang, Chris Douglas, Natacha Crooks, Joseph M. Hellerstein
Proc. ACM Manag. Data2
2021 QuiCK: A Queuing System in CloudKit
abstract
We present QuiCK, a queuing system built for managing asynchronous tasks in CloudKit, Apple's storage backend service. QuiCK stores queued messages along with user data in CloudKit, and supports CloudKit's tenancy model including isolation, fair resource allocation, observability, and tenant migration. QuiCK is built on the FoundationDB Record Layer, an open source transactional DBMS. It employs massive two-level sharding, with tens of billions of queues on the first level (separately storing the queued items for each user of every CloudKit app), and hundreds of queues on a second level (one per FoundationDB cluster used by CloudKit). Our evaluation demonstrates that QuiCK scales linearly with additional consumer resources, effectively avoids contention, provides fairness across CloudKit tenants, and executes deferred tasks with low latency.
Kfir Lev-Ari, Yizuo Tian, Alexander Shraer, Chris Douglas, Andrey Andreev, Kevin Beranek, Scott Dugas, Alec Grieser, Jeremy Hemmo
SIGMOD Conference4
2019 Hydra: a federated resource manager for data-center scale analytics
Carlo Curino, Subru Krishnan, Konstantinos Karanasos, Sriram Rao, Giovanni Matteo Fumarola, Botong Huang, Kishore Chaliparambil, Arun Suresh, Young Chen, Solom Heddaya, Roni Burd, Sarvesh Sakalanaga, Chris Douglas, Bill Ramsey, Raghu Ramakrishnan 0001
NSDI13
2018 Netco: Cache and I/O Management for Analytics over Disaggregated Stores
abstract
We consider a common setting where storage is disaggregated from the compute in data-parallel systems. Colocating caching tiers with the compute machines can reduce load on the interconnect but doing so leads to new resource management challenges. We design a system Netco, which prefetches data into the cache (based on workload predictability), and appropriately divides the cache space and network bandwidth between the prefetches and serving ongoing jobs. Netco makes various decisions (what content to cache, when to cache and how to apportion bandwidth) to support end-to-end optimization goals such as maximizing the number of jobs that meet their service-level objectives (e.g., deadlines). Our implementation of these ideas is available within the open-source Apache HDFS project. Experiments on a public cloud, with production-trace inspired workloads, show that Netco uses up to 5x less remote I/O compared to existing techniques and increases the number of jobs that meet their deadlines up to 80%.
Virajith Jalaparti, Chris Douglas, Mainak Ghosh, Ashvin Agrawal, Avrilia Floratou, Srikanth Kandula, Ishai Menache, Joseph Naor, Sriram Rao
SoCC2
2017 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics
abstract
Azure Data Lake Store (ADLS) is a fully-managed, elastic, scalable, and secure file system that supports Hadoop distributed file system (HDFS) and Cosmos semantics. It is specifically designed and optimized for a broad spectrum of Big Data analytics that depend on a very high degree of parallel reads and writes, as well as collocation of compute and data for high bandwidth and low-latency access. It brings together key components and features of Microsoft?s Cosmos file system-long used by internal customers at Microsoft and HDFS, and is a unified file storage solution for analytics on Azure. Internal and external workloads run on this unified platform. Distinguishing aspects of ADLS include its design for handling multiple storage tiers, exabyte scale, and comprehensive security and data sharing features. We present an overview of ADLS architecture, design points, and performance.
Raghu Ramakrishnan 0001, Baskar Sridharan, John R. Douceur, Pavan Kasturi, Balaji Krishnamachari-Sampath, Karthick Krishnamoorthy, Mitica Manu, Spiro Michaylov, Rogério Ramos, Neil Sharman, Zee Xu, Youssef Barakat, Chris Douglas, Richard Draves, Shrikant S. Naidu, Shankar Shastry, Atul Sikaria, Simon Sun, Ramarathnam Venkatesan
SIGMOD Conference14
2017 Apache REEF: Retainable Evaluator Execution Framework
abstract
Resource Managers like YARN and Mesos have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault tolerance, task scheduling and coordination) and reimplement common mechanisms (e.g., caching, bulk-data transfers). This article presents REEF, a development framework that provides a control plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource reuse for data caching and state management abstractions that greatly ease the development of elastic data processing pipelines on cloud platforms that support a Resource Manager service. We illustrate the power of REEF by showing applications built atop: a distributed shell application, a machine-learning framework, a distributed in-memory caching system, and a port of the CORFU system. REEF is currently an Apache top-level project that has attracted contributors from several institutions and it is being used to develop several commercial offerings such as the Azure Stream Analytics service.
Byung-Gon Chun, Tyson Condie, Yingda Chen, Carlo Curino, Chris Douglas, Matteo Interlandi, Beomyeol Jeon, Joo Seong Jeong, Gyewon Lee, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Mariia Mykhailova, Shravan M. Narayanamurthy, Joseph Noor, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Taegeon Um, Julia Wang, Markus Weimer, Youngseok Yang
ACM Trans. Comput. Syst.7
2015 Blind men and an elephant coalescing open-source, academic, and industrial perspectives on BigData
abstract
This tutorial is organized in two parts. In the first half, we will present an overview of applications and services in the BigData ecosystem. We will use known distributed database and systems literature as landmarks to orient the attendees in this fast-evolving space. Throughout, we will contrast models of resource management, performance, and the constraints that shape the architectures of prominent systems. We will also discuss the role of academia and industry in the development of open-source infrastructure, with an emphasis on open problems and strategies for collaboration. We assume only basic familiarity with distributed systems. In the second half, we will delve into Apache Hadoop YARN. YARN (Yet Another Resource Negotiator) transformed Hadoop from a MapReduce engine to a general-purpose cluster scheduler. Since its introduction, it has been deployed in production and extended to support use cases beyond large-scale batch processing. The tutorial will present the active research and development supporting such heterogeneous workloads, with particular attention to multi-tenant scheduling. Topics include security, resource isolation, protocols, and preemption. This portion will be detailed, but accessible to anyone with a background in distributed systems and all attendees of the first half of the tutorial.
Chris Douglas, Carlo Curino
ICDE1
2015 REEF: Retainable Evaluator Execution Framework
abstract
Resource Managers like Apache YARN have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low-level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault-tolerance, task scheduling and coordination) and re-implement common mechanisms (e.g., caching, bulk-data transfers). This paper presents REEF, a development framework that provides a control-plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource re-use for data caching, and state management abstractions that greatly ease the development of elastic data processing work-flows on cloud platforms that support a Resource Manager service. REEF is being used to develop several commercial offerings such as the Azure Stream Analytics service. Furthermore, we demonstrate REEF development of a distributed shell application, a machine learning algorithm, and a port of the CORFU [4] system. REEF is also currently an Apache Incubator project that has attracted contributors from several instititutions.
Markus Weimer, Yingda Chen, Byung-Gon Chun, Tyson Condie, Carlo Curino, Chris Douglas, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Shravan M. Narayanamurthy, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Julia Wang
SIGMOD Conference6
2015 Mercury: Hybrid Centralized and Distributed Scheduling in Large Shared Clusters
Konstantinos Karanasos, Sriram Rao, Carlo Curino, Chris Douglas, Kishore Chaliparambil, Giovanni Matteo Fumarola, Solom Heddaya, Raghu Ramakrishnan 0001, Sarvesh Sakalanaga
USENIX ATC4
2014 Reservation-based Scheduling: If You're Late Don't Blame Us!
abstract
The continuous shift towards data-driven approaches to business, and a growing attention to improving return on investments (ROI) for cluster infrastructures is generating new challenges for big-data frameworks. Systems originally designed for big batch jobs now handle an increasingly complex mix of computations. Moreover, they are expected to guarantee stringent SLAs for production jobs and minimize latency for best-effort jobs.
Carlo Curino, Djellel Eddine Difallah, Chris Douglas, Subru Krishnan, Raghu Ramakrishnan 0001, Sriram Rao
SoCC3
2013 Apache Hadoop YARN: yet another resource negotiator
abstract
The initial design of Apache Hadoop [1] was tightly focused on running massive, MapReduce jobs to process a web crawl. For increasingly diverse companies, Hadoop has become the data and computational agorá---the de facto place where data and computational resources are shared and accessed. This broad adoption and ubiquitous usage has stretched the initial design well beyond its intended target, exposing two key shortcomings: 1) tight coupling of a specific programming model with the resource management infrastructure, forcing developers to abuse the MapReduce programming model, and 2) centralized handling of jobs' control flow, which resulted in endless scalability concerns for the scheduler.
Vinod Kumar Vavilapalli, Arun C. Murthy, Chris Douglas, Sharad Agarwal, Mahadev Konar, Robert Evans, Thomas Graves, Jason Lowe, Hitesh Shah, Siddharth Seth, Bikas Saha, Carlo Curino, Owen O'Malley, Sanjay Radia, Benjamin C. Reed, Eric Baldeschwieler
SoCC3
2012 True elasticity in multi-tenant data-intensive compute clusters
abstract
Data-intensive computing (DISC) frameworks scale by partitioning a job across a set of fault-tolerant tasks, then diffusing those tasks across large clusters. Multi-tenanted clusters must accommodate service-level objectives (SLO) in their resource model, often expressed as a maximum latency for allocating the desired set of resources to every job. When jobs are partitioned into tasks statically, a cluster cannot meet its SLOs while maintaining both high utilization and efficiency. Ideally, we want to give resources to jobs when they are free but would expect to reclaim them instantaneously when new jobs arrive, without losing work. DISC frameworks do not support such elasticity because interrupting running tasks incurs high overheads. Amoeba enables lightweight elasticity in DISC frameworks by identifying points at which running tasks of over-provisioned jobs can be safely exited, committing their outputs, and spawning new tasks for the remaining work. Effectively, tasks of DISC jobs are now sized dynamically in response to global resource scarcity or abundance. Simulation and deployment of our prototype shows that Amoeba speeds up jobs by 32% without compromising utilization or efficiency.
Ganesh Ananthanarayanan, Chris Douglas, Raghu Ramakrishnan 0001, Sriram Rao, Ion Stoica
SoCC2
2012 Walnut: a unified cloud object store
abstract
Walnut is an object-store being developed at Yahoo! with the goal of serving as a common low-level storage layer for a variety of cloud data management systems including Hadoop (a MapReduce system), MObStor (a multimedia serving system), and PNUTS (an extended key-value serving system). Thus, a key performance challenge is to meet the latency and throughput requirements of the wide range of workloads commonly observed across these diverse systems. The motivation for Walnut is to leverage a carefully optimized low-level storage system, with support for elasticity and high-availability, across all of Yahoo!'s data clouds. This would enable sharing of hardware resources across hitherto siloed clouds of different types, offering greater potential for intelligent load balancing and efficient elastic operation, and simplify the operational tasks related to data storage.
Jianjun Chen 0001, Chris Douglas, Michi Mutsuzaki, Patrick Quaid, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears
SIGMOD Conference2