Scott Michelson

dblp:189/0126 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 92% Distributed systems · 8%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
1.222023
Global Capacity Management With Flux · OSDI 2023
RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation · SOSP 2021
Cloud and datacenter computing
resource allocation
0.722023
RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation · SOSP 2021
Global Capacity Management With Flux · OSDI 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.722020
Twine: A Unified Cluster Management System for Shared Infrastructure · OSDI 2020
Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services · OSDI 2016
Cloud and datacenter computing
resource management
0.412020
Twine: A Unified Cluster Management System for Shared Infrastructure · OSDI 2020
Cloud and datacenter computing
datacenter operations
0.422021
Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services · OSDI 2016
RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation · SOSP 2021
Cloud and datacenter computing › datacenter operations
datacenter reliability
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Distributed systems
fault tolerance
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018

Methods — techniques the papers use, named apart from their topics

mixed integer programming · 0.5
YearPublicationVenuePosition
2023 Global Capacity Management With Flux
Marius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall, Po-Yen Chou, Richard Cornew, Adela Kabiljo, Ranjith Kumar S., Maroo Lieuw, Justin Meza, Scott Michelson, Thomas Rohloff, Hayley Russell, Jeff Qin, Chunqiang Tang
OSDI11
2021 RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation
abstract
Capacity reservation is a common offering in public clouds and on-premise infrastructure. However, no prior work provides capacity reservation with SLO guarantees that takes into account random and correlated hardware failures, datacenter maintenance, and heterogeneous hardware. In this paper, we describe how Facebook's region-scale Resource Allowance System (RAS) addresses these issues and provides guaranteed capacity. RAS uses a capacity abstraction called reservation to represent a set of servers dynamically assigned to a logical cluster. We take a two-level approach to scale resource allocation to all datacenters in a region, where a mixed-integer-programming solver continuously optimizes server-to-reservation assignments off the critical path, and a traditional container allocator does real-time placement of containers on servers in a reservation. As a relatively new component of Facebook's 10-year old cluster manager Twine, RAS has been running in production for almost two years, continuously optimizing the allocation of millions of servers to thousands of reservations. We describe the design of RAS and share our experience of deploying it at scale.
Andrew Newell, Dimitrios Skarlatos 0002, Maxim Khutornenko, Mayank Pundir, Yuanlai Liu, Linh Le, Brendon Daugherty, Apurva Samudra, Prashasti Baid, James Kneeland, Igor Kabiljo, Dmitry Shchukin, Andre Rodrigues, Scott Michelson, Ben Christensen, Kaushik Veeraraghavan, Chunqiang Tang
SOSP18
2020 Twine: A Unified Cluster Management System for Shared Infrastructure
Chunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor, Scott Michelson, Thawan Kooburat, Aravind Anbudurai, Kabir Gogia, Ben Christensen, Alex Gartrell, Maxim Khutornenko, Sachin Kulkarni, Marcin Pawlowski 0003, Tuomas Pelkonen, Andre Rodrigues, Rounak Tibrewal, Vaishnavi Venkatesan, Peter Zhang
OSDI5
2018 Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently
Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankaralingam Panneerselvam, Alex Gyori, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, Tianyin Xu
OSDI3
2016 Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services
Kaushik Veeraraghavan, Justin Meza, Wonho Kim, Sonia Margulis, Scott Michelson, Rajesh Nishtala, Daniel Obenshain, Dmitri Perelman, Yee Jiun Song
OSDI6