Luke Marshall

dblp:207/0065 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0003-0633-4004ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Cloud and datacenter computing · 55% Distributed systems · 10% High-performance computing · 9%
Artificial intelligence
2 papers
Reinforcement learning · 79% Efficient and distributed learning · 21%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
1.332023
Kerveros: Efficient and Scalable Cloud Admission Control · OSDI 2023
Protean: VM Allocation Service at Scale · OSDI 2020
Hindsight Learning for MDPs with Exogenous Inputs · ICML 2023
High-performance computing
collective communication
0.812024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Parallel and multicore computing › parallel scheduling
communication scheduling
0.812024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Electronic design automation › physical design › routing
multicommodity flow
0.812024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.712023
Hindsight Learning for MDPs with Exogenous Inputs · ICML 2023
Distributed and cloud data management › cloud database
database-as-a-service
0.712023
Solver-In-The-Loop Cluster Resource Management for Database-as-a-Service · Proc. VLDB Endow. 2023
Embedded and real-time systems › real-time scheduling
admission control
0.712023
Kerveros: Efficient and Scalable Cloud Admission Control · OSDI 2023
Cloud and datacenter computing
cloud platform
0.712023
Hyrax: Fail-in-Place Server Operation in Cloud Platforms · OSDI 2023
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.712023
Solver-In-The-Loop Cluster Resource Management for Database-as-a-Service · Proc. VLDB Endow. 2023
Distributed systems
fault tolerance
0.712023
Hyrax: Fail-in-Place Server Operation in Cloud Platforms · OSDI 2023
Cloud and datacenter computing › multi-tenancy
tenant placement
0.712023
Solver-In-The-Loop Cluster Resource Management for Database-as-a-Service · Proc. VLDB Endow. 2023
Cloud and datacenter computing › resource provisioning
virtual machine provisioning
0.622023
Protean: VM Allocation Service at Scale · OSDI 2020
Hindsight Learning for MDPs with Exogenous Inputs · ICML 2023
Cloud and datacenter computing
virtualization
0.412020
Protean: VM Allocation Service at Scale · OSDI 2020
Machine learning › Efficient and distributed learning
distributed training
0.212024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Distributed systems
distributed coordination
0.112020
Protean: VM Allocation Service at Scale · OSDI 2020
Cloud and datacenter computing
resource allocation
0.112020
Protean: VM Allocation Service at Scale · OSDI 2020

Methods — techniques the papers use, named apart from their topics

traffic engineering · 1.5hindsight learning · 1.3counterfactual reasoning · 1.3constraint solving · 1.3combinatorial optimization · 1.3multicommodity flow · 0.8multi-commodity flow · 0.8scheduling · 0.4VM allocation · 0.4
YearPublicationVenuePosition
2024 Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem
abstract
Cloud operators utilize collective communication optimizers to enhance the efficiency of the single-tenant, centrally managed training clusters they manage. However, current optimizers struggle to scale for such use cases and often compromise solution quality for scalability. Our solution, TE-CCL, adopts a traffic-engineering-based approach to collective communication. Compared to a state-of-the-art optimizer, TACCL, TE-CCL produced schedules with 2× better performance on topologies TACCL supports (and its solver took a similar amount of time as TACCL's heuristic-based approach). TECCL additionally scales to larger topologies than TACCL. On our GPU testbed, TE-CCL outperformed TACCL by 2.14× and RCCL by 3.18× in terms of algorithm bandwidth.
Xuting Liu 0003, Behnaz Arzani, Siva Kesava Reddy K., Liangyu Zhao, Vincent Liu 0001, Srikanth Kandula, Luke Marshall
SIGCOMM8
2023 Hindsight Learning for MDPs with Exogenous Inputs
abstract
Many resource management problems require sequential decision-making under uncertainty, where the only uncertainty affecting the decision outcomes are exogenous variables outside the control of the decision-maker. We model these problems as Exo-MDPs (Markov Decision Processes with Exogenous Inputs) and design a class of data-efficient algorithms for them termed Hindsight Learning (HL). Our HL algorithms achieve data efficiency by leveraging a key insight: having samples of the exogenous variables, past decisions can be revisited in hindsight to infer counterfactual consequences that can accelerate policy improvements. We compare HL against classic baselines in the multi-secretary and airline revenue management problems. We also scale our algorithms to a business-critical cloud resource management problem – allocating Virtual Machines (VMs) to physical machines, and simulate their performance with real datasets from a large public cloud provider. We find that HL algorithms outperform domain-specific heuristics, as well as state-of-the-art reinforcement learning methods.
Sean R. Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall, Hugo Barbalho, Jennifer Neville, Ishai Menache, Adith Swaminathan
ICML4
2023 Hyrax: Fail-in-Place Server Operation in Cloud Platforms
Jialun Lyu, Marisa You, Celine Irvene, Mark Jung, Tyler Narmore, Jacob Shapiro, Luke Marshall, Savyasachi Samal, Ioannis Manousakis, Lisa Hsu, Preetha Subbarayalu, Ashish Raniwala, Brijesh Warrier, Ricardo Bianchini, Bianca Schroeder, Daniel S. Berger
OSDI7
2023 Kerveros: Efficient and Scalable Cloud Admission Control
Sultan Mahmud Sajal, Luke Marshall, Beibin Li, Shandan Zhou, Abhisek Pan, Konstantina Mellou, Deepak Narayanan, Timothy Zhu, David Dion, Thomas Moscibroda, Ishai Menache
OSDI2
2023 Solver-In-The-Loop Cluster Resource Management for Database-as-a-Service
abstract
In Database-as-a-Service (DBaaS) clusters, resource management is a complex optimization problem that assigns tenants to nodes, subject to various constraints and objectives. Tenants share resources within a node, however, their resource demands can change over time and exhibit high variance. As tenants may accumulate large state, moving them to a different node becomes disruptive, making intelligent placement decisions crucial to avoid service disruption. Placement decisions need to account for dynamic changes in tenant resource demands, different causes of service disruption, and various placement constraints, giving rise to a complex search space. In this paper, we show how to bring combinatorial solvers to bear on this problem, formulating the objective of minimizing service disruption as an optimization problem amenable to fast solutions. We implemented our approach in the Service Fabric cluster manager codebase. Experiments show significant reductions in constraint violations and tenant moves, compared to the previous state-of-the-art, including the unmodified Service Fabric cluster manager, as well as recent research on DBaaS tenant placement.
Arnd Christian König, Karan Newatia, Luke Marshall, Vivek R. Narasayya
Proc. VLDB Endow.4
2020 Optimizing Onsite Food Services at Scale
abstract
Large food-service companies typically support a wide range of operations (catering, vending machines, repairs), each with different operational characteristics (manpower, vehicles, tools, timing constraints, etc.). While the advances in Internet-based technologies facilitate the adoption of automated scheduling systems, the complexity and heterogeneity of the different operations hinders the design of comprehensive optimization solutions. Indeed, our collaboration with Compass Group, one of the largest food-service companies in the world, reveals that many of its workforce assignments are done manually due to the lack of scheduling solutions that can accommodate the complexity of operational constraints. Further, the diversity in the nature of operations prevents collaboration and sharing of resources among various services such as catering and beverage distribution, leading to an inflated fleet size.
Konstantina Mellou, Luke Marshall, Krishna Chintalapudi, Patrick Jaillet, Ishai Menache
SIGSPATIAL/GIS2
2020 Protean: VM Allocation Service at Scale
Ori Hadary, Luke Marshall, Ishai Menache, Abhisek Pan, Esaias E. Greeff, David Dion, Star Dorminey, Shailesh Joshi, Mark Russinovich, Thomas Moscibroda
OSDI2
2019 Multi-Itinerary Optimization as Cloud Service (Industrial Paper)
abstract
In this paper, we describe Multi-Itinerary Optimization (MIO) - a novel Bing maps service that automates the process of building itineraries for multiple agents while optimizing their routes to save travel time or distance. MIO can be used by organizations with a fleet of vehicles and drivers, mobile salesforce, or a team of personnel in the field in order to maximize workforce efficiency. MIO accounts for service time windows, duration, and priority, as well as traffic conditions between locations, resulting in challenging algorithmic problems at multiple levels (e.g., calculating travel-time distance matrices at scale, scheduling services for multiple agents).
Alexandru Cristian, Luke Marshall, Mihai Negrea, Flavius Stoichescu, Peiwei Cao, Ishai Menache
SIGSPATIAL/GIS2