Brian D. Friedman

dblp:31/9986 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0003-3534-8787ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 100%
Artificial intelligence
2 papers
Efficient and distributed learning · 100%
Computer networks
1 paper
Edge and fog computing · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 77% Data mining · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
inference serving
0.812024
Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling · ASPLOS (1) 2024
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
An efficient and non-intrusive GPU scheduling framework for deep learning training systems · SC 2020
Edge and fog computing
edge inference
0.212024
Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling · ASPLOS (1) 2024
Computational social science and digital humanities
social network analysis
0.212013
Enterprise social network analysis and modeling: A tale of two graphs · INFOCOM 2013
Recommender systems › user modeling
user interaction modeling
0.212013
Enterprise social network analysis and modeling: A tale of two graphs · INFOCOM 2013
Machine learning › Efficient and distributed learning
distributed training
0.112020
An efficient and non-intrusive GPU scheduling framework for deep learning training systems · SC 2020

Methods — techniques the papers use, named apart from their topics

model variant selection · 2.3joint optimization · 2.3adaptive batching · 2.3sidecar process · 0.9kubernetes plugins · 0.9adaptive scheduling · 0.9statistical modeling · 0.3graph analysis · 0.3
YearPublicationVenuePosition
2024 Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling
abstract
Existing machine learning inference-serving systems largely rely on hardware scaling by adding more devices or using more powerful accelerators to handle increasing query demands. However, hardware scaling might not be feasible for fixed-size edge clusters or private clouds due to their limited hardware resources. A viable alternate solution is accuracy scaling, which adapts the accuracy of ML models instead of hardware resources to handle varying query demands. This work studies the design of a high-throughput inference-serving system with accuracy scaling that can meet throughput requirements while maximizing accuracy. To achieve the goal, this work proposes to identify the right amount of accuracy scaling by jointly optimizing three sub-problems: how to select model variants, how to place them on heterogeneous devices, and how to assign query workloads to each device. It also proposes a new adaptive batching algorithm to handle variations in query arrival times and minimize SLO violations. Based on the proposed techniques, we build an inference-serving system called Proteus and empirically evaluate it on real-world and synthetic traces. We show that Proteus reduces accuracy drop by up to 3× and latency timeouts by 2--10× with respect to baseline schemes, while meeting throughput requirements.
Sohaib Ahmad, Hui Guan 0001, Brian D. Friedman, Thomas Williams, Ramesh K. Sitaraman, Thomas Y. C. Woo
ASPLOS (1)3
2020 An efficient and non-intrusive GPU scheduling framework for deep learning training systems
abstract
Efficient GPU scheduling is the key to minimizing the execution time of the Deep Learning (DL) training workloads. DL training system schedulers typically allocate a fixed number of GPUs to each job, which inhibits high resource utilization and often extends the overall training time. The recent introduction of schedulers that can dynamically reallocate GPUs has achieved better cluster efficiency. This dynamic nature, however, introduces additional overhead by terminating and restarting jobs or requires modification to the DL training frameworks.We propose and develop an efficient, non-intrusive GPU scheduling framework that employs a combination of an adaptive GPU scheduler and an elastic GPU allocation mechanism to reduce the completion time of DL training workloads and improve resource utilization. Specifically, the adaptive GPU scheduler includes a scheduling algorithm that uses training job progress information to determine the most efficient allocation and reallocation of GPUs for incoming and running jobs at any given time. The elastic GPU allocation mechanism works in concert with the scheduler. It offers a lightweight and nonintrusive method to reallocate GPUs based on a “SideCar” process that temporarily stops and restarts the job's DL training process with a different number of GPUs. We implemented the scheduling framework as plugins in Kubernetes and conducted evaluations on two 16-GPU clusters with multiple training jobs based on TensorFlow. Results show that our proposed scheduling framework reduces the overall execution time and the average job completion time by up to 45% and 63%, respectively, compared to the Kubernetes default scheduler. Compared to a termination based scheduler, our framework reduces the overall execution time and the average job completion time by up to 20% and 37%, respectively.
Oscar J. Gonzalez, Xiaobo Zhou 0002, Thomas Williams, Brian D. Friedman, Martin Havemann, Thomas Y. C. Woo
SC5
2013 Enterprise social network analysis and modeling: A tale of two graphs
abstract
Like their public counterpart such as Facebook and Twitter, enterprise social networks are poised to revolutionize how people interact in the workplace. There is a pressing need to understand how people are using these social networks. Unlike the public social networks like Facebook or Twitter which are normally characterized using the social graph or the interaction graph, enterprise social networks are also governed by an organization graph. Based on a six month dataset collected from May through October 2011 of a large enterprise social network, we study the characteristics of activities of its enterprise social network. We observe that the user attributes in the organization graph such as geographic location (eg. country) and his/her rank in the company hierarchy have a significant impact on how the user uses the social network and how user interacts with each other. We then build formal statistical models of user interaction graphs in enterprise social network and quantify effects of user attributes from organization graphs on these interactions. Furthermore, as the enterprise social network medium bring users from diverse locations and social status forming ad-hoc communities, our statistical model can be further enhanced by including these ad-hoc communities.
Jin Cao 0002, Li Erran Li, Brian D. Friedman
INFOCOM4