Ryan D. Friese

dblp:202/6653 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-4121-2195ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 45% Parallel and multicore computing · 17% Energy-efficient computing · 17%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
resource management
0.522016
Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques · IEEE Trans. Parallel Distributed Syst. 2016
Utility Functions and Resource Management in an Oversubscribed Heterogeneous Computing Environment · IEEE Trans. Computers 2015
Energy-efficient computing
energy-aware scheduling
0.212016
Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques · IEEE Trans. Parallel Distributed Syst. 2016
Electronic design automation › high-level synthesis › scheduling
makespan minimization
0.212016
Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques · IEEE Trans. Parallel Distributed Syst. 2016
Parallel and multicore computing
task scheduling
0.212016
Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques · IEEE Trans. Parallel Distributed Syst. 2016
Cloud and datacenter computing › job scheduling › economic scheduling
utility-based scheduling
0.212015
Utility Functions and Resource Management in an Oversubscribed Heterogeneous Computing Environment · IEEE Trans. Computers 2015
GPUs and heterogeneous computing
heterogeneous computing systems
0.112016
Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques · IEEE Trans. Parallel Distributed Syst. 2016

Methods — techniques the papers use, named apart from their topics

pareto front analysis · 0.2linear programming · 0.2utility function · 0.2task dropping heuristics · 0.2simulation · 0.2
YearPublicationVenuePosition
2024 Automatic Extraction of Network Configurations for Realistic Simulation and Validation
abstract
In this work, we propose a framework to auto-tune the multiple network models' simulation configurations within SST/macro using Tree-structured Parzen Estimator-based Bayesian optimization to observe the effect on simulation accuracy across different message regimes. These regimes consist of small to large message sizes and latency to bandwidth-bound messages. We provide a detailed analysis of the simulation error for four representative HPC systems. Our Bayesian optimization-based autotuning framework for network models achieves a maximum of 5x improvement in accuracy over best-effort manual configurations based on available hardware specifications.
Joshua Suetterlein, Stephen J. Young, Jesun Sahariar Firoz, Joseph B. Manzano, Ryan D. Friese, Nathan R. Tallent, Kevin J. Barker, Timothy Stavenger
ISPASS5
2023 Surveillance mission scheduling with unmanned aerial vehicles in dynamic heterogeneous environments
Dylan Machovec, Howard Jay Siegel, James A. Crowder, Sudeep Pasricha, Anthony A. Maciejewski, Ryan D. Friese
J. Supercomput.6
2022 Towards Supporting Semiring in MLIR-Based COMET Compiler
abstract
Semirings are widely used in large-scale scientific applications of high-dimensional data and graph analytics for linear algebra computations. In this work, we propose a semiring compiler for today's high-performance computing (HPC) systems, often armed with heterogeneous devices, as an alternative to library-based approaches. In particular, we extend a domain-specific language (DSL) and compiler framework to automatically generate kernel code for semiring operations within the COMpiler for Extreme Targets (COMET) based on the Multi-Level Intermediate Representation (MLIR) framework. We provide a high-level programming abstraction representing various semiring operations with the familiar Einstein notation. We also build a semiring dialect and efficient code generation based on MLIR's extensible framework that can process a variety of semiring operators. By leveraging high-level semantics information and progressive lowering in code generation, we achieved better performance with up to 3.8x speedup compared with operations in the LAGraph library.
Luanzheng Guo, Rizwan A. Ashraf, Ryan D. Friese, Gokcen Kestor
PACT3
2020 Effectively Using Remote I/O For Work Composition in Distributed Workflows
abstract
Distributed scientific workflows are becoming more important with the interest in incorporating AI into their loops. A critical programming and performance question is how to compose workflow tasks when data is produced on one system but must be consumed on another. Since the dominant technique is composition with remote I/O, this paper explores its performance expectations. We describe BigFlowSim, a workflow I/O simulator that captures key implementation choices for remote I/O, including intensity, reuse, locality, access pattern, and data movement. With BigFlowSim, we generate a synthetic benchmark. We quantify the effects of each parameter with a performance sensitivity study. We explain trends in terms of data movement reduction and show that, under certain conditions, it is possible to establish a total order among most parameters.
Ryan D. Friese, Burcu Ozcelik Mutlu, Nathan R. Tallent, Joshua Suetterlein, Jan Strube 0001
IEEE BigData1
2020 Rapid Memory Footprint Access Diagnostics
abstract
Footprint and reuse distance measure temporal locality and therefore do not capture the significance of access patterns (spacial locality). A strided access pattern has the largest possible footprint but usually has the best performance. To highlight exposed memory latency, we separate footprint into strided (prefetchable) and irregular (non-prefetchable) access components and calculate the growth rate of each. To rapidly compute these footprint access diagnostics, we present two methods, whole-program and precise. Current footprint analyses can cause 200× or more slowdown with realistic inputs and are therefore impractical. Our whole-program method reduces the overhead to 10% by computing upper bounds, but still yields inter-procedural insight through a call path profile. Our precise method uses additional static analysis and profiling to refine the upper bounds for intra-procedural loop nests. We evaluate our approaches using benchmarks that vary access patterns (strided vs. unpredictable), sparsity (all words in a cache line vs. some), and reuse (varying and repeated accesses per element). Notably, for loop nests with unpredictable accesses, the precise method's accuracy is within 10% of ideal. The whole-program method has sufficient accuracy to diagnose bottlenecks.
Ozgur O. Kilic, Nathan R. Tallent, Ryan D. Friese
ISPASS3
2019 TAZeR: Hiding the Cost of Remote I/O in Distributed Scientific Workflows
abstract
Many scientific workflows access data derived from specialized instruments. When the data is analyzed, it is accessed over wide area networks, creating bottlenecks from long access latencies. We ask the question: assuming that data must be accessed remotely, can latencies be hidden without application change? We present TAZeR, a remote I/O framework that reduces effective data access latency. TAZeR transparently converts POSIX I/O into operations that interleave application work with data transfer, i.e., read prefetching and write stage-out. TAZeR ensures read data moves directly to application memory without synchronous intervention (soft zero-copy). TAZeR uses distributed bandwidth-aware staging to exploit data reuse across application tasks and to manage the capacity constraints of fast hierarchical storage. We evaluate TAZeR on a High Energy Physics workflow where two 1 Gb/s WAN links request remote data at 48 Gb/s using non-streaming access patterns. TAZeR is 12× and 22× faster than XRootD (state-of-the-art) and file copies (current approach), respectively; and within 7% of optimal. We explore conditions under which TAZeR can hide I/O accesses by showing performance as effective staging sizes change.
Joshua Suetterlein, Ryan D. Friese, Nathan R. Tallent, Malachi Schram
IEEE BigData2
2019 Rapidly Measuring Loop Footprints
abstract
Knowing a loop's footprint - the unique data items it accesses - enables important locality and capacity analysis. Unfortunately, current methods for computing footprint cause integer factors of slowdown and are therefore difficult to use with realistic inputs. Current methods are slow because they are based on software tracing of load instructions and data addresses. We present an approach that combines lightweight measurements and static binary analysis. We use static analysis to reason about each load's expected reuse. We then calculate the footprint for each loop using inference rules informed by measurements from common performance counters. We validate our method on benchmarks that vary access patterns (strided vs. unpredictable), reuse (varying and repeated accesses per element), and sparsity (all words in a cache line vs. some). For strided patterns, errors are within 1%; for unpredictable ones, errors are 5-10%. Our overheads are under 10%. A tool based on software tracing has an error of less than 1%, but either introduces at least 130× overhead or hangs.
Ozgur O. Kilic, Nathan R. Tallent, Ryan D. Friese
CLUSTER3
2018 Optimizing Distributed Data-Intensive Workflows
abstract
We present techniques for optimizing the performance of data-intensive workflows that execute on geographically distributed and heterogeneous resources. We optimize for both throughput and response time. Optimizing for throughput, we alleviate data-transfer bottlenecks. To hide access times of accessing remote data, we transparently introduce prefetching (overlapping data transfer and computation), without changing workflow source code. Optimizing for response time, we introduce intelligent scheduling for a set of high-priority tasks. We replace a greedy scheduler that assigns tasks without accounting for differing performance on heterogeneous resources, leading to long latencies. Intelligent scheduling rapidly selects a near-optimal solution for a bi-objective optimization problem. One objective is a good task assignment; the other objective is minimize I/O contention by distributing load across resources and time. To reason about task completion times, we use modeling tools to generate accurate predictions of execution times. We show performance results for Belle II workflow for high energy physics. The combination of these techniques can improve throughput over production Belle II configurations by 20-40%. Our work is general and adaptable to other distributed workflows.
Ryan D. Friese, Nathan R. Tallent, Malachi Schram, Mahantesh Halappanavar, Kevin J. Barker
CLUSTER1
2018 Stochastic Programming Approach for Resource Selection Under Demand Uncertainty
Tanveer Hossain Bhuiyan, Mahantesh Halappanavar, Ryan D. Friese, Hugh R. Medal, Luis de la Torre 0001, Arun V. Sathanur, Nathan R. Tallent
JSSPP3
2017 Pushing the Limits of Irregular Access Patterns on Emerging Network Architecture: A Case Study
abstract
Irregular applications pose considerable challenges to modern computer systems, especially in distributed environments, where traditional high-performance networks are optimized for large message transfers. In this work, we analyze performance of an irregular application proxy benchmark running over traditional MPI/Infiniband as well as over the Data Vortex network, an emerging network architecture optimized for small packets. Our goal is to identify the challenges posed by irregular applications and explore how emerging highperformance communication networks can cope with some of them. In particular, we analyze the performance of the Data Vortex interconnection network using the HPCC RandomAccess benchmark with different look-ahead buffer sizing to simulate varying degrees of irregularity. We use multiple implementations of the HPCC RandomAccess benchmark for both MPI/Infiniband and Data Vortex networks, with various degrees of optimizations. We compare Data Vortex results to a traditional high-bandwidth Infiniband network at scale and show that Data Vortex networks provide linear scalability and high performance.
Roberto Gioiosa, Thomas E. Warfel, Antonino Tumeo, Ryan D. Friese
CLUSTER4
2017 Generating Performance Models for Irregular Applications
abstract
Many applications have irregular behavior - e.g., input-dependent solvers, irregular memory accesses, or unbiased branches - that cannot be captured using today's automated performance modeling techniques. We describe new hierarchical critical path analyses for the Palm model generation tool. To obtain a good tradeoff between model accuracy, generality, and generation cost, we combine static and dynamic analysis. To create a model's outer structure, we capture tasks along representative MPI critical paths. We create a histogram of critical tasks with parameterized task arguments and instance counts. To model each task, we identify hot instruction-level paths and model each path based on data flow, data locality, and microarchitectural constraints. We describe application models that generate accurate predictions for strong scaling when varying CPU speed, cache and memory speed, microarchitecture, and (with supervision) input data class. Our models' errors are usually below 8%; and always below 13%.
Ryan D. Friese, Nathan R. Tallent, Abhinav Vishnu, Darren J. Kerbyson, Adolfy Hoisie
IPDPS1
2017 Towards Efficient Resource Allocation for Distributed Workflows Under Demand Uncertainties
Ryan D. Friese, Mahantesh Halappanavar, Arun V. Sathanur, Malachi Schram, Darren J. Kerbyson, Luis de la Torre 0001
JSSPP1
2016 HPC node performance and energy modeling with the co-location of applications
Daniel Dauwe, Eric Jonardi, Ryan D. Friese, Sudeep Pasricha, Anthony A. Maciejewski, David A. Bader, Howard Jay Siegel
J. Supercomput.3
2016 Energy and Makespan Tradeoffs in Heterogeneous Computing Systems using Efficient Linear Programming Techniques
abstract
Resource management for large-scale high performance computing systems pose difficult challenges to system administrators. The extreme scale of these modern systems require task scheduling algorithms that are capable of handling at least millions of tasks and thousands of machines. These large computing systems consume vast amounts of electricity leading to high operating costs. System administrators try to simultaneously reduce operating costs and offer state-of-the-art performance; however, these are often conflicting objectives. Highly scalable algorithms are necessary to schedule tasks efficiently and to help system administrators gain insight into energy/performance trade-offs of the system. System administrators can examine this trade-off space to quantify how much a difference in the performance level will cost in electricity, or analyze how much performance can be expected within an energy budget. In this study, we design a novel linear programming based resource allocation algorithm for a heterogeneous computing system to efficiently compute high quality solutions for simultaneously minimizing energy and makespan. These solutions are used to bound the Pareto front to easily trade-off energy and performance. The new algorithms are highly scalable in both solution quality and computation time compared to existing algorithms, especially as the problem size increases.
Kyle M. Tarplee, Ryan D. Friese, Anthony A. Maciejewski, Howard Jay Siegel, Edwin K. P. Chong
IEEE Trans. Parallel Distributed Syst.2
2015 Scalable linear programming based resource allocation for makespan minimization in heterogeneous computing systems
Kyle M. Tarplee, Ryan D. Friese, Anthony A. Maciejewski, Howard Jay Siegel
J. Parallel Distributed Comput.2
2015 Utility Functions and Resource Management in an Oversubscribed Heterogeneous Computing Environment
abstract
We model an oversubscribed heterogeneous computing system where tasks arrive dynamically and a scheduler maps the tasks to machines for execution. The environment and workloads are based on those being investigated by the Extreme Scale Systems Center at Oak Ridge National Laboratory. Utility functions that are designed based on specifications from the system owner and users are used to create a metric for the performance of resource allocation heuristics. Each task has a time-varying utility (importance) that the enterprise will earn based on when the task successfully completes execution. We design multiple heuristics, which include a technique to drop low utility-earning tasks, to maximize the total utility that can be earned by completing tasks. The heuristics are evaluated using simulation experiments with two levels of oversubscription. The results show the benefit of having fast heuristics that account for the importance of a task and the heterogeneity of the environment when making allocation decisions in an oversubscribed environment. The ability to drop low utility-earning tasks allow the heuristics to tolerate the high oversubscription as well as earn significant utility.
Bhavesh Khemka, Ryan D. Friese, Luis Diego Briceno, Howard Jay Siegel, Anthony A. Maciejewski, Gregory A. Koenig, Chris Groër, Gene Okonski, Marcia Hilton, Jendra Rambharos, Stephen W. Poole
IEEE Trans. Computers2
2013 Efficient and Scalable Computation of the Energy and Makespan Pareto Front for Heterogeneous Computing Systems
Kyle M. Tarplee, Ryan D. Friese, Anthony A. Maciejewski, Howard Jay Siegel
FedCSIS2
2013 Efficient and Scalable Pareto Front Generation for Energy and Makespan in Heterogeneous Computing Systems
Kyle M. Tarplee, Ryan D. Friese, Anthony A. Maciejewski, Howard Jay Siegel
WCO@FedCSIS2