Barry Sly-Delgado

dblp:289/4131 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-2101-1396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SciWIND: Effectively Exploiting Node-Local Storage for Data-Intensive High-Energy Physics Workflows
Colin Thomas, Barry Sly-Delgado, Connor Moore, Benjamín Tovar, Kevin Lannon, Douglas Thain
IPDPS3
2024 Reshaping High Energy Physics Applications for Near-Interactive Execution Using TaskVine
abstract
High energy physics experiments produce petabytes of data annually that must be reduced to gain insight into the laws of nature. Early-stage reduction executes long-running, high-throughput workflows across thousands of nodes spanning multiple facilities to produce shared datasets. Later stages are typically written by individuals or small groups and must be refined and re-run many times for correctness. Reducing iteration times of later stages is key to accelerating discovery. We demonstrate our experience reshaping late-stage analysis applications on thousands of nodes. It is not enough merely to increase scale: it is necessary to make changes throughout the stack, including storage systems, data management, task scheduling, and application design. We demonstrate these changes when applied to two analysis applications built on open source data analysis frameworks (Coffea, Dask, TaskVine). We evaluate the performance of the applications on opportunistic campus clusters, showing effective scaling up to 7200 cores, thus producing significant speedup.
Barry Sly-Delgado, Benjamín Tovar, Douglas Thain
SC1
2022 Dynamic Task Shaping for High Throughput Data Analysis Applications in High Energy Physics
abstract
Distributed data analysis frameworks are widely used for processing large datasets generated by instruments in scientific fields such as astronomy, genomics, and particle physics. Such frameworks partition petabyte-size datasets into chunks and execute many parallel tasks to search for common patterns, locate unusual signals, or compute aggregate properties. When well-configured, such frameworks make it easy to churn through large quantities of data on large clusters. However, configuring frameworks presents a challenge for end users, who must select a variety of parameters such as the blocking of the input data, the number of tasks, the resources allocated to each task, and the size of nodes on which they run. If poorly configured, the result may perform many orders of magnitude worse than optimal, or the application may even fail to make progress at all. Even if a good configuration is found through painstaking observations, the performance may change drastically when the input data or analysis kernel changes. This paper considers the problem of automatically configuring a data analysis application for high energy physics (TopEFT) built upon standard frameworks for physics analysis (Coffea) and distributed tasking (Work Queue). We observe the inherent variability within the application, demonstrate the problems of poor configuration, and then develop several techniques for automatically sizing tasks to meet goals of resource consumption, and overall application completion.
Benjamín Tovar, Ben Lyons, Kelci Mohrman, Barry Sly-Delgado, Kevin Lannon, Douglas Thain
IPDPS4
2020 Leveraging Location Based Services for Analysis of Suicides in Rural Counties
Barry Sly-Delgado, Meysam Ghaffari, Aastha Gunjan, Cynthia Brandt, Joseph L. Goulet, Karen H. Wang, Ashok Srinivasan
AMIA1