EDBT 2026 Demo / reviewers in the wild / expert
Joshua Suetterlein
dblp:132/6950 · also Joshua D. Suetterlein
· DBLP profile ↗
13ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0003-0871-3307ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL MemoryabstractWe share early experiments with software development for shared CXL memory on the H3 Falcon C5022 CXL switch. We describe how the different stages in the development of a commercial RAG pipeline were impacted by the presence of CXL memory. All of Wikipedia is split into 46 million text passages which are embedded into a high-dimensional embedding space and subjected to heavy load in single-and multi-host experiments. We observe significant performance advantages available to DuckDB and Faiss without changing the software; we measure up to 12× latency reduction with real workloads and 40× reduction with synthetic workloads on CXL, compared to the same queries run with demand paging on NVMe. We explain the current challenges of working with two disjoint cache coherency domains and explain how the Fabric-Attached Memory File System (famfs) and famfs producer-consumer queues provide synchronization patterns without atomic operations on current ×86 CPUs. We then discuss open challenges remaining for high performance atomics and locking mechanisms in shared memory. Lastly we show how famfs page-level interleaving enables near-linear throughput scaling when a second host serves queries from the same Faiss index on shared fabric-attached memory and how shared memory allocations can be orchestrated by Kubernetes in a full vertical RAG deployment. Alfred Bratterud, Gisle Dankel, Amin Farajianzadeh, John Groves, Chengyi Juan, Joshua Suetterlein, Andrés Márquez 0001, Petter Gustad, Arnt Emil Ingulstad |
IEEE Trans. Computers | 6 |
| 2025 | Scaling Laws for the Workload Throughput of Emerging Heterogeneous ClustersabstractNext-generation HPC clusters are evolving into highly heterogeneous systems that integrate traditional computing resources with emerging accelerator technologies such as quantum processors, neuromorphic units, dataflow architectures, and specialized AI accelerators within a unified infrastructure. These advanced systems enable workloads to dynamically utilize different accelerators during various computation phases, creating complex execution patterns. The performance of the workloads can therefore be impacted by many factors, including how the accelerators are shared, their utilization, and their placement within the system. Moreover, effects such as the system and network state due to the overall system load can significantly impact the job completion rate. Understanding, identifying, and quantifying the impact of the most critical factors (e.g., the number of allocated accelerators) will help decide the investment decisions for accelerator acquisition and deployment that can improve the overall system throughput. This paper extensively studies these complex interactions among advanced accelerators within an HPC cluster and various workloads. We introduce a novel analytical model which predicts the speedup of a workload given an accelerator/system configuration. This model can be used to quantify the effect of augmenting additional accelerators on job performance running on an HPC cluster. We validate the model using both simulated and real environments. Akhil Alasandagutti, Joshua Suetterlein, Jesun Sahariar Firoz, Stephen J. Young, Joseph B. Manzano, Jason R. Stewart, Patrick G. Bridges, Trilce Estrada, Kevin J. Barker |
CCGrid | 2 |
| 2024 | Graph Analytics on Jellyfish topologyabstractBecause large unstructured datasets are important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the "best" network topologies for irregular communication, unstructured graph-based interconnects can be more suitable, due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.In addition to two common stencil-based mini-applications (LULESH and Sweep3D), we analyze three popular graph workloads – clustering, pattern enumeration (triangle counting), and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, considering relevant network routing schemes. Using packet-level simulations, we report average improvements of about 4–20% and 3–26% between equivalent Jellyfish vs. Dragonfly and Jellyfish vs. Fat tree topologies across diverse input graphs and applications. Md Nahid Newaz, Joshua Suetterlein, Nathan R. Tallent, Md Atiqul Mollah, Ming Hua 0003 |
IPDPS | 3 |
| 2024 | Automatic Extraction of Network Configurations for Realistic Simulation and ValidationabstractIn this work, we propose a framework to auto-tune the multiple network models' simulation configurations within SST/macro using Tree-structured Parzen Estimator-based Bayesian optimization to observe the effect on simulation accuracy across different message regimes. These regimes consist of small to large message sizes and latency to bandwidth-bound messages. We provide a detailed analysis of the simulation error for four representative HPC systems. Our Bayesian optimization-based autotuning framework for network models achieves a maximum of 5x improvement in accuracy over best-effort manual configurations based on available hardware specifications. Joshua Suetterlein, Stephen J. Young, Jesun Sahariar Firoz, Joseph B. Manzano, Ryan D. Friese, Nathan R. Tallent, Kevin J. Barker, Timothy Stavenger |
ISPASS | 1 |
| 2022 | Extending an asynchronous runtime system for high throughput applications: A case study
Joshua Suetterlein, Joseph B. Manzano, Andrés Márquez 0001, Guang R. Gao |
J. Parallel Distributed Comput. | 1 |
| 2021 | MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU serversabstractMulti-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-intensive workloads. In addition, these accelerators are increasingly being inter-connected in complex topologies and workloads are exhibiting a wider variety of inter-accelerator communication patterns. However, existing allocation policies are ill-suited for these emerging use-cases. Specifically, this work identifies that multi-accelerator workloads are commonly fragmented leading to reduced bandwidth and increased latency for inter-accelerator communication. Kiran Ranganath, Joshua Suetterlein, Joseph B. Manzano, Shuaiwen Song, Daniel Wong 0001 |
SC | 2 |
| 2020 | Effectively Using Remote I/O For Work Composition in Distributed WorkflowsabstractDistributed scientific workflows are becoming more important with the interest in incorporating AI into their loops. A critical programming and performance question is how to compose workflow tasks when data is produced on one system but must be consumed on another. Since the dominant technique is composition with remote I/O, this paper explores its performance expectations. We describe BigFlowSim, a workflow I/O simulator that captures key implementation choices for remote I/O, including intensity, reuse, locality, access pattern, and data movement. With BigFlowSim, we generate a synthetic benchmark. We quantify the effects of each parameter with a performance sensitivity study. We explain trends in terms of data movement reduction and show that, under certain conditions, it is possible to establish a total order among most parameters. Ryan D. Friese, Burcu Ozcelik Mutlu, Nathan R. Tallent, Joshua Suetterlein, Jan Strube 0001 |
IEEE BigData | 4 |
| 2020 | On the Marriage of Asynchronous Many Task Runtimes and Big Data: A GlanceabstractThe rise of the accelerator-based architectures and reconfigurable computing have showcased the weakness of software stack toolchains that still maintain a static view of the hardware instead of relying on a symbiotic relationship between static (e.g., compilers) and dynamic tools (e.g., runtimes). In the past decades, this need has given rise to adaptive runtimes with increasingly finer computational tasks. These finer tasks help to take advantage of the hardware by switching out when a long latency operation is encountered (because of the deeper memory hierarchies and new memory technologies that might target streaming instead of random access), thus trading off idle time for unrelated work. Examples of these finer task runtimes are Asynchronous Many Task (AMT) runtimes, in which highly efficient computational graphs run on a variety of hardware. Due to its inherent latency tolerant characteristics, latency-sensitive applications, such as Graph Analytics and Big Data can effectively use these runtimes. This paper aims to present an example of how the careful design of an AMT can exploit the hardware substrate when faced with high latency applications such as the ones given in the Big Data domain. Moreover, with its introspection and adaptive capabilities, we aim to show the power of these runtimes when facing the changing requirements of application workloads. We use the Performance Open Community Runtime (P-OCR) as our vehicle to demonstrate the concepts presented here. Joshua Suetterlein, Joseph B. Manzano, Andrés Márquez 0001, Guang R. Gao |
HiPC | 1 |
| 2019 | TAZeR: Hiding the Cost of Remote I/O in Distributed Scientific WorkflowsabstractMany scientific workflows access data derived from specialized instruments. When the data is analyzed, it is accessed over wide area networks, creating bottlenecks from long access latencies. We ask the question: assuming that data must be accessed remotely, can latencies be hidden without application change? We present TAZeR, a remote I/O framework that reduces effective data access latency. TAZeR transparently converts POSIX I/O into operations that interleave application work with data transfer, i.e., read prefetching and write stage-out. TAZeR ensures read data moves directly to application memory without synchronous intervention (soft zero-copy). TAZeR uses distributed bandwidth-aware staging to exploit data reuse across application tasks and to manage the capacity constraints of fast hierarchical storage. We evaluate TAZeR on a High Energy Physics workflow where two 1 Gb/s WAN links request remote data at 48 Gb/s using non-streaming access patterns. TAZeR is 12× and 22× faster than XRootD (state-of-the-art) and file copies (current approach), respectively; and within 7% of optimal. We explore conditions under which TAZeR can hide I/O accesses by showing performance as effective staging sizes change. Joshua Suetterlein, Ryan D. Friese, Nathan R. Tallent, Malachi Schram |
IEEE BigData | 1 |
| 2019 | A Parallel Graph Environment for Real-World Data Analytics WorkflowsabstractEconomic competitiveness and national security depend increasingly on the insightful analysis of large data sets. The diversity of real-world data sources and analytic workflows impose challenging hardware and software requirements for parallel graph platforms. The irregular nature of graph methods is not supported well by the deep memory hierarchies of conventional distributed systems, requiring new processor and runtime system designs to tolerate memory and synchronization latencies. Moreover, the efficiency of relational table operations and matrix computations are not attainable when data is stored in common graph data structures. In this paper, we present HAGGLE, a high-performance, scalable data analytics platform. The platform's hybrid data model supports a variety of distributed, thread-safe data structures, parallel programming constructs, and persistent and streaming data. An abstract runtime layer enables us to map the stack to conventional, distributed computer systems with accelerators. The runtime uses multithreading, active messages, and data aggregation to hide memory and synchronization latencies on large-scale systems. Vito Giovanni Castellana, Maurizio Drocco, John Feo, Jesun Sahariar Firoz, Thejaka Amila Kanewala, Andrew Lumsdaine, Joseph B. Manzano, Andrés Márquez 0001, Marco Minutoli, Joshua Suetterlein, Antonino Tumeo, Marcin Zalewski |
DATE | 10 |
| 2018 | Adaptive Runtime Features for Distributed Graph AlgorithmsabstractThe following topics are dealt with: parallel processing; learning (artificial intelligence); graphics processing units; graph theory; parallel algorithms; scheduling; application program interfaces; parallel architectures; storage management; parallel machines. Jesun Sahariar Firoz, Marcin Zalewski, Joshua Suetterlein, Andrew Lumsdaine |
HiPC | 3 |
| 2016 | Extending the Roofline Model for Asynchronous Many-Task RuntimesabstractA common practice for application developers is to experimentally determine the granularity of a task after a code has been parallelized based on the observed overhead of a runtime. Instead, we propose a new methodology based on an extended Roofline model to provide practical upper bounds on the throughput performance of an application. First, we extend the Roofline model to support not only latency hiding analysis, but also a multidimensional amortized analysis. By combining this new methodology with a serial application and an Asynchronous Many Task (AMT) runtime implementation, we can predict the worst case runtime overhead attribution of individual runtime features prior to the development of parallel code. Joshua Suetterlein, Joshua Landwehr, Andrés Márquez 0001, Joseph B. Manzano, Guang R. Gao |
CLUSTER | 1 |
| 2013 | An Implementation of the Codelet Model
Joshua Suetterlein, Stéphane Zuckerman, Guang R. Gao |
Euro-Par | 1 |