EDBT 2026 Demo / reviewers in the wild / expert
Joshua Dennis Booth
dblp:28/9637
· DBLP profile ↗
11ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0001-7269-6367ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural acceleration of incomplete factorization preconditioning
Joshua Dennis Booth, Hongyang Sun 0001, Trevor Garnett |
Neural Comput. Appl. | 1 |
| 2024 | To Protect or Not To Protect: Probability-Aware Selective Protection for Sparse Iterative SolversabstractWith the increasing scale of high-performance computing (HPC) systems, transient bit-flip errors are now more likely than ever, posing a threat to long-running scientific applications. A substantial portion of these applications involve simulation of partial differential equations (PDEs), modeling physical processes over discretized spatial and temporal domains, with some requiring solving sparse linear systems of equations. While these applications are often paired with system-level application-agnostic resilience techniques, such as checkpointing and replication, using these techniques imposes significant overhead. In this work, we present a probability-aware framework that produces low-overhead selective protection schemes for the widely used Preconditioned Conjugate Gradient (PCG) method, whose performance can heavily degrade due to error propagation through the sparse matrix-vector multiplication (SpMV) operation. Through the use of a straightforward mathematical model and an optimized machine learning model, our selective protection schemes incorporate error probability to protect only certain crucial operations. An experimental evaluation using 15 matrices from the SuiteSparse Matrix Collection demonstrates that our protection schemes effectively reduce resilience overheads, outperforming two baseline and two existing protection schemes across all error probabilities. Daniel Ryley Johnson, Hongyang Sun 0001, Joshua Dennis Booth, Padma Raghavan |
SBAC-PAD | 3 |
| 2024 | A NUMA-Aware Version of an Adaptive Self-Scheduling Loop SchedulerabstractParallelizing code in a shared-memory environment is commonly done utilizing loop scheduling (LS) in a fork-join manner as in OpenMP. This manner of parallelization is popular due to its ease to code, but the choice of the LS method is important when the workload per iteration is highly variable. Currently, the shared-memory environment is evolving in high-performance computing as larger chiplet-based processors with high core counts and segmented L3 cache are introduced. These processors have a stronger non-uniform memory access (NUMA) effect than the previous generation of x86-64 processors. This work attempts to modify the adaptive self-scheduling loop scheduler known as iCh ( i rregular Ch unk) for these NUMA environments while analyzing the impact of these systems on default OpenMP LS methods. In particular, iCh is as a default LS method for irregular applications (i.e., applications where the workload per iteration is highly variable) that guarantees “good” performance without tuning. The modified version, named NiCh , is demonstrated over multiple irregular applications to show the variation in performance. The work demonstrates that NiCh is able to better handle architectures with stronger NUMA effects, and particularly is better than iCh when the number of threads is greater than the number of cores. However, NiCh also comes with being less universally “good” than iCh and a set of parameters that are hardware dependent. Joshua Dennis Booth, Phillip Allen Lane |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | Heterogeneous sparse matrix-vector multiplication via compressed sparse row format
Phillip Allen Lane, Joshua Dennis Booth |
Parallel Comput. | 2 |
| 2022 | HTS: A Threaded Multilevel Sparse Hybrid Solver
Joshua Dennis Booth |
IPDPS | 1 |
| 2022 | An adaptive self-scheduling loop schedulerabstractAbstract Many shared‐memory parallel irregular applications, such as sparse linear algebra and graph algorithms, depend on efficient loop scheduling (LS) in a fork‐join manner despite that the work per loop iteration can greatly vary depending on the application and the input. Because of the importance of LS, many different methods (e.g., workload‐aware self‐scheduling) and parameters (e.g., chunk size) have been explored to achieve reasonable performance, and many of these methods require expert prior knowledge about the application and input before runtime. This work proposes a new LS method that requires little to no expert knowledge to achieve speedups close to those of tuned LS methods by self‐managing chunk size based on a heuristic of throughput and using work‐stealing to recover from workload imbalances. This method, named iCh, is implemented into libgomp for testing. It is evaluated against OpenMP's guided, dynamic, and taskloop methods and is evaluated against BinLPT and generic work‐stealing on an array of applications that includes: a synthetic benchmark, breadth‐first search, K‐Means, the molecular dynamics code LavaMD, and sparse matrix‐vector multiplication. On a 28 thread Intel system, iCh is the only method to always be one of the top three LS methods. On average across all applications, iCh is within of the best method and is even able to outperform other LS methods for breadth‐first search and K‐Means. Joshua Dennis Booth, Phillip Allen Lane |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | An on-node scalable sparse incomplete LU factorization for a many-core iterative solver with Javelin
Joshua Dennis Booth, Gregory Bolet |
Parallel Comput. | 1 |
| 2017 | Basker: Parallel sparse LU factorization utilizing hierarchical parallelism and data layouts
Joshua Dennis Booth, Nathan D. Ellingwood, Heidi Thornquist, Sivasankaran Rajamanickam |
Parallel Comput. | 1 |
| 2015 | Phase Detection with Hidden Markov Models for DVFS on Many-Core ProcessorsabstractThe energy concerns of many-core processors are increasing with the number of cores. We provide a new method that reduces energy consumption of an application on many-core processors by identifying unique segments to apply dynamic voltage and frequency scaling (DVFS). Our method, phase-based voltage and frequency scaling (PVFS), hinges on the identification of phases, i.e., Segments of code with unique performance and power attributes, using hidden Markov Models. In particular, we demonstrate the use of this method to target hardware components on many-core processors such as Network-on-Chip (NoC). PVFS uses these phases to construct a static power schedule that uses DVFS to reduce energy with minimal performance penalty. This general scheme can be used with a variety of performance and power metrics to match the needs of the system and application. More importantly, the flexibility in the general scheme allows for targeting of the unique hardware components of future many-core processors. We provide an in-depth analysis of PVFS applied to five threaded benchmark applications, and demonstrate the advantage of using PVFS for 4 to 32 cores in a single socket. Empirical results of PVFS show a reduction of up to 10.1% of total energy while only impacting total time by at most 2.7% across all core counts. Furthermore, PVFS outperforms standard coarse-grain time-driven DVFS, while scaling better in terms of energy savings with increasing core counts. Joshua Dennis Booth, Jagadish Kotra, Hui Zhao 0013, Mahmut T. Kandemir, Padma Raghavan |
ICDCS | 1 |
| 2015 | STS-k: a multilevel sparse triangular solution scheme for NUMA multicoresabstractWe consider techniques to improve the performance of parallel sparse triangular solution on non-uniform memory architecture multicores by extending earlier coloring and level set schemes for single-core multiprocessors. We develop STS-k, where k represents a small number of transformations for latency reduction from increased spatial and temporal locality of data accesses. We propose a graph model of data reuse to inform the development of STS-k and to prove that computing an optimal cost schedule is NP-complete. We observe significant speed-ups with STS-3 on 32-core Intel Westmere-Ex and 24-core AMD `MagnyCours' processors. Incremental gains solely from the 3-level transformations in STS-3 for a fixed ordering, correspond to reductions in execution times by factors of 1.4(Intel) and 1.5(AMD) for level sets and 2(Intel) and 2.2(AMD) for coloring. On average, execution times are reduced by a factor of 6(Intel) and 4(AMD) for STS-3 with coloring compared to a reference implementation using level sets. Humayun Kabir, Joshua Dennis Booth, Guillaume Pallez, Anne Benoit, Yves Robert, Padma Raghavan |
SC | 2 |
| 2014 | A multilevel compressed sparse row format for efficient sparse computations on multicore processorsabstractWe seek to improve the performance of sparse matrix computations on multicore processors with non-uniform memory access (NUMA). Typical implementations use a bandwidth reducing ordering of the matrix to increase locality of accesses with a compressed storage format to store and operate only on the non-zero values. We propose a new multilevel storage format and a companion ordering scheme as an explicit adaptation to map to NUMA hierarchies. More specifically, we propose CSR-k, a multilevel form of the popular compressed sparse row (CSR) format for a multicore processor with k > 1 well-differentiated levels in the memory subsystem. Additionally, we develop Band-k, a modified form of a traditional bandwidth reduction scheme, to convert a matrix represented in CSRto our proposed CSR-k. We evaluate the performance of the widely-used and important sparse matrix-vector multiplication (SpMV) kernel using CSR-2 on Intel Westmere processors for a test suite of 12 large sparse matrices with row densities in the range 3 to 45. On 32 cores, on average across all matrices in the test suite, the execution time for SpMV with CSR-2is less than 42% of the time taken by the state-of-the-art automatically tuned SpMV resulting in energy savings of approximately 56%. Additionally, on average, the parallel speed-up on 32 cores of the automatically tuned SpMV relative to its 1-core performance is 8.18 compared to a value of 19.71 for CSR-2. Our analysis indicates that the higher performance of SpMV with CSR-2 comes from achieving higher reuse of x in the shared L3 cache without incurring overheads from fill-in of original zeroes. Furthermore, the pre-processing costs of SpMV with CSR-2 can be amortized on average over 97 iterations of SpMV using CSR and are substantially lower than the 513 iterations required for the automatically tuned implementation. Based on these results, CSR-k appears to be a promising multilevel formulation of CSR for adapting sparse computations to multicore processors with NUMA memory hierarchies. Humayun Kabir, Joshua Dennis Booth, Padma Raghavan |
HiPC | 2 |