EDBT 2026 Demo / reviewers in the wild / expert
Brandon Cook 0001
dblp:186/2732 · also Brandon G. Cook
· DBLP profile ↗
15ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-4203-4079ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Job Scheduling in High Performance Computing Systems with Disaggregated Memory ResourcesabstractDisaggregated memory promises to meet growing memory requirements of applications while improving system resource utilization in high-performance computing (HPC) systems. Compared to traditional systems-where expensive resources such as CPUs, GPUs, and memory, are assigned to jobs in units of nodes-systems with disaggregated memory introduce memory pools that can be shared among jobs; this introduces new optimization metrics to the job scheduler. In this paper, we propose a data-driven approach to evaluate job scheduling and resource configuration in HPC systems with disaggregated memory. To incorporate the memory requirements of jobs for both local and disaggregated memory resources and improve system efficiency in open-science HPC systems, we introduce a novel job scheduling algorithm called FM (Fair Memory). Our simulation results show that FM outperforms commonly-used job schedulers in terms of jobs' bounded slowdown when the shared memory pool capacity is limited, and in terms of fairness under all conditions. Jie Li 0057, George Michelogiannakis, Samuel A. Maloney, Brandon Cook 0001, Estela Suarez, John Shalf, Yong Chen 0001 |
CLUSTER | 4 |
| 2023 | Efficient Intra-Rack Resource Disaggregation for HPC Using Co-Packaged DWDM PhotonicsabstractThe diversity of workload requirements and increasing hardware heterogeneity in emerging high performance computing (HPC) systems motivate resource disaggregation. Resource disaggregation allows compute and memory resources to be allocated individually as required to each workload. However, it is unclear how to efficiently realize this capability and cost-effectively meet the stringent bandwidth and latency requirements of HPC applications. To that end, we describe how modern photonics can be co-designed with modern HPC racks to implement flexible intra-rack resource disaggregation and fully meet the bit error rate (BER) and high escape bandwidth of all chip types in modern HPC racks. Our photonic-based disaggregated rack provides an average application speedup of 11% (46% maximum) for 25 CPU and 61% for 24 GPU benchmarks compared to a similar system that instead uses modern electronic switches for disaggregation. Using observed resource usage from a production system, we estimate that an iso-performance intra-rack disaggregated HPC system using photonics would require 4× fewer memory modules and 2× fewer NICs than a non-disaggregated baseline. George Michelogiannakis, Yehia Arafa, Brandon Cook 0001, Liang Yuan Dai, Abdel-Hameed A. Badawy, Madeleine Glick, Yuyang Wang 0003, Keren Bergman, John Shalf |
CLUSTER | 3 |
| 2023 | Fault-Tolerant LOBPCG for Nuclear CI CalculationsabstractExascale computing platforms with millions of compute units and with thousands of nodes are predicted to experience frequent faults which interrupt applications’ execution. In this context resilience against faults becomes important. We examine user and software level fault mitigation strategies in a distributed LOBPCG algorithm targeting nuclear CI calculations. In particular, we present and evaluate one strategy that keeps the total number of fault-tolerant LOBPCG iterations close to that of the standard LOBPCG algorithm ran on a fault-free machine. Meiyue Shao, Dossay Oryspayev, Chao Yang 0001, Pieter Maris, Brandon Cook 0001 |
HPC Asia | 5 |
| 2023 | Not all applications have boring communication patterns: Profiling message matching with BMMabstractSummary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach. Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | A Case For Intra-rack Resource Disaggregation in HPCabstractThe expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization. George Michelogiannakis, Benjamin Klenk, Brandon Cook 0001, Min Yee Teh, Madeleine Glick, Larry Dennison, Keren Bergman, John Shalf |
ACM Trans. Archit. Code Optim. | 3 |
| 2020 | Quantifying the impact of network congestion on application performance and network metricsabstractIn modern high-performance computing (HPC) systems, network congestion is an important factor that contributes to performance degradation. However, how network congestion impacts application performance is not fully understood. As Aries network, a recent HPC network architecture featuring a dragonfly topology, is equipped with network counters measuring packet transmission statistics on each router, these network metrics can potentially be utilized to understand network performance. In this work, by experiments on a large HPC system, we quantify the impact of network congestion on various applications' performance in terms of execution time, and we correlate application performance with network metrics. Our results demonstrate diverse impacts of network congestion: while applications with intensive MPI operations (such as HACC and MILC) suffer from more than 40% extension in their execution times under network congestion, applications with less intensive MPI operations (such as Graph500 and HPCG) are mostly not affected. We also demonstrate that a stall-to-flit ratio metric derived from Aries network counters is positively correlated with performance degradation and, thus, this metric can serve as an indicator of network congestion in HPC systems. Yijia Zhang 0002, Taylor L. Groves, Brandon Cook 0001, Nicholas J. Wright, Ayse K. Coskun |
CLUSTER | 3 |
| 2020 | Scaling of Union of Intersections for Inference of Granger Causal Networks from Observational DataabstractThe development of advanced recording and measurement devices in scientific fields is producing high-dimensional time series data. Vector autoregressive (VAR) models are well suited for inferring Granger-causal networks from high dimensional time series data sets, but accurate inference at scale remains a central challenge. We have recently introduced a flexible and scalable statistical machine learning framework, Union of Intersections (UoI), which enables low false-positive and low false-negative feature selection along with low bias and low variance estimation, enhancing interpretation and predictive accuracy. In this paper, we scale the UoI framework for VAR models (algorithm UoIV AR) to infer network connectivity from large time series data sets (TBs). To achieve this, we optimize distributed convex optimization and introduce novel strategies for improved data read and data distribution times. We study the strong and weak scaling of the algorithm on a Xeon-phi based supercomputer (100,000 cores). These advances enable us to estimate the largest VAR model as known (1000 nodes, corresponding to 1M parameters) and apply it to large time series data from neurophysiology (192 neurons) and finance (470 companies). Mahesh Balasubramanian 0001, Trevor D. Ruiz, Brandon Cook 0001, Prabhat, Sharmodeep Bhattacharyya, Aviral Shrivastava, Kristofer E. Bouchard |
IPDPS | 3 |
| 2020 | The Case of Performance Variability on Dragonfly-based SystemsabstractPerformance of a parallel code running on a large supercomputer can vary significantly from one run to another even when the executable and its input parameters are left unchanged. Such variability can occur due to perturbation of the computation and/or communication in the code. In this paper, we investigate the case of performance variability arising due to network effects on supercomputers that use a dragonfly topology - specifically, Cray XC systems equipped with the Aries interconnect. We perform post-mortem analysis of network hardware counters, profiling output, job queue logs, and placement information, all gathered from periodic representative application runs. We investigate the causes of performance variability using deviation prediction and recursive feature elimination. Additionally, using time-stepped performance data of individual applications, we train machine learning models that can forecast the execution time of future time steps. Abhinav Bhatele, Jayaraman J. Thiagarajan, Taylor L. Groves, Rushil Anirudh, Staci A. Smith, Brandon Cook 0001, David K. Lowenthal |
IPDPS | 6 |
| 2020 | Tuning floating-point precision using dynamic program information and temporal localityabstractWe present a methodology for precision tuning of full applications. These techniques must select a search space composed of either variables or instructions and provide a scalable search strategy. In full application settings one cannot assume compiler support for practical reasons. Thus, an additional important challenge is enabling code refactoring. We argue for an instruction-based search space and we show: 1) how to exploit dynamic program information based on call stacks; and 2) how to exploit the iterative nature of scientific codes, combined with temporal locality. We applied the methodology to tune the implementation of scientific codes written in a combination of Python, CUDA, C++ and Fortran, tuning calls to math exp library functions. The iterative search refinement always reduces the search complexity and the number of steps to solution. Dynamic program information increases search efficacy. Using this approach, we obtain application runtime performance improvements up to 27%. Hugo Brunie, Costin Iancu, Khaled Z. Ibrahim, Philip Brisk, Brandon Cook 0001 |
SC | 5 |
| 2019 | Eigensolver performance comparison on Cray XC systemsabstractSummary Hermitian (symmetric) eigenvalue solvers are the core constituents of electronic structure, quantum‐chemistry, and other HPC applications such as Quantum ESPRESSO, VASP, CP2K, and NWChem to name a few. Our understanding of the performance of symmetric eigenvalue algorithms on various hardware is clearly important to the quantum chemistry or condensed matter physics community but in fact goes beyond that community. For instance, big data analytics is increasingly utilizing eigenvalues solvers, in the study of randomized singular value decomposition (SVD) or principal component analysis (PCA). Noise, vibration, and harshness (NVH) is another field where fast and efficient eigenvalue solvers are required. Most eigenvalue solver packages feature numerous different parameters which can be tuned for performance, eg, the number of nodes, number of total ranks, the decomposition of the matrix, etc. In this paper, we investigate the performance of different packages as well as the influence of these knobs on the solver performance. Brandon Cook 0001, Thorsten Kurth, Jack Deslippe, Pierre Luc Carrier, Nick Hill, Nathan Wichmann |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Evaluating the networking characteristics of the Cray XC-40 Intel Knights Landing-based Cori supercomputer at NERSCabstractSummary There are many potential issues associated with deploying the Intel Xeon PhiTM (code named Knights Landing [KNL]) manycore processor in a large‐scale supercomputer. One in particular is the ability to fully utilize the high‐speed communications network, given that the serial performance of a Xeon PhiTM core is a fraction of a Xeon®core. In this paper, we take a look at the trade‐offs associated with allocating enough cores to fully utilize the Aries high‐speed network versus cores dedicated to computation, eg, the trade‐off between MPI and OpenMP. In addition, we evaluate new features of Cray MPI in support of KNL, such as internode optimizations. We also evaluate one‐sided programming models such as Unified Parallel C. We quantify the impact of the above trade‐offs and features using a suite of National Energy Research Scientific Computing Center applications. Douglas Doerfler, Brian Austin, Brandon Cook 0001, Jack Deslippe, Krishna Kandalla, Peter Mendygral |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | Preparing NERSC users for Cori, a Cray XC40 system with Intel many integrated coresabstractSummary The newest NERSC supercomputer Cori is a Cray XC40 system consisting of 2,388 Intel Xeon Haswell nodes and 9,688 Intel Xeon‐Phi “Knights Landing” (KNL) nodes. Compared to the Xeon‐based clusters NERSC users are familiar with, optimal performance on Cori requires consideration of KNL mode settings; process, thread, and memory affinity; fine‐grain parallelization; vectorization; and use of the high‐bandwidth MCDRAM memory. This paper describes our efforts preparing NERSC users for KNL through the NERSC Exascale Science Application Program, Web documentation, and user training. We discuss how we configured the Cori system for usability and productivity, addressing programming concerns, batch system configurations, and default KNL cluster and memory modes. System usage data, job completion analysis, programming and running jobs issues, and a few successful user stories on KNL are presented. Yun (Helen) He, Brandon Cook 0001, Jack Deslippe, Brian Friesen, Richard A. Gerber, Rebecca Hartman-Baker, Alice E. Koniges, Thorsten Kurth, Stephen Leak, Woo-Sun Yang, Zhengji Zhao, Eddie Baron, Peter Hauschildt |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Performance Characterization of De Novo Genome Assembly on Leading Parallel Systems
Marquita Ellis, Evangelos Georganas, Rob Egan, Steven Hofmeyr, Aydin Buluç, Brandon Cook 0001, Leonid Oliker, Katherine A. Yelick |
Euro-Par | 6 |
| 2017 | Multi-source sensor fusion for small unmanned aircraft systems using fuzzy logicabstractAs the applications for using small Unmanned Aircraft Systems (sUAS) beyond visual line of sight (BVLOS) continue to grow in the coming years, it is imperative that intelligent sensor fusion techniques be explored. In BVLOS scenarios the vehicle position must accurately be tracked over time to ensure no two vehicles collide with one another, no vehicle crashes into surrounding structures, and to identify off-nominal scenarios. In this study, an intelligent systems approach is used to estimate the position of sUAS given a variety of sensor platforms, including GPS, radar, and onboard detection hardware. Common research challenges include multiple sensor platforms and sensor reliability. In an effort to resolve these challenges, techniques such as a Maximum a Posteriori estimation and Fuzzy Logic based sensor confidence determination are used. Brandon Cook 0001, Kelly Cohen |
FUZZ-IEEE | 1 |
| 2016 | MPI usage at NERSC: Present and FutureabstractIn this poster, we describe how MPI is used at the National Energy Research Scientific Computing Center (NERSC) NERSC is the production high-performance computing center for the US Department of Energy, with more than 5000 users and 800 distinct projects. Through a variety of tools (e.g., User Survey, application team collaborations, etc.), we determine how MPI is used on our latest systems, with a particular focus on advanced features and how early applications intend to use MPI on NERSC's upcoming Intel Knights Landing (KNL) many-core system1 - one of the first to be deployed. In the poster, we also compare the usage of MPI to exascale developmental programming models such as UPC++ and HPX, with an eye on what features and extensions to MPI are plausible and useful for NERSC users. We also discuss perceived shortcomings of MPI, and why certain groups use other parallel programming models on the systems. In addition to a broad survey of the NERSC HPC population, we follow the evolution of a few key application codes2 that are being highly optimized for the KNL architecture using advanced OpenMP techniques. We study how these highly optimized on-node proxy apps and full applications start to make the transition to using full hybrid MPI+OpenMP implementations on the self-hosted KNL system. Alice E. Koniges, Brandon Cook 0001, Jack Deslippe, Thorsten Kurth, Hongzhang Shan |
EuroMPI | 2 |