EDBT 2026 Demo / reviewers in the wild / expert
Noel Keen
dblp:20/6115
· DBLP profile ↗
8ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0003-3607-3554ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Energy-Efficient HPC: Insights from Power Profiling a Cloud-Resolving Earth System Model
Zhengji Zhao, Noel Keen, Oscar Antepara, Samuel Williams 0001, Luca Bertagna, Naser Mahfouz, James B. White III, Leonid Oliker, Brian Austin, Nicholas J. Wright |
IPDPS | 2 |
| 2023 | The Simple Cloud-Resolving E3SM Atmosphere Model Running on the Frontier Exascale System
Peter M. Caldwell, Luca Bertagna, Conrad Clevenger, Aaron Donahue, James G. Foucar, Oksana Guba, Benjamin R. Hillman, Noel Keen, Jayesh Krishna, Matthew R. Norman, Sarat Sreepathi, Christopher Terai, James B. White III, Andrew G. Salinger, Renata B. McCoy, L. Ruby Leung, David C. Bader, Danqing Wu |
SC | 9 |
| 2023 | Not all applications have boring communication patterns: Profiling message matching with BMMabstractSummary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach. Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Data Elevator: Low-Contention Data Movement in Hierarchical Storage SystemabstractHierarchical storage subsystems that include multiple layers of burst buffers (BB) and disk-based parallel file systems (PFS), are becoming an essential part of HPC systems to address the I/O performance gap. However, the state-of-the-art software for managing these hierarchical storage subsystems, such as Cray DataWarp, requires user involvement in moving data among storage layers. Such manual data movement may experience poor performance because of resource contention on the I/O servers of a layer for serving data movement in the hierarchy as well as regular read/write requests. In this paper, we propose a new system, named Data Elevator, for transparently and efficiently moving data in hierarchical storage. Users specify the final destination for their data, typically a PFS. Data Elevator intercepts the I/O calls, stages data on a fast persistent storage layer (for example, an SSD-based burst buffer), and then asynchronously transfers the data to the final destination in the background. Data Elevator reduces the resource contention on BB servers via offloading the data movement from a fixed number of BB server nodes to compute nodes. The number of the compute nodes is configurable based on the data movement load. Data Elevator also allows optimizations, such as overlapping read and write operations, choosing I/O modes, and aligning buffer boundaries. In our tests with large-scale scientific applications, Data Elevator is as much as 4.2X faster than Cray DataWarp, and 4X faster than directly writing data to PFS. Bin Dong 0002, Surendra Byna, Kesheng Wu, Prabhat, Hans Johansen, Jeffrey N. Johnson, Noel Keen |
HiPC | 7 |
| 2011 | Petascale Block-Structured AMR Applications without Distributed Meta-data
Brian van Straalen, Phillip Colella, Daniel T. Graves, Noel Keen |
Euro-Par (2) | 4 |
| 2010 | Parallel I/O performance: From events to ensemblesabstractParallel I/O is fast becoming a bottleneck to the research agendas of many users of extreme scale parallel computers. The principle cause of this is the concurrency explosion of high-end computation, coupled with the complexity of providing parallel file systems that perform reliably at such scales. More than just being a bottleneck, parallel I/O performance at scale is notoriously variable, being influenced by numerous factors inside and outside the application, thus making it extremely difficult to isolate cause and effect for performance events. In this paper, we propose a statistical approach to understanding I/O performance that moves from the analysis of performance events to the exploration of performance ensembles. Using this methodology, we examine two I/O-intensive scientific computations from cosmology and climate science, and demonstrate that our approach can identify application and middleware performance deficiencies - resulting in more than 4× run time improvement for both examined applications. Andrew Uselton, Mark Howison, Nicholas J. Wright, David Skinner, Noel Keen, John Shalf, Karen L. Karavanic, Leonid Oliker |
IPDPS | 5 |
| 2009 | Scalability challenges for massively parallel AMR applicationsabstractPDE solvers using Adaptive Mesh Refinement on block structured grids are some of the most challenging applications to adapt to massively parallel computing environments. We describe optimizations to the Chombo AMR framework that enable it to scale efficiently to thousands of processors on the Cray XT4. The optimization process also uncovered OS-related performance variations that were not explained by conventional OS interference benchmarks. Ultimately the variability was traced back to complex interactions between the application, system software, and the memory hierarchy. Once identified, software modifications to control the variability improved performance by 20% and decreased the variation in computation time across processors by a factor of 3. These newly identified sources of variation will impact many applications and suggest new benchmarks for OS-services be developed. Brian van Straalen, John Shalf, Terry J. Ligocki, Noel Keen, Woo-Sun Yang |
IPDPS | 4 |
| 2007 | An adaptive mesh refinement benchmark for modern parallel programming languagesabstractWe present an Adaptive Mesh Refinement benchmark for evaluating programmability and performance of modern parallel programming languages. Benchmarks employed today by language developing teams, originally designed for performance evaluation of computer architectures, do not fully capture the complexity of state-of-the-art computational software systems running on today’s parallel machines or to be run on the emerging ones from the multi-cores to the peta-scale High Productivity Computer Systems. This benchmark, extracted from a real application framework, presents challenges for a programming language in both expressiveness and performance. It consists of an infrastructure for finite difference calculations on block-structured adaptive meshes and a solver for elliptic Partial Differential Equations built on this infrastructure. Adaptive Mesh Refinement algorithms are challenging to implement due to the irregularity introduced by local mesh refinement. We describe those challenges posed by this benchmark through two reference implementations (C++/Fortran/MPI and Titanium) and in the context of three programming models. Categories and Subject Descriptors Tong Wen, Jimmy Su, Phillip Colella, Katherine A. Yelick, Noel Keen |
SC | 5 |