Edgar A. León

dblp:44/6007 · DBLP profile ↗
← Back
16ranked-venue papers
11as first author
4since 2021 · last 2026
0000-0001-5805-9046ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 9 first-author · 4 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Compute Overlap Stall (COS): Predicting Performance of Power Management for Shared Memory Codes When Throttling Processors, Memory, and Thread Concurrency
Alexandra K. McCoy, Bo Li 0032, Gregory Bolet, Hayden Estes, Edgar A. León, Kirk W. Cameron
IEEE Trans. Parallel Distributed Syst.5
2025 Breaking the System Noise Barrier at Exascale
abstract
To meet the increasing demands of parallel scientific applications, supercomputers continue to grow in both scale and complexity. The fastest supercomputer in the world, El Capitan, features over a million CPU cores and tens of thousands of GPUs. Applications running on such large-scale systems are particularly susceptible to system noise or interference caused by the operating system (OS) and other services running on the same compute nodes as the application.
Edgar A. León, Joseph Glenski, Mark J. Stock, Kim H. McMahon, William Loewe, Clark Snyder, Larry Kaplan, Srinath Vadlamani, Timothy I. Mattox, Trent D'Hooge, Brian Behlendorf, Nathan Hanford, Ramesh Pankajakshan, Matthew L. Leininger
SC1
2024 Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System
abstract
Molecular dynamics (MD) simulations have transformed our understanding of the nanoscale, driving breakthroughs in materials science, computational chemistry, and several other fields, including biophysics and drug design. Even on exascale supercomputers, however, runtimes are excessive for systems and timescales of scientific interest. Here, we demonstrate strong scaling of MD simulations on the Cerebras Wafer-Scale Engine. By dedicating a processor core for each simulated atom, we demonstrate a 457-fold improvement in timesteps per second versus the Frontier GPU-based Exascale platform, along with a large improvement in timesteps per unit energy. Reducing every year of runtime to less than a day unlocks currently inaccessible timescales of slow microstructure transformation processes that are critical for understanding material behavior and function.Our dataflow algorithm runs Embedded Atom Method (EAM) simulations at rates over 699k timesteps per second for problems with up to 800k atoms. This demonstrated performance is unprecedented for general-purpose processing cores.
Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian, Ilya Sharapov, Aidan P. Thompson, Delyan Kalchev, Danny Perez, Robert Schreiber, Scott Pakin, Edgar A. León, James H. Laros III, Michael James 0002, Sivasankaran Rajamanickam
SC11
2021 On-the-Fly, Robust Translation of MPI Libraries
abstract
Most parallel scientific applications rely on third-party libraries, some of which may have multiple implementations including open-source and vendor-proprietary. While sharing an application programming interface (API), many of these implementations do not have a shared application binary interface (ABI) and require recompiling applications to change the library implementation used. For many applications, recompiling is a long and complex process and sometimes not even an option when the application is shipped binary only. ABI incompatibility strikes at the heart of portability, productivity, and performance by (1) impeding application execution across different systems; (2) adding developer hours rebuilding an application; and (3) not taking advantage of host-optimized libraries.In this paper, we present a methodology and framework to solve ABI incompatibility across MPI libraries, which follow a well-defined API. The proposed framework called Wi4MPI translates the ABI dynamically from the MPI library used to build the application to a different MPI library available at run time. We show Wi4MPI works robustly on a wide spectrum of architectures, networks, and MPI libraries. Furthermore, we demonstrate its usefulness on several use cases highlighting significant portability, performance, and productivity benefits.
Edgar A. León, Marc Joos, Nathan Hanford, Adrien Cotte, Tony Delforge, François Diakhaté, Vincent Ducrot, Ian Karlin, Marc Pérache
CLUSTER1
2020 TOSS-2020: a commodity software stack for HPC
abstract
The simulation environment of any HPC platform is key to the performance, portability, and productivity of scientific applications. This environment has traditionally been provided by platform vendors, presenting challenges for HPC centers and users including platform-specific software that tend to stagnate over the lifetime of the system. In this paper, we present the Tri-Laboratory Operating System Stack (TOSS), a production simulation environment based on Linux and open source software, with proprietary software components integrated as needed. TOSS, focused on mid-to-large scale commodity HPC systems, provides a common simulation environment across system architectures, reduces the learning curve on new systems, and benefits from a lineage of past experience and bug fixes. To further the scope and applicability of TOSS, we demonstrate its feasibility and effectiveness on a leadership-class supercomputer architecture. Our evaluation, relative to the vendor stack, includes an analysis of resource manager complexity, system noise, networking, and application performance.
Edgar A. León, Trent D'Hooge, Nathan Hanford, Ian Karlin, Ramesh Pankajakshan, Jim Foraker, Christopher M. Chambreau, Matthew L. Leininger
SC1
2017 COS: A Parallel Performance Model for Dynamic Variations in Processor Speed, Memory Speed, and Thread Concurrency
abstract
Highly-parallel, high-performance scientific applications must maximize performance inside of a power envelope while maintaining scalability. Emergent parallel and distributed systems offer a growing number of operating modes that provide unprecedented control of processor speed, memory latency, and memory bandwidth. Optimizing these systems for performance and power requires an understanding of the combined effects of these modes and thread concurrency on execution time. In this paper, we describe how an analytical performance model that separates pure computation time (C) and pure stall time (S) from computation-memory overlap time (O) can accurately capture these combined effects. We apply the COS model to predict the performance of thread and power mode combinations to within 7% and 17% for parallel applications (e.g. LULESH) on Intel x86 and IBM BG/Q architectures, respectively. The key insight of the COS model is that the combined effects of processor and memory throttling and concurrency on overlap trend differently than the combined effects on pure computation and pure stall time. The COS model is novel in that it enables independent approximation of overlap which leads to capabilities and accuracies that are as good or better than the best available approaches.
Bo Li 0032, Edgar A. León, Kirk W. Cameron
HPDC2
2017 Predicting the performance impact of different fat-tree configurations
abstract
The fat-tree topology is one of the most commonly used network topologies in HPC systems. Vendors support several options that can be configured when deploying fat-tree networks on production systems, such as link bandwidth, number of rails, number of planes, and tapering. This paper showcases the use of simulations to compare the impact of these design options on representative production HPC applications, libraries, and multi-job workloads. We present advances in the TraceR-CODES simulation framework that enable this analysis and evaluate its prediction accuracy against experiments on a production fat-tree network. In order to understand the impact of different network configurations on various anticipated scenarios, we study workloads with different communication patterns, computation-to-communication ratios, and scaling characteristics. Using multi-job workloads, we also study the impact of inter-job interference on performance and compare the cost-performance tradeoffs.
Abhinav Bhatele, Louis H. Howell, David Böhme, Ian Karlin, Edgar A. León, Misbah Mubarak, Noah Wolfe, Todd Gamblin, Matthew L. Leininger
SC6
2016 System Noise Revisited: Enabling Application Scalability and Reproducibility with SMT
abstract
Despite significant advances in reducing system noise, the scalability and performance of scientific applications running on production commodity clusters today continue to suffer from the effects of noise. Unlike custom and expensive leadership systems, the Linux ecosystem provides a rich set of services that application developers utilize to increase productivity and to ease porting. The cost is the overhead that these services impose on a running application, negatively impacting its scalability and performance reproducibility. In this work, we propose and evaluate a simple yet effective way to isolate an application from system processes by leveraging Simultaneous Multi-Threading (SMT), a pervasive architectural feature on current systems. Our method requires no changes to the operating system or to the application. We quantify its effectiveness on a diverse set of scientific applications of interest to the U. S. Department of Energy showing performance improvements of up to 2.4 times at 16,384 tasks for a high-order finite elements shock hydrodynamics application. Finally, we provide guidance to system and application developers on how to best leverage SMT under different application characteristics and scales.
Edgar A. León, Ian Karlin, Adam Moody
IPDPS1
2016 Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree
abstract
Understanding the characteristics and requirements of applications that run on commodity clusters is key to properly configuring current machines and, more importantly, procuring future systems effectively. There are only a few studies, however, that are current and characterize realistic workloads. For HPC practitioners and researchers, this limits our ability to design solutions that will have an impact on real systems. We present a systematic study that characterizes applications with an emphasis on communication requirements. It includes cluster utilization data, identifying a representative set of applications from a U.S. Department of Energy laboratory, and characterizing their communication requirements. The driver for this work is understanding application sensitivity to a tapered fat-tree network. These results provided key insights into the procurement of our next generation commodity systems. We believe this investigation can provide valuable input to the HPC community in terms of workload characterization and requirements from a large supercomputing center.
Edgar A. León, Ian Karlin, Abhinav Bhatele, Steve H. Langer, Christopher M. Chambreau, Louis H. Howell, Trent D'Hooge, Matthew L. Leininger
SC1
2016 Program optimizations: The interplay between power, performance, and energy
Edgar A. León, Ian Karlin, Ryan E. Grant, Matthew G. F. Dosanjh
Parallel Comput.1
2015 Optimizing Explicit Hydrodynamics for Power, Energy, and Performance
abstract
Practical considerations for future supercomputer designs will impose limits on both instantaneous power consumption and total energy consumption. Working within these constraints while providing the maximum possible performance, application developers will need to optimize their code for speed alongside power and energy concerns. This paper analyzes the effectiveness of several code optimizations including loop fusion, data structure transformations, and global allocations. A per component measurement and analysis of different architectures is performed, enabling the examination of code optimizations on different compute subsystems. Using an explicit hydrodynamics proxy application from the U.S. Department of Energy, LULESH, we show how code optimizations impact different computational phases of the simulation. This provides insight for simulation developers into the best optimizations to use during particular simulation compute phases when optimizing code for future supercomputing platforms. We examine and contrast both x86 and Blue Gene architectures with respect to these optimizations.
Edgar A. León, Ian Karlin, Ryan E. Grant
CLUSTER1
2015 A Container-Based Approach to OS Specialization for Exascale Computing
abstract
Future exascale systems will impose several conflicting challenges on the operating system (OS) running on the compute nodes of such machines. On the one hand, the targeted extreme scale requires the kind of high resource usage efficiency that is best provided by lightweight OSes. At the same time, substantial changes in hardware are expected for exascale systems. Compute nodes are expected to host a mix of general-purpose and special-purpose processors or accelerators tailored for serial, parallel, compute-intensive, or I/O-intensive workloads. Similarly, the deeper and more complex memory hierarchy will expose multiple coherence domains and NUMA nodes in addition to incorporating nonvolatile RAM. That expected workload and hardware heterogeneity and complexity is not compatible with the simplicity that characterizes high performance lightweight kernels. In this work, we describe the Argo Exascale node OS, which is our approach to providing in a single kernel the required OS environments for the two aforementioned conflicting goals. We resort to multiple OS specializations on top of a single Linux kernel coupled with multiple containers.
Judicael A. Zounmevo, Swann Perarnau, Kamil Iskra, Kazutomo Yoshii, Roberto Gioiosa, Brian Van Essen, Maya B. Gokhale, Edgar A. León
IC2E8
2011 Cache injection for parallel applications
abstract
For two decades, the memory wall has affected many applications in their ability to benefit from improvements in processor speed. Cache injection addresses this disparity for I/O by writing data into a processor's cache directly from the I/O bus. This technique reduces data latency and, unlike data prefetching, improves memory bandwidth utilization. These improvements are significant for data-intensive applications whose performance is dominated by compulsory cache misses.
Edgar A. León, Rolf Riesen, Kurt B. Ferreira, Arthur B. Maccabe
HPDC1
2009 Instruction-level simulation of a cluster at scale
abstract
Instruction-level simulation is necessary to evaluate new architectures. However, single-node simulation cannot predict the behavior of a parallel application on a supercomputer. We present a scalable simulator that couples a cycle-accurate node simulator with a supercomputer network model. Our simulator executes individual instances of IBM's Mambo PowerPC simulator on hundreds of cores. We integrated a NIC emulator into Mambo and model the network instead of fully simulating it. This decouples the individual node simulators and makes our design scalable.
Edgar A. León, Rolf Riesen, Arthur B. Maccabe, Patrick G. Bridges
SC1
2005 An infrastructure for the development of kernel network services proof of concept: fast UDP
abstract
Scientific applications demand tremendous amounts of computational capabilities. In the last few years, the computational power provided by clusters of workstations and SMPs has become popular as a cost-effective alternative to supercomputers. Nodes in these systems are interconnected using high-speed networks via "smart" Network Interface Controllers (NIC). These controllers allow the overlap of computation and communication by processing communication tasks on the NIC and computational tasks on the host processor(s).
Edgar A. León, Michal Ostrowski
SOSP1
2002 Instrumenting LogP Parameters in GM: Implementation and Validation
abstract
This paper describes an apparatus which can be used to vary communication performance parameters for MPI applications, and provides a tool to analyze the impact of communication performance on parallel applications. Our apparatus is based on Myrinet (along with GM). We use an extension of the LogP model to allow higher flexibility in determining the parameter(s) to which parallel applications may be more sensitive. We show that individual communication parameters can be controlled within a small percentage error, and that the other parameters remain unchanged.
Edgar A. León, Arthur B. Maccabe, Ron Brightwell
LCN1