Eric J. Bohm

dblp:50/6079 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Interconnection networks and networks-on-chip · 35% High-performance computing · 34% Parallel and multicore computing · 24%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip
network topology
0.322014
Mapping to Irregular Torus Topologies and Other Techniques for Petascale Biomolecular Simulation · SC 2014
Topology aware task mapping techniques: an api and case study · PPoPP 2009
High-performance computing
performance optimization at scale
0.212014
Mapping to Irregular Torus Topologies and Other Techniques for Petascale Biomolecular Simulation · SC 2014
Interconnection networks and networks-on-chip › network topology
topology-aware mapping
0.212014
Mapping to Irregular Torus Topologies and Other Techniques for Petascale Biomolecular Simulation · SC 2014
Parallel and multicore computing
parallel programming runtimes
0.122011
Enabling and scaling biomolecular simulations of 100 million atoms on petascale machines with a multicore-optimized message-driven runtime · SC 2011
Poster reception - Charm++ simplifies coding for the cell processor · SC 2006
High-performance computing › scientific computing systems
molecular dynamics simulation
0.112012
Optimizing fine-grained communication in a biomolecular simulation application on Cray XK6 · SC 2012
High-performance computing › scientific computing systems
biomolecular simulation
0.112011
Enabling and scaling biomolecular simulations of 100 million atoms on petascale machines with a multicore-optimized message-driven runtime · SC 2011
Parallel and multicore computing › parallel programming runtimes
message-driven runtime
0.112011
Enabling and scaling biomolecular simulations of 100 million atoms on petascale machines with a multicore-optimized message-driven runtime · SC 2011
Parallel and multicore computing
task allocation
0.112009
Topology aware task mapping techniques: an api and case study · PPoPP 2009
Hardware accelerators and domain-specific architectures
accelerator programming models
0.112006
Poster reception - Charm++ simplifies coding for the cell processor · SC 2006
Processor architecture and microarchitecture › chip multiprocessor
cell processor
0.112006
Poster reception - Charm++ simplifies coding for the cell processor · SC 2006
Computational science and engineering › computational chemistry › molecular simulation
molecular dynamics
0.112014
Mapping to Irregular Torus Topologies and Other Techniques for Petascale Biomolecular Simulation · SC 2014
High-performance computing
parallel i/o
0.012011
Enabling and scaling biomolecular simulations of 100 million atoms on petascale machines with a multicore-optimized message-driven runtime · SC 2011
Parallel and multicore computing
parallel programming models
0.012006
Poster reception - Charm++ simplifies coding for the cell processor · SC 2006

Methods — techniques the papers use, named apart from their topics

topology adaptation · 0.4spatial decomposition · 0.4runtime optimization · 0.1performance analysis · 0.1node-aware optimization · 0.1hierarchical load balancing · 0.1charm++ runtime · 0.1wormhole routing · 0.1runtime system · 0.1offload API · 0.1
YearPublicationVenuePosition
2014 Overcoming the Scalability Challenges of Epidemic Simulations on Blue Waters
abstract
Modeling dynamical systems represents an important application class covering a wide range of disciplines including but not limited to biology, chemistry, finance, national security, and health care. Such applications typically involve large-scale, irregular graph processing, which makes them difficult to scale due to the evolutionary nature of their workload, irregular communication and load imbalance. EpiSimdemics is such an application simulating epidemic diffusion in extremely large and realistic social contact networks. It implements a graph-based system that captures dynamics among co-evolving entities. This paper presents an implementation of EpiSimdemics in Charm++ that enables future research by social, biological and computational scientists at unprecedented data and system scales. We present new methods for application-specific processing of graph data and demonstrate the effectiveness of these methods on a Cray XE6, specifically NCSA's Blue Waters system.
Jae-Seung Yeom, Abhinav Bhatele, Keith R. Bisset, Eric J. Bohm, Abhishek Gupta 0002, Laxmikant V. Kalé, Madhav V. Marathe, Dimitrios S. Nikolopoulos, Martin Schulz 0001, Lukasz Wesolowski
IPDPS4
2014 Mapping to Irregular Torus Topologies and Other Techniques for Petascale Biomolecular Simulation
abstract
Currently deployed petascale supercomputers typically use toroidal network topologies in three or more dimensions. While these networks perform well for topology-agnostic codes on a few thousand nodes, leadership machines with 20,000 nodes require topology awareness to avoid network contention for communication-intensive codes. Topology adaptation is complicated by irregular node allocation shapes and holes due to dedicated input/output nodes or hardware failure. In the context of the popular molecular dynamics program NAMD, we present methods for mapping a periodic 3-D grid of fixed-size spatial decomposition domains to 3-D Cray Gemini and 5-D IBM Blue Gene/Q toroidal networks to enable hundred-million atom full machine simulations, and to similarly partition node allocations into compact domains for smaller simulations using multiple-copy algorithms. Additional enabling techniques are discussed and performance is reported for NCSA Blue Waters, ORNL Titan, ANL Mira, TACC Stampede, and NERSC Edison.
James C. Phillips, Yanhua Sun, Eric J. Bohm, Laxmikant V. Kalé
SC4
2012 Optimizing fine-grained communication in a biomolecular simulation application on Cray XK6
abstract
Achieving good scaling for fine-grained communication intensive applications on modern supercomputers remains challenging. In our previous work, we have shown that such an application -- NAMD -- scales well on the full Jaguar XT5 without long-range interactions; Yet, with them, the speedup falters beyond 64K cores. Although the new Gemini interconnect on Cray XK6 has improved network performance, the challenges remain, and are likely to remain for other such networks as well. We analyze communication bottlenecks in NAMD and its CHARM++ runtime, using the Projections performance analysis tool. Based on the analysis, we optimize the runtime, built on the uGNI library for Gemini. We present several techniques to improve the fine-grained communication. Consequently, the performance of running 92224-atom Apoa1 with GPUs on TitanDev is improved by 36%. For 100-million-atom STMV, we improve upon the prior Jaguar XT5 result of 26 ms/step to 13 ms/step using 298,992 cores on Jaguar XK6.
Yanhua Sun, Gengbin Zheng, Eric J. Bohm, James C. Phillips, Laximant V. Kalé, Terry R. Jones
SC4
2011 Simulation-Based Performance Analysis and Tuning for a Two-Level Directly Connected System
abstract
Hardware and software co-design is becoming increasingly important due to complexities in supercomputing architectures. Simulating applications before there is access to the real hardware can assist machine architects in making better design decisions that can optimize application performance. At the same time, the application and runtime can be optimized and tuned beforehand. BigSim is a simulation-based performance prediction framework designed for these purposes. It can be used to perform packet-level network simulations of parallel applications using existing parallel machines. In this paper, we demonstrate the utility of BigSim in analyzing and optimizing parallel application performance for future systems based on the PERCS network. We present simulation studies using benchmarks and real applications expected to run on future supercomputers. Future petascale systems will have more than 100,000 cores, and we present simulations at that scale.
Ehsan Totoni, Abhinav Bhatele, Eric J. Bohm, Celso L. Mendes, Ryan M. Mokos, Gengbin Zheng, Laxmikant V. Kalé
ICPADS3
2011 Enabling and scaling biomolecular simulations of 100 million atoms on petascale machines with a multicore-optimized message-driven runtime
abstract
A 100-million-atom biomolecular simulation with NAMD is one of the three benchmarks for the NSF-funded sustainable petascale machine. Simulating this large molecular system on a petascale machine presents great challenges, including handling I/O, large memory footprint and getting good strong-scaling results. In this paper, we present parallel I/O techniques to enable the simulation. A new SMP model is designed to efficiently utilize ubiquitous wide multicore clusters by extending the Charm++ asynchronous message-driven runtime. We exploit node-aware techniques to optimize both the application and the underlying SMP runtime. Hierarchical load balancing is further exploited to scale NAMD to the full Jaguar PF Cray XT5 (224,076 cores) at Oak Ridge National Laboratory, both with and without PME full electrostatics, achieving 93% parallel efficiency (vs 6720 cores) at 9 ms per step for a simple cutoff calculation. Excellent scaling is also obtained on 65,536 cores of the Intrepid Blue Gene/P at Argonne National Laboratory.
Yanhua Sun, Gengbin Zheng, Eric J. Bohm, Laxmikant V. Kalé, James C. Phillips, Chris Harrison 0001
SC4
2011 Optimizing communication for Charm++ applications by reducing network contention
abstract
Abstract Optimal network performance is critical for efficient parallel scaling of communication‐bound applications on large machines. No‐load latencies do not increase significantly with the number of hops traveled when wormhole routing is deployed. Yet, we and others have recently shown that in the presence of contention, message latencies can grow substantially large. Hence, task mapping strategies should take the topology of the machine into account on large machines. In this paper, we present topology aware mapping as a technique to optimize communication on three‐dimensional mesh interconnects and hence improve the performance. Our methodology is facilitated by the idea of object‐based decomposition used in Charm++ which separates the processes of decomposition from mapping of computation to processors and allows a more flexible mapping based on communication patterns between objects. Exploiting this and the topology of the allocated job partition, we present mapping strategies for a production code, OpenAtom to improve the overall performance and scaling. OpenAtom presents complex communication scenarios of interaction involving multiple groups of objects and makes the mapping task a challenge. Results are presented for OpenAtom on up to 16 384 processors of Blue Gene/L, 8192 processors of Blue Gene/P and 2048 processors of Cray XT3. Copyright © 2010 John Wiley & Sons, Ltd.
Abhinav Bhatele, Eric J. Bohm, Laxmikant V. Kalé
Concurr. Comput. Pract. Exp.2
2010 Simulating Large Scale Parallel Applications Using Statistical Models for Sequential Execution Blocks
abstract
Predicting sequential execution blocks of a large scale parallel application is an essential part of accurate prediction of the overall performance of the application. When simulating a future machine, or a prototype system only available at a small scale, it becomes a significant challenge. Using hardware simulators may not be feasible due to excessively slowed down execution times and insufficient resources. The difficulty of these challenges increases proportionally with the scale of the simulation. In this paper, we propose an approach based on statistical models to accurately predict the performance of the sequential execution blocks that comprise a parallel application. We deployed these techniques in a trace-driven simulation framework to capture both the detailed behavior of the application as well as the overall predicted performance. The technique is validated using both synthetic benchmarks and the NAMD application.
Gengbin Zheng, Gagan Raj Gupta 0001, Eric J. Bohm, Isaac Dooley, Laxmikant V. Kalé
ICPADS3
2009 A Case Study of Communication Optimizations on 3D Mesh Interconnects
Abhinav Bhatele, Eric J. Bohm, Laxmikant V. Kalé
Euro-Par2
2009 Topology aware task mapping techniques: an api and case study
abstract
Optimal network performance is critical to efficient parallel scaling for communication-bound applications on large machines. With wormhole routing, no-load latencies do not increase significantly with number of hops traveled. Yet, we, and others have recently shown that in presence of contention, message latencies can grow substantially large. Hence task mapping strategies should take the topology of the machine into account on large machines. This poster presents a uniform API which provides topology information on 3D tori like IBM Blue Gene and Cray XT machines. We present techniques to use this API to improve performance. The API can be used by user-level codes to obtain information about allocated partitions at runtime which is essential for mapping.We motivate why it is important to consider network topology, using a simple 3D Stencil kernel. We then present mapping strategies for a production code, OpenAtom, running on three-dimensional torus and mesh topologies. OpenAtom presents complex communication scenarios of interaction between multiple groups of objects. Results are presented in the context of 3D Stencil and OpenAtom on up to 16,384 processors of Blue Gene/L, 8,192 processors of Blue Gene/P and 2,048 processors of Cray XT3.
Abhinav Bhatele, Eric J. Bohm, Laxmikant V. Kalé
PPoPP2
2006 Poster reception - Charm++ simplifies coding for the cell processor
abstract
While the Cell processor, jointly developed by IBM, Sony, and Toshiba, has great computational power, it also presents many challenges including portability and ease of programming. We have been adapting the Charm++ Runtime System to utilize the Cell. We believe that the Charm++ model fits well with the Cell for many reasons: encapsulation of data, effective prefetching, the ability to peak ahead in message queues, virtualization, etc. To these ends, we have developed the Offload API (an independent code) which allows Charm++ applications to easily take advantage of the Cell. Our goal is to allow Charm++ programs to run on Cell-based and non-Cell-based platforms without modification to application code. Example Charm++ programs using the Offload API already exist. We have also begun modifying NAMD, a popular molecular dynamics code, to use the Cell. In this poster, we plan to present current progress and future plans for this work.
David M. Kunzman, Gengbin Zheng, Eric J. Bohm, James C. Phillips, Laxmikant V. Kalé
SC3