Steve H. Langer

dblp:120/5487 · also Steven H. Langer · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0001-5297-1165ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 38% Parallel and multicore computing · 26% Performance modeling and evaluation · 23%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
application porting
0.412019
Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019
Parallel and multicore computing
programming models
0.412019
Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019
Interconnection networks and networks-on-chip › network topology › tree networks
fat-tree network
0.212016
Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree · SC 2016
Performance modeling and evaluation
workload characterization
0.212016
Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree · SC 2016
Performance modeling and evaluation
performance variability
0.212013
There goes the neighborhood: performance degradation due to nearby jobs · SC 2013
Visualization and visual analytics › graph visualization
network traffic visualization
0.112012
Visualizing Network Traffic to Understand the Performance of Massively Parallel Simulations · IEEE Trans. Vis. Comput. Graph. 2012
Visualization and visual analytics › software visualization
performance visualization
0.112012
Visualizing Network Traffic to Understand the Performance of Massively Parallel Simulations · IEEE Trans. Vis. Comput. Graph. 2012
High-performance computing
communication characterization
0.112012
Visualizing Network Traffic to Understand the Performance of Massively Parallel Simulations · IEEE Trans. Vis. Comput. Graph. 2012
Parallel and multicore computing › parallel computing
parallel application performance
0.112012
Visualizing Network Traffic to Understand the Performance of Massively Parallel Simulations · IEEE Trans. Vis. Comput. Graph. 2012
Parallel and multicore computing
task allocation
0.112012
Mapping applications with collectives over sub-communicators on torus networks · SC 2012
Performance modeling and evaluation
benchmarking
0.112019
Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019
High-performance computing
collective communication
0.012012
Mapping applications with collectives over sub-communicators on torus networks · SC 2012
Interconnection networks and networks-on-chip › network topology
torus network
0.012012
Mapping applications with collectives over sub-communicators on torus networks · SC 2012

Methods — techniques the papers use, named apart from their topics

linked 2d and 3d views · 0.3case study · 0.3trace analysis · 0.2empirical study · 0.2performance measurement · 0.2topology-aware mapping · 0.1
YearPublicationVenuePosition
2022 Enabling machine learning-ready HPC ensembles with Merlin
Jayson Luc Peterson, Benjamin Bay, Joe Koning, Peter B. Robinson, Jessica Semler, Jeremy White, Rushil Anirudh, Kevin Athey, Peer-Timo Bremer, Francesco Di Natale, Jim Gaffney, Sam Ade Jacobs, Bhavya Kailkhura, Bogdan Kustowski, Steve H. Langer, Brian K. Spears, Jayaraman J. Thiagarajan, Brian Van Essen, Jae-Seung Yeom
Future Gener. Comput. Syst.16
2019 Preparation and optimization of a diverse workload for a large-scale heterogeneous system
abstract
Productivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience.
Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz
SC16
2016 Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree
abstract
Understanding the characteristics and requirements of applications that run on commodity clusters is key to properly configuring current machines and, more importantly, procuring future systems effectively. There are only a few studies, however, that are current and characterize realistic workloads. For HPC practitioners and researchers, this limits our ability to design solutions that will have an impact on real systems. We present a systematic study that characterizes applications with an emphasis on communication requirements. It includes cluster utilization data, identifying a representative set of applications from a U.S. Department of Energy laboratory, and characterizing their communication requirements. The driver for this work is understanding application sensitivity to a tapered fat-tree network. These results provided key insights into the procurement of our next generation commodity systems. We believe this investigation can provide valuable input to the HPC community in terms of workload characterization and requirements from a large supercomputing center.
Edgar A. León, Ian Karlin, Abhinav Bhatele, Steve H. Langer, Christopher M. Chambreau, Louis H. Howell, Trent D'Hooge, Matthew L. Leininger
SC4
2014 Optimizing the performance of parallel applications on a 5D torus via task mapping
abstract
Six of the ten fastest supercomputers in the world in 2014 use a torus interconnection network for message passing between compute nodes. Torus networks provide high bandwidth links to near-neighbors and low latencies over multiple hops on the network. However, large diameters of such networks necessitate a careful placement of parallel tasks on the compute nodes to minimize network congestion. This paper presents a methodological study of optimizing application performance on a five-dimensional torus network via the technique of topology-aware task mapping. Task mapping refers to the placement of processes on compute nodes while carefully considering the network topology between the nodes and the communication behavior of the application. We focus on the IBM Blue Gene/Q machine and two production applications - a laser-plasma interaction code called pF3D and a lattice QCD application called MILC. Optimizations presented in the paper improve the communication performance of pF3D by 90% and that of MILC by up to 47%.
Abhinav Bhatele, Katherine E. Isaacs, Ronak Buch, Todd Gamblin, Steve H. Langer, Laxmikant V. Kalé
HiPC6
2013 There goes the neighborhood: performance degradation due to nearby jobs
abstract
Predictable performance is important for understanding and alleviating application performance issues; quantifying the effects of source code, compiler, or system software changes; estimating the time required for batch jobs; and determining the allocation requests for proposals. Our experiments show that on a Cray XE system, the execution time of a communication-heavy parallel application ranges from 28% faster to 41% slower than the average observed performance. Blue Gene systems, on the other hand, demonstrate no noticeable run-to-run variability. In this paper, we focus on Cray machines and investigate potential causes for performance variability such as OS jitter, shape of the allocated partition, and interference from other jobs sharing the same network links. Reducing such variability could improve overall throughput at a computer center and save energy costs.
Abhinav Bhatele, Kathryn Mohror, Steve H. Langer, Katherine E. Isaacs
SC3
2012 Mapping applications with collectives over sub-communicators on torus networks
abstract
The placement of tasks in a parallel application on specific nodes of a supercomputer can significantly impact performance. Traditionally, this task mapping has focused on reducing the distance between communicating tasks on the physical network. This minimizes the number of hops that point-to-point messages travel and thus reduces link sharing between messages and contention. However, for applications that use collectives over sub-communicators, this heuristic may not be optimal. Many collectives can benefit from an increase in bandwidth even at the cost of an increase in hop count, especially when sending large messages. For example, placing communicating tasks in a cube configuration rather than a plane or a line on a torus network increases the number of possible paths messages might take. This increases the available bandwidth which can lead to significant performance gains. We have developed Rubik, a tool that provides a simple and intuitive interface to create a wide variety of mappings for structured communication patterns. Rubik supports a number of elementary operations such as splits, tilts, or shifts, that can be combined into a large number of unique patterns. Each operation can be applied to disjoint groups of processes involved in collectives to increase the effective bandwidth. We demonstrate the use of Rubik for improving performance of two parallel codes, pF3D and Qbox, which use collectives over sub-communicators.
Abhinav Bhatele, Todd Gamblin, Steve H. Langer, Peer-Timo Bremer, Erik W. Draeger, Bernd Hamann, Katherine E. Isaacs, Aaditya G. Landge, Joshua A. Levine, Valerio Pascucci, Martin Schulz 0001, Charles H. Still
SC3
2012 Visualizing Network Traffic to Understand the Performance of Massively Parallel Simulations
abstract
The performance of massively parallel applications is often heavily impacted by the cost of communication among compute nodes. However, determining how to best use the network is a formidable task, made challenging by the ever increasing size and complexity of modern supercomputers. This paper applies visualization techniques to aid parallel application developers in understanding the network activity by enabling a detailed exploration of the flow of packets through the hardware interconnect. In order to visualize this large and complex data, we employ two linked views of the hardware network. The first is a 2D view, that represents the network structure as one of several simplified planar projections. This view is designed to allow a user to easily identify trends and patterns in the network traffic. The second is a 3D view that augments the 2D view by preserving the physical network topology and providing a context that is familiar to the application developers. Using the massively parallel multi-physics code pF3D as a case study, we demonstrate that our tool provides valuable insight that we use to explain and optimize pF3D's performance on an IBM Blue Gene/P system.
Aaditya G. Landge, Joshua A. Levine, Abhinav Bhatele, Katherine E. Isaacs, Todd Gamblin, Martin Schulz 0001, Steve H. Langer, Peer-Timo Bremer, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.7