Robert Preissl

dblp:28/5392 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 33% Emerging computing paradigms · 19% Performance modeling and evaluation · 19%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
0.112012
Compass: a scalable simulator for an architecture for cognitive computing · SC 2012
Performance modeling and evaluation
simulation
0.112012
Compass: a scalable simulator for an architecture for cognitive computing · SC 2012
Distributed systems
communication optimization
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
Parallel and multicore computing › parallel computing › parallel communication
one-sided communication
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
Parallel and multicore computing
parallel programming models
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
High-performance computing › supercomputer architecture
blue gene/q
0.012012
Compass: a scalable simulator for an architecture for cognitive computing · SC 2012
High-performance computing › large-scale simulation
massively parallel simulation
0.012012
Compass: a scalable simulator for an architecture for cognitive computing · SC 2012
Computational science and engineering
computational physics
0.012011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011

Methods — techniques the papers use, named apart from their topics

one-sided CAF · 0.2PGAS · 0.2OpenMP · 0.2MPI · 0.2parallel compiler · 0.1multithreaded simulation · 0.1PGAS communication · 0.1
YearPublicationVenuePosition
2012 Compass: a scalable simulator for an architecture for cognitive computing
abstract
Inspired by the function, power, and volume of the organic brain, we are developing TrueNorth, a novel modular, non-von Neumann, ultra-low power, compact architecture. TrueNorth consists of a scalable network of neurosynaptic cores, with each core containing neurons, dendrites, synapses, and axons. To set sail for TrueNorth, we developed Compass, a multi-threaded, massively parallel functional simulator and a parallel compiler that maps a network of long-distance pathways in the macaque monkey brain to TrueNorth. We demonstrate near-perfect weak scaling on a 16 rack IBM® Blue Gene®/Q (262144 CPUs, 256 TB memory), achieving an unprecedented scale of 256 million neurosynaptic cores containing 65 billion neurons and 16 trillion synapses running only 388x slower than real time with an average spiking rate of 8.1 Hz. By using emerging PGAS communication primitives, we also demonstrate 2x better real-time performance over MPI primitives on a 4 rack Blue Gene/P (16384 CPUs, 16 TB memory).
Robert Preissl, Theodore M. Wong, Pallab Datta, Myron Flickner, Raghavendra Singh, Steven K. Esser, William P. Risk, Horst D. Simon, Dharmendra S. Modha
SC1
2011 Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms
abstract
We present novel parallel language constructs for the communication intensive part of a magnetic fusion simulation code. The focus of this work is the shift phase of charged particles of a tokamak simulation code in toroidal geometry. We introduce new hybrid PGAS/OpenMP implementations of highly optimized hybrid MPI/OpenMP based communication kernels. The hybrid PGAS implementations use an extension of standard hybrid programming techniques, enabling the distribution of high communication work loads of the underlying kernel among OpenMP threads. Building upon lightweight one-sided CAF (Fortran 2008) communication techniques, we also show the benefits of spreading out the communication over a longer period of time, resulting in a reduction of bandwidth requirements and a more sustained communication and computation overlap. Experiments on up to 130560 processors are conducted on the NERSC Hopper system, which is currently the largest HPC platform with hardware support for one-sided communication and show performance improvements of 52 % at highest concurrency.
Robert Preissl, Nathan Wichmann, Bill Long, John Shalf, Stéphane Ethier, Alice E. Koniges
SC1
2010 Exploitation of Dynamic Communication Patterns through Static Analysis
abstract
Collective operations can have a large impact on the performance of parallel applications. However, the ideal implementation of a particular collective communication often depends on both the application and the targeted machine structure. Our approach combines dynamic and static analysis techniques to identify common collective communication patterns expressed as point-to-point calls and transforms them into equivalent MPI collectives. We first detect potential collective communication patterns in runtime traces and associate them with the corresponding source code regions. If our static analysis verifies that the introduction of collectives is safe for any program flow, we then replace the original communication primitives with their collective counterpart. In this paper we introduce the necessary algorithms to determine the safety of these transformations and we demonstrate several use cases, including automatic use of new extensions to the MPI standard such as nonblocking collective operations. The use of dynamic analysis significantly reduces compile times, resulting in a speed-up of about 50 for source transformations of HPL due to more directed analysis capabilities and also dramatically decreases complexity of the underlying static analysis.
Robert Preissl, Bronis R. de Supinski, Martin Schulz 0001, Daniel J. Quinlan, Dieter Kranzlmüller, Thomas Panas
ICPP1
2010 Transforming MPI source code based on communication patterns
Robert Preissl, Martin Schulz 0001, Dieter Kranzlmüller, Bronis R. de Supinski, Daniel J. Quinlan
Future Gener. Comput. Syst.1
2008 Detecting Patterns in MPI Communication Traces
abstract
Since processor counts in supercomputers are increasing dramatically, efficient interprocessor communication is becoming even more important for the applications that run on them. A high level, abstract understanding of an application's communication behavior would not only simplify debugging of that communication but would also support more directed performance optimization. We explore automated identification of communication patterns to provide that high level abstraction. We introduce an algorithm to extract communication patterns from MPI traces automatically. Our algorithm first finds locally repeating sequences and then iteratively grows them into global patterns. We demonstrate our technique on three realistic codes using traces from up to 128 processors. Our results show that our approach detects the underlying communication pattern within reasonable time andmemory constraints, even for large trace sizes.
Robert Preissl, Thomas Köckerbauer, Martin Schulz 0001, Dieter Kranzlmüller, Bronis R. de Supinski, Daniel J. Quinlan
ICPP1