Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Timothy C. Germann

dblp:27/4143 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2018
0000-0002-6813-238XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 70% GPUs and heterogeneous computing · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
heterogeneous supercomputing
0.112008
369 Tflop/s molecular dynamics simulations on the Roadrunner general-purpose heterogeneous supercomputer · SC 2008
High-performance computing › scientific computing systems
molecular dynamics simulation
0.112008
369 Tflop/s molecular dynamics simulations on the Roadrunner general-purpose heterogeneous supercomputer · SC 2008
High-performance computing
scientific computing systems
0.112008
369 Tflop/s molecular dynamics simulations on the Roadrunner general-purpose heterogeneous supercomputer · SC 2008
High-performance computing › large-scale simulation
parallel molecular dynamics
0.012008
369 Tflop/s molecular dynamics simulations on the Roadrunner general-purpose heterogeneous supercomputer · SC 2008

Methods — techniques the papers use, named apart from their topics

lennard-jones potential · 0.1MPI · 0.1
YearPublicationVenuePosition
2018 The basic matrix library (BML) for quantum chemistry
Nicolas Bock, Christian F. A. Negre, Susan M. Mniszewski, Jamaludin Mohd-Yusof, Bálint Aradi, Jean-Luc Fattebert, Daniel Osei-Kuffuor, Timothy C. Germann, Anders M. N. Niklasson
J. Supercomput.8
2009 369 Tflop/s molecular dynamics simulations on the petaflop hybrid supercomputer 'Roadrunner'
abstract
Abstract We describe the implementation of a short‐range parallel molecular dynamics (MD) code, SPaSM, on the heterogeneous general‐purpose Roadrunner supercomputer. Each Roadrunner ‘TriBlade’ compute node consists of two AMD Opteron dual‐core microprocessors and four IBM PowerXCell 8i enhanced Cell microprocessors (each consisting of one PPU and eight SPU cores), so that there are four MPI ranks per node, each with one Opteron and one Cell. We will briefly describe the Roadrunner architecture and some of the initial hybrid programming approaches that have been taken, focusing on the SPaSM application as a case study. An initial ‘evolutionary’ port, in which the existing legacy code runs with minor modifications on the Opterons and the Cells are only used to compute interatomic forces, achieves roughly a 2× speedup over the unaccelerated code. On the other hand, our ‘revolutionary’ implementation adopts a Cell‐centric view, with data structures optimized for, and living on, the Cells. The Opterons are mainly used to direct inter‐rank communication and perform I/O‐heavy periodic analysis, visualization, and checkpointing tasks. The performance measured for our initial implementation of a standard Lennard–Jones pair potential benchmark reached a peak of 369 Tflop/s double‐precision floating‐point performance on the full Roadrunner system (27.7% of peak), nearly 10× faster than the unaccelerated (Opteron‐only) version. Copyright © 2009 John Wiley & Sons, Ltd.
Timothy C. Germann, Kai Kadau, Sriram Swaminarayan
Concurr. Comput. Pract. Exp.1
2008 369 Tflop/s molecular dynamics simulations on the Roadrunner general-purpose heterogeneous supercomputer
abstract
We present timing and performance numbers for a short-range parallel molecular dynamics (MD) code, SPaSM, that has been rewritten for the heterogeneous Roadrunner supercomputer. Each Roadrunner compute node consists of two AMD Opteron dualcore microprocessors and four PowerXCell 8i enhanced Cell microprocessors, so that there are four MPI ranks per node, each with one Opteron and one Cell. The interatomic forces are computed on the Cells (each with one PPU and eight SPU cores), while the Opterons are used to direct inter-rank communication and perform I/O-heavy periodic analysis, visualization, and checkpointing tasks. The performance measured for our initial implementation of a standard Lennard-Jones pair potential benchmark reached a peak of 369 Tflop/s double-precision floating-point performance on the full Roadrunner system (27.7% of peak), corresponding to 124 MFlop/Watt/s at a price of approximately 3.69 MFlops/dollar. We demonstrate an initial target application, the jetting and ejection of material from a shocked surface.
Sriram Swaminarayan, Kai Kadau, Timothy C. Germann, Gordon C. Fossum
SC3