Mark A. Taylor

dblp:22/5969 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0002-9267-2554ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 70% Parallel and multicore computing · 16% Interconnection networks and networks-on-chip · 12%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.622020
A performance-portable nonhydrostatic atmospheric dycore for the energy exascale earth system model running at cloud-resolving resolutions · SC 2020
Performance of the community earth system model · SC 2011
High-performance computing
scientific computing systems
0.622020
A performance-portable nonhydrostatic atmospheric dycore for the energy exascale earth system model running at cloud-resolving resolutions · SC 2020
Performance of the community earth system model · SC 2011
High-performance computing › scientific computing systems
atmospheric modeling
0.412020
A performance-portable nonhydrostatic atmospheric dycore for the energy exascale earth system model running at cloud-resolving resolutions · SC 2020
High-performance computing › performance engineering
performance portability
0.412020
A performance-portable nonhydrostatic atmospheric dycore for the energy exascale earth system model running at cloud-resolving resolutions · SC 2020
Interconnection networks and networks-on-chip
network congestion
0.412019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
Parallel and multicore computing
task allocation
0.412019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
High-performance computing › scientific computing systems
climate modeling
0.112011
Performance of the community earth system model · SC 2011
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.112019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
benchmarking
0.012011
Performance of the community earth system model · SC 2011

Methods — techniques the papers use, named apart from their topics

kokkos · 0.4c++ refactoring · 0.4GPU acceleration · 0.4graph partitioning · 0.4geometric partitioning · 0.4performance tuning · 0.1numerical algorithm evaluation · 0.1
YearPublicationVenuePosition
2020 A performance-portable nonhydrostatic atmospheric dycore for the energy exascale earth system model running at cloud-resolving resolutions
abstract
We present an effort to port the nonhydrostatic atmosphere dynamical core of the Energy Exascale Earth System Model (E3SM) to efficiently run on a variety of architectures, including conventional CPU, many-core CPU, and GPU. We specifically target cloud-resolving resolutions of 3 km and 1 km. To express on-node parallelism we use the C++ library Kokkos, which allows us to achieve a performance portable code in a largely architecture-independent way. Our C++ implementation is at least as fast as the original Fortran implementation on IBM Power9 and Intel Knights Landing processors, proving that the code refactor did not compromise the efficiency on CPU architectures. On the other hand, when using the GPUs, our implementation is able to achieve 0.97 Simulated Years Per Day, running on the full Summit supercomputer. To the best of our knowledge, this is the most achieved to date by any global atmosphere dynamical core running at such resolutions.
Luca Bertagna, Oksana Guba, Mark A. Taylor, James G. Foucar, Jeffrey M. Larkin, Andrew M. Bradley, Sivasankaran Rajamanickam, Andrew G. Salinger
SC3
2019 Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks
abstract
We present a new method for reducing parallel applications' communication time by mapping their MPI tasks to processors in a way that lowers the distance messages travel and the amount of congestion in the network. Assuming geometric proximity among the tasks is a good approximation of their communication interdependence, we use a geometric partitioning algorithm to order both the tasks and the processors, assigning task parts to the corresponding processor parts. In this way, interdependent tasks are assigned to “nearby” cores in the network. We also present a number of algorithmic optimizations that exploit specific features of the network or application to further improve the quality of the mapping. We specifically address the case of sparse node allocation, where the nodes assigned to a job are not necessarily located in a contiguous block nor within close proximity to each other in the network. However, our methods generalize to contiguous allocations as well, and results are shown for both contiguous and non-contiguous allocations. We show that, for the structured finite difference mini-application MiniGhost, our mapping methods reduced communication time up to 75 percent relative to MiniGhost's default mapping on 128K cores of a Cray XK7 with sparse allocation. For the atmospheric modeling code E3SM/HOMME, our methods reduced communication time up to 31% on 16K cores of an IBM BlueGene/Q with contiguous allocation.
Mehmet Deveci, Karen D. Devine, Kevin T. Pedretti, Mark A. Taylor, Sivasankaran Rajamanickam, Ümit V. Çatalyürek
IEEE Trans. Parallel Distributed Syst.4
2011 Performance of the community earth system model
abstract
The Community Earth System Model (CESM), released in June 2010, incorporates new physical process and new numerical algorithm options, significantly enhancing simulation capabilities over its predecessor, the June 2004 release of the Community Climate System Model. CESM also includes enhanced performance tuning options and performance portability capabilities. This paper describes performance and performance scaling on both the Cray XT5 and the IBM BG/P for four representative production simulations, varying both problem size and enabled physical processes. The paper also describes preliminary performance results for high resolution simulations using over 200,000 processor cores, indicating the promise of ongoing work in numerical algorithms and where further work is required.
Patrick H. Worley, Arthur A. Mirin, Anthony P. Craig, Mark A. Taylor, John M. Dennis, Mariana Vertenstein
SC4
2004 Architecture of LA-MPI, A Network-Fault-Tolerant MPI
abstract
Summary form only given. We discuss the unique architectural elements of the Los Alamos message passing interface (LA-MPI), a high-performance, network-fault-tolerant, thread-safe MPI library. LA-MPI is designed for use on terascale clusters which are inherently unreliable due to their sheer number of system components and trade-offs between cost and performance. We examine in detail the design concepts used to implement LA-MPI. These include reliability features, such as application-level checksumming, message retransmission, and automatic message rerouting. Other key performance enhancing features, such as concurrent message routing over multiple, diverse network adapters and protocols, and communication-specific optimizations (e.g., shared memory) are examined.
Rob T. Aulwes, David J. Daniel, Nehal N. Desai, Richard L. Graham, L. Dean Risinger, Mark A. Taylor, Timothy S. Woodall, Mitchel W. Sukalski
IPDPS6