Michael J. Aftosmis

dblp:86/4874 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 50% High-performance computing · 38% Interconnection networks and networks-on-chip · 12%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
parallel system performance
0.112005
High Resolution Aerospace Applications using the NASA Columbia Supercomputer · SC 2005
Interconnection networks and networks-on-chip › cluster interconnect
infiniband
0.012005
High Resolution Aerospace Applications using the NASA Columbia Supercomputer · SC 2005
Performance modeling and evaluation › network performance analysis
interconnect performance
0.012005
High Resolution Aerospace Applications using the NASA Columbia Supercomputer · SC 2005

Methods — techniques the papers use, named apart from their topics

reynolds-averaged navier-stokes · 0.1multigrid · 0.1
YearPublicationVenuePosition
2011 Performance Analysis of CFD Application Cart3D Using MPInside and Performance Monitor Unit Data on Nehalem and Westmere Based Supercomputers
abstract
Cart3D is a computational fluid dynamics (CFD) application aimed at conceptual and preliminary design of aerospace vehicles with complex geometries. It is widely used by design engineers at NASA, Department of Defense and aerospace companies in the USA. We present detailed performance analysis of Cart3D using two tools SGI MPInside and op_scope that collects hardware counter data from Intel Performance Monitoring Unit (PMU) on supercomputers based on Nehalem micro-architecture. Using these tools, we have done dynamic profiling of Cart3D (compute time, communication time and I/O time), along with dynamic profiling of MPI functions (MPI_Sendrecv, MPI_Bcast, MPI_Isend, MPI_Irecv, MPI_Allreduce, MPI_Barrier, etc.) with respect to message size of each rank and time consumed by each function. MPI communication is further analyzed by studying the performance of MPI functions used in this application as a function of message size and number of cores. Using these tools we have also studied efficiency of the processor to measure its effective utilization, efficiency of the floating-point units, percentage of vectorization and percentage of data coming from L2 cache, L3 cache, and main memory. This study was performed on two computing sub-systems based on quad-core Nehalem-EP and hex-core West mere-EP processors that are part of Pleiades an SGI Altix ICE at NASA Ames Research Center.
Subhash Saini, Piyush Mehrotra, Kenichi Taylor, Michael J. Aftosmis, Rupak Biswas
HPCC4
2005 High Resolution Aerospace Applications using the NASA Columbia Supercomputer
abstract
This paper focuses on the parallel performance of two high-performance aerodynamic simulation packages on the newly installed NASA Columbia supercomputer. These packages include both a high-fidelity, unstructured, Reynolds-averaged Navier-Stokes solver, and a fully-automated inviscid flow package for cut-cell Cartesian grids. The complementary combination of these two simulation codes enables high-fidelity characterization of aerospace vehicle design performance over the entire flight envelope through extensive parametric analysis and detailed simulation of critical regions of the flight envelope. Both packages are industrial-level codes designed for complex geometry and incorporate customized multigrid solution algorithms. The performance of these codes on Columbia is examined using both MPI and OpenMP and using both the NUMAlink and InfiniBand interconnect fabrics. Numerical results demonstrate good scalability on up to 2016 cpus using the NUMAlink4 interconnect, with measured computational rates in the vicinity of 3 TFLOP/s, while InfiniBand showed some performance degradation at high CPU counts, particularly with multigrid. Nonetheless, the results are encouraging enough to indicate that larger test cases using combined MPI/OpenMP communication should scale well on even more processors.
Dimitri J. Mavriplis, Michael J. Aftosmis, Marsha J. Berger
SC2
2005 Performance of a new CFD flow solver using a hybrid programming paradigm
Marsha J. Berger, Michael J. Aftosmis, D. D. Marshall, Scott M. Murman
J. Parallel Distributed Comput.2