John M. Dennis

dblp:29/1090 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
0since 2021 · last 2017
0000-0002-2119-8242ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-authorArtificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 70% Performance modeling and evaluation · 17% Storage systems · 13%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Environmental and earth informatics · 90% Computational science and engineering · 10%
Computer networks
1 paper
Network measurement and analytics · 50% Routing and switching · 50%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.322015
Improving the scalability of the ocean barotropic solver in the community earth system model · SC 2015
Performance of the community earth system model · SC 2011
Environmental and earth informatics
climate modeling
0.322015
Improving the scalability of the ocean barotropic solver in the community earth system model · SC 2015
A methodology for evaluating the impact of data compression on climate simulation data · HPDC 2014
High-performance computing
performance optimization at scale
0.222015
Performance of the community earth system model · SC 2011
Improving the scalability of the ocean barotropic solver in the community earth system model · SC 2015
Storage systems
data compression
0.212014
A methodology for evaluating the impact of data compression on climate simulation data · HPDC 2014
High-performance computing
lossy compression
0.212014
A methodology for evaluating the impact of data compression on climate simulation data · HPDC 2014
Performance modeling and evaluation
benchmarking
0.222012
Establishing a Miniapp as a programmability proxy · PPoPP 2012
Performance of the community earth system model · SC 2011
High-performance computing › scientific computing systems
climate modeling
0.222011
Performance of the community earth system model · SC 2011
Terascale spectral element dynamical core for atmospheric general circulation models · SC 2001
Network measurement and analytics
traffic characterization
0.112012
On using virtual circuits for GridFTP transfers · SC 2012
Routing and switching › packet switching
virtual circuit
0.112012
On using virtual circuits for GridFTP transfers · SC 2012
High-performance computing › data transfer
scientific data transfer
0.012012
On using virtual circuits for GridFTP transfers · SC 2012
High-performance computing › large-scale simulation
atmospheric general circulation model
0.012001
Terascale spectral element dynamical core for atmospheric general circulation models · SC 2001

Methods — techniques the papers use, named apart from their topics

preconditioned chebyshev-type iterative method · 0.4error vector propagation preconditioner · 0.4compression algorithms · 0.4log analysis · 0.3performance tuning · 0.1numerical algorithm evaluation · 0.1semi-implicit time stepping · 0.1spectral-element method · 0.0spectral element method · 0.0
YearPublicationVenuePosition
2017 Assessing Representativeness of Kernels Using Descriptive Statistics
abstract
A kernel or mini-app is a self-contained small application that retains certain characteristics of the original application [7]. Working on a kernel or mini-app in the place of the original application can dramatically reduce the resources and effort required for performing software tasks such as performance optimization and porting to new platforms. However, using kernel as a proxy is based on the assumption that it represents the original application in the context of how it is being used. In this paper, we introduce an extension to the Fortran Kernel Generator (KGen) which is an automated kernel extraction tool [1]. The extension allows comparison of the execution characteristics between the original application and the generated kernel using descriptive statistics. From the comparison, the user is provided with statistics that provide information on the degree and context of representativeness of the kernel. KGen also utilizes the information generated to help it to automatically improve representativeness of the kernels whilst reducing the size of the workload generated. We applied this extension to three kernels. One is generated from a Fortran scientific library and the remaining two are generated from an earth system model. We have demonstrated that the descriptive statistics provided in the enhancement provide not only quantitative metrics and context of representativeness but also a way to improve the quality of representativeness of the kernels generated.
Youngsung Kim, John M. Dennis, Christopher Kerr
CLUSTER2
2016 A new parallel python tool for the standardization of earth system model data
abstract
We have developed a new parallel Python tool for the standardization of Earth System Model (ESM) data for publication as part of Model Intercomparison Projects (MIPs). It was specifically designed to aid Community Earth System Model (CESM) scientists at the National Center for Atmospheric Research (NCAR) in preparation for the Coupled Model Intercomparison Project, Phase 6 (CMIP6), expected to start in early 2017. However, the tool is general to any and all MIPs and ESMs. The tool is implemented with MPI parallelism using mpi4py, and it performs the data standardization computation with a directed acyclic graph (DAG) data structure capable of streaming data from ESM input data to standardized output files. In this paper, we describe the tool, its design and testing.
Kevin Paul, Sheri A. Mickelson, John M. Dennis
IEEE BigData3
2016 Minimal Aggregated Shared Memory Messaging on Distributed Memory Supercomputers
abstract
Many high-performance distributed memory applications rely on point-to-point messaging using the Message Passing Interface (MPI). Due to the latency of the network, and other costs, this communication can limit the scalability of an application when run on high node counts of distributed memory supercomputers. Communication costs are further increased on modern multi-and many-core architectures, when using more than one MPI process per node, as each process sends and receives messages independently, inducing multiple latencies and contention for resources. In this paper, we use shared memory constructs available in the MPI 3.0 standard to implement an aggregated communication method to minimize the number of inter-node messages to reduce these costs. We compare the performance of this Minimal Aggregated SHared Memory (MASHM) messaging to the standard point-to-point implementation on large-scale supercomputers, where we see that MASHM leads to enhanced strong scalability of a weighted Jacobi relaxation. For this application, we also see that the use of shared memory parallelism through MASHM and MPI 3.0 can be more efficient than using Open Multi-Processing (OpenMP). We then present a model for the communication costs of MASHM which shows that this method achieves its goal of reducing latency costs while also reducing bandwidth costs. Finally, we present MASHM as an open source library to facilitate the integration of this efficient communication method into existing distributed memory applications.
Ben Jamroz, John M. Dennis
IPDPS2
2015 Light-weight parallel Python tools for earth system modeling workflows
abstract
In the last 30 years, earth system modeling has become increasingly data-intensive. The Community Earth System Model (CESM) response to the next Intergovernmental Panel on Climate Change (IPCC) assessment report (AR6) may require close to 1 Billion CPU hours of computation and generate up to 12 PB of raw data for post-processing. Existing post-processing tools are serial-only and impossibly slow with this much data. To improve the post-processing performance, our team has adopted a strategy of targeted replacement of the "bottleneck software" with light-weight parallel Python alternatives. This allows maximum impact with the least disruption to the CESM community and the shortest delivery time. We developed two light-weight parallel Python tools: one to convert model output from time-slice to time-series format, and one to perform fast time-averaging of time-series data. We present the motivation, approach, and results of these two tools, and our plans for future research and development.
Kevin Paul, Sheri A. Mickelson, John M. Dennis, Haiying Xu, David Brown 0006
IEEE BigData3
2015 Improving the scalability of the ocean barotropic solver in the community earth system model
abstract
High-resolution climate simulations are increasingly in demand and require tremendous computing resources. In the Community Earth SystemModel (CESM), the Parallel Ocean Model (POP) is computationally expensive for high-resolution grids (e.g., 0.1°) and is frequently the least scalable component of CESM for certain production simulations. In particular, the modified Preconditioned Conjugate Gradient (PCG), used to solve the elliptic system of equations in the barotropic mode, scales poorly at the high core counts, which is problematic for high-resolution simulations. In this work, we demonstrate that the communication costs in the barotropic solver occupy an increasing portion of the total POP execution time as core counts are increased. To mitigate this problem, we implement a preconditioned Chebyshev-type iterative method in POP (called P-CSI), which requires far fewer global reductions than PCG. We also develop an effective block preconditioner based on the Error Vector Propagation Method to attain a competitive convergence rate for P-CSI. We demonstrate that the improved scalability of P-CSI results in a 5.2x speedup of the barotropic mode in high-resolution POP on 16,875 cores, which yields a 1.7x speedup of the overall POP simulation. Further, we ensure that the new solver produces an ocean climate consistent with the original one via an ensemble-based statistical method.
Xiaomeng Huang, Allison H. Baker, Yu-heng Tseng, Frank O. Bryan, John M. Dennis, Guangwen Yang 0002
SC6
2014 A methodology for evaluating the impact of data compression on climate simulation data
abstract
High-resolution climate simulations require tremendous computing resources and can generate massive datasets. At present, preserving the data from these simulations consumes vast storage resources at institutions such as the National Center for Atmospheric Research (NCAR). The historical data generation trends are economically unsustainable, and storage resources are already beginning to limit science objectives. To mitigate this problem, we investigate the use of data compression techniques on climate simulation data from the Community Earth System Model. Ultimately, to convince climate scientists to compress their simulation data, we must be able to demonstrate that the reconstructed data reveals the same mean climate as the original data, and this paper is a first step toward that goal. To that end, we develop an approach for verifying the climate data and use it to evaluate several compression algorithms. We find that the diversity of the climate data requires the individual treatment of variables, and, in doing so, the reconstructed data can fall within the natural variability of the system, while achieving compression rates of up to 5:1.
Allison H. Baker, Haiying Xu, John M. Dennis, Michael N. Levy, Doug Nychka, Sheri A. Mickelson, Jim Edwards, Mariana Vertenstein, Al Wegener
HPDC3
2012 Establishing a Miniapp as a programmability proxy
abstract
Miniapps serve as test beds for prototyping and evaluating new algorithms, data structures, and programming models before incorporating such changes into larger applications. For the miniapp to accurately predict how a prototyped change would affect a larger application it is necessary that the miniapp be shown to serve as a proxy for that larger application. Although many benchmarks claim to proxy the performance for a set of large applications, little work has explored what criteria must be met for a benchmark to serve as a proxy for examining programmability. In this poster we describe criteria that can be used to establish that a miniapp serves as a performance and programmability proxy.
Andrew Stone, John M. Dennis, Michelle Mills Strout
PPoPP2
2012 On using virtual circuits for GridFTP transfers
abstract
The goal of this work is to characterize scientific data transfers and to determine the suitability of dynamic virtual circuit service for these transfers instead of the currently used IP-routed service. Specifically, logs collected by servers executing a commonly used scientific data transfer application, GridFTP, are obtained from three US super-computing/scientific research centers, NERSC, SLAC, and NCAR, and analyzed. Dynamic virtual circuit (VC) service, a relatively new offering from providers such as ESnet and Internet2, allows for the selection of a path on which a rate-guaranteed connection is established prior to data transfer. Given VC setup overhead, the first analysis of the GridFTP transfer logs characterizes the duration of sessions, where a session consists of multiple back-to-back transfers executed in batch mode between the same two GridFTP servers. Of the NCAR-NICS sessions analyzed, 56% of all sessions (90% of all transfers) would have been long enough to be served with dynamic VC service. An analysis of transfer logs across four paths, NCAR-NICS, SLAC-BNL, NERSC-ORNL and NERSC-ANL, shows significant throughput variance, where NICS, BNL, ORNL, and ANL are other US national laboratories. For example, on the NERSC-ORNL path, the inter-quartile range was 695 Mbps, with a maximum value of 3.64 Gbps and a minimum value of 758 Mbps. An analysis of the impact of various factors that are potential causes of this variance is also presented.
Zhengyang Liu 0005, Malathi Veeraraghavan, Chris Tracy, Jing Tie, Ian T. Foster, John M. Dennis, Jason Hick, Yee-Ting Li
SC7
2011 Performance of the community earth system model
abstract
The Community Earth System Model (CESM), released in June 2010, incorporates new physical process and new numerical algorithm options, significantly enhancing simulation capabilities over its predecessor, the June 2004 release of the Community Climate System Model. CESM also includes enhanced performance tuning options and performance portability capabilities. This paper describes performance and performance scaling on both the Cray XT5 and the IBM BG/P for four representative production simulations, varying both problem size and enabled physical processes. The paper also describes preliminary performance results for high resolution simulations using over 200,000 processor cores, indicating the promise of ongoing work in numerical algorithms and where further work is required.
Patrick H. Worley, Arthur A. Mirin, Anthony P. Craig, Mark A. Taylor, John M. Dennis, Mariana Vertenstein
SC5
2007 Inverse Space-Filling Curve Partitioning of a Global Ocean Model
abstract
In this paper, we describe how inverse space-filling curve partitioning is used to increase the simulation rate of a global ocean model. Space-filling curve partitioning allows for the elimination of load imbalance in the computational grid due to land points. Improved load balance combined with code modifications within the conjugate gradient solver significantly increase the simulation rate of the parallel ocean program at high resolution. The simulation rate for a high resolution model nearly doubled from 4.0 to 7.9 simulated years per day on 28,972 IBM Blue Gene/L processors. We also demonstrate that our techniques increase the simulation rate on 7545 Cray XT3 processors from 6.3 to 8.1 simulated years per day. Our results demonstrate how minor code modifications can have significant impact on resulting performance for very large processor counts.
John M. Dennis
IPDPS1
2001 Terascale spectral element dynamical core for atmospheric general circulation models
abstract
Climate modeling is a grand challenge problem where scientific progress is measured not in terms of the largest problem that can be solved but by the highest achievable integration rate. These models have been notably absent in previous Gordon Bell competitions due to their inability to scale to large processor counts. A scalable and efficient spectral element atmospheric model is presented. A new semi-implicit time stepping scheme accelerates the integration rate relative to an explicit model by a factor of two, achieving 130 years per day at T63L30 equivalent resolution. Execution rates are reported for the standard shallow water and Held-Suarez climate benchmarks on IBM SP clusters. The explicit T170 equivalent multi-layer shallow water model sustains 343 Gflops at NERSC, 206 Gflops at NPACI (SDSC) and 127 Gflops at NCAR. An explicit Held-Suarez integration sustains 369 Gflops on 128 16-way IBM nodes at NERSC.
Richard D. Loft, Stephen J. Thomas, John M. Dennis
SC3
1995 Implementation and Performance Issues of a Massively Parallel Atmospheric Model
Steven W. Hammond, Richard D. Loft, John M. Dennis, Richard K. Sato
Parallel Comput.3