Stéphane Ethier

dblp:62/3113 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
High-performance computing · 60% Parallel and multicore computing · 11% Storage systems · 11%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.762016
Extreme scale plasma turbulence simulations on top supercomputers worldwide · SC 2016
Kinetic turbulence simulations at extreme scale on leadership-class systems · SC 2013
Gyrokinetic toroidal simulations on leading multi- and manycore HPC systems · SC 2011
High-performance computing
performance optimization at scale
0.532016
Extreme scale plasma turbulence simulations on top supercomputers worldwide · SC 2016
Kinetic turbulence simulations at extreme scale on leadership-class systems · SC 2013
Gyrokinetic toroidal simulations on leading multi- and manycore HPC systems · SC 2011
Visualization and visual analytics › information visualization › statistical graphics
distribution visualization
0.312017
Scalable Visualization of Time-varying Multi-parameter Distributions Using Spatially Organized Histograms · IEEE Trans. Vis. Comput. Graph. 2017
Visualization and visual analytics
scientific visualization
0.312017
Scalable Visualization of Time-varying Multi-parameter Distributions Using Spatially Organized Histograms · IEEE Trans. Vis. Comput. Graph. 2017
Visualization and visual analytics › temporal data visualization
time-varying data visualization
0.312017
Scalable Visualization of Time-varying Multi-parameter Distributions Using Spatially Organized Histograms · IEEE Trans. Vis. Comput. Graph. 2017
High-performance computing
scientific data compression
0.322012
ISOBAR Preconditioner for Effective and High-throughput Lossless Data Compression · ICDE 2012
S-preconditioner for Multi-fold Data Reduction with Guaranteed User-Controlled Accuracy · ICDM 2011
High-performance computing › large-scale simulation
extreme-scale simulation
0.212016
Extreme scale plasma turbulence simulations on top supercomputers worldwide · SC 2016
Storage systems
data compression
0.222012
ISOBAR Preconditioner for Effective and High-throughput Lossless Data Compression · ICDE 2012
ISABELA-QA: query-driven analytics with ISABELA-compressed extreme-scale scientific data · SC 2011
High-performance computing › parallel i/o
i/o
0.112012
Byte-precision level of detail processing for variable precision analytics · SC 2012
Storage systems › data compression
lossless compression
0.112012
ISOBAR Preconditioner for Effective and High-throughput Lossless Data Compression · ICDE 2012
Parallel and multicore computing › parallel architecture
multicore optimization
0.112012
Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms · IEEE Trans. Parallel Distributed Syst. 2012
High-performance computing
scientific computing
0.112012
Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms · IEEE Trans. Parallel Distributed Syst. 2012
High-performance computing
scientific data analysis
0.112012
Byte-precision level of detail processing for variable precision analytics · SC 2012
Performance modeling and evaluation
benchmarking
0.132005
Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005
Scientific Computations on Modern Parallel Vector Systems · SC 2004
Evaluation of Cache-based Superscalar and Cacheless Vector Architectures for Scientific Computations · SC 2003
Performance modeling and evaluation › parallel performance evaluation
scientific application performance
0.132005
Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005
Scientific Computations on Modern Parallel Vector Systems · SC 2004
Evaluation of Cache-based Superscalar and Cacheless Vector Architectures for Scientific Computations · SC 2003
Processor architecture and microarchitecture
vector processor
0.132005
Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005
Scientific Computations on Modern Parallel Vector Systems · SC 2004
Evaluation of Cache-based Superscalar and Cacheless Vector Architectures for Scientific Computations · SC 2003
Distributed systems
communication optimization
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
Storage systems
data reduction
0.112011
S-preconditioner for Multi-fold Data Reduction with Guaranteed User-Controlled Accuracy · ICDM 2011
High-performance computing › lossy compression
error-bounded lossy compression
0.112011
S-preconditioner for Multi-fold Data Reduction with Guaranteed User-Controlled Accuracy · ICDM 2011
Parallel and multicore computing › parallel computing › parallel communication
one-sided communication
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
Parallel and multicore computing
parallel programming models
0.112011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011
High-performance computing
scientific data management
0.112011
ISABELA-QA: query-driven analytics with ISABELA-compressed extreme-scale scientific data · SC 2011
Parallel and multicore computing
parallel algorithms
0.112009
Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009
High-performance computing › scientific computing systems
particle-in-cell simulation
0.112009
Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009
High-performance computing › scientific visualization
in situ visualization
0.112017
Scalable Visualization of Time-varying Multi-parameter Distributions Using Spatially Organized Histograms · IEEE Trans. Vis. Comput. Graph. 2017
Computational science and engineering
plasma physics
0.012013
Kinetic turbulence simulations at extreme scale on leadership-class systems · SC 2013
Distributed systems
concurrent data transfer
0.012003
Grid -Based Parallel Data Streaming implemented for the Gyrokinetic Toroidal Code · SC 2003
Distributed systems
data streaming
0.012003
Grid -Based Parallel Data Streaming implemented for the Gyrokinetic Toroidal Code · SC 2003
Distributed systems
grid computing
0.012003
Grid -Based Parallel Data Streaming implemented for the Gyrokinetic Toroidal Code · SC 2003
Computational science and engineering
computational physics
0.012011
Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms · SC 2011

Methods — techniques the papers use, named apart from their topics

particle-in-cell · 0.3multithreading · 0.3domain decomposition · 0.3isosurfaces · 0.3iso surfaces · 0.3in-situ processing · 0.3in situ processing · 0.3performance modeling · 0.2one-sided communication · 0.2MPI 3.0 · 0.2k-means · 0.1fourier analysis · 0.1wavelet transform · 0.1preconditioning · 0.1one-sided CAF · 0.1PGAS · 0.1OpenMP · 0.1MPI · 0.1
YearPublicationVenuePosition
2019 Scalable Performance Awareness for In Situ Scientific Applications
abstract
Part of the promise of exascale computing and the next generation of scientific simulation codes is the ability to bring together time and spatial scales that have traditionally been treated separately. This enables creating complex coupled simulations and in situ analysis pipelines, encompassing such things as "whole device" fusion models or the simulation of cities from sewers to rooftops. Unfortunately, the HPC analysis tools that have been built up over the preceding decades are ill suited to the debugging and performance analysis of such computational ensembles. In this paper, we present a new vision for performance measurement and understanding of HPC codes, MonitoringAnalytics (MONA). MONA is designed to be a flexible, high performance monitoring infrastructure that can perform monitoring analysis in place or in transit by embedding analytics and characterization directly into the data stream, without relying upon delivering all monitoring information to a central database for post-processing. It addresses the trade-offs between the prohibitively expensive capture of all performance characteristics and not capturing enough to detect the features of interest. We demonstrate several uses of MONA; capturing and indexing multi-executable performance profiles to enable later processing, extraction of performance primitives to enable the generation of customizable benchmarks and performance skeletons, and extracting communication and application behaviors to enable better control and placement for the current and future runs of the science ensemble. Relevant performance information based on a system for MONA built from ADIOS and SOSflow technologies is provided for DOE science applications and leadership machines.
Matthew Wolf, Julien Dominski, Gabriele Merlo, Jong Choi 0001, Greg Eisenhauer, Stéphane Ethier, Kevin A. Huck, Scott Klasky, Jeremy Logan, Allen D. Malony, Chad Wood
eScience6
2018 OpenACC vs the Native Programming on Sunway TaihuLight: A Case Study with GTC-P
abstract
Sunway TaihuLight is China's recent top-ranked supercomputer worldwide that was the first to be built entirely with home-grown processors. This supercomputer can be programmed with two approaches: directive-based OpenACC and native programming. These approaches are studied here using GTC-P, a particle-in-cell code for investigating micro-turbulence in magnetic fusion plasmas. We have compared the performance and programming efforts between the OpenACC and the native version of GTC-P. Associated results show that in the OpenACC version, the kernel with irregular memory access becomes the main performance bottleneck due to poor data locality. To address this issue, we have applied two optimizations on the native version: (1) register level communication (RLC); and (2) an "asynchronization" strategy. With these two optimizations, the native version can achieve up to 2.5X speedup for the memory-bound kernel compared with the OpenACC version. In addition, we have now scaled GTC-P on 4,259,840 cores of TaihuLight and demonstrate performance comparisons with several world-leading supercomputers.
Linjin Cai, Yichao Wang 0001, William Tang 0002, Bei Wang 0002, Stéphane Ethier, James Lin 0001
CLUSTER5
2018 Coupling Exascale Multiphysics Applications: Methods and Lessons Learned
abstract
With the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration.
Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood
eScience14
2017 Scalable Visualization of Time-varying Multi-parameter Distributions Using Spatially Organized Histograms
abstract
Visualizing distributions from data samples as well as spatial and temporal trends of multiple variables is fundamental to analyzing the output of today's scientific simulations. However, traditional visualization techniques are often subject to a trade-off between visual clutter and loss of detail, especially in a large-scale setting. In this work, we extend the use of spatially organized histograms into a sophisticated visualization system that can more effectively study trends between multiple variables throughout a spatial domain. Furthermore, we exploit the use of isosurfaces to visualize time-varying trends found within histogram distributions. This technique is adapted into both an on-the-fly scheme as well as an in situ scheme to maintain real-time interactivity at a variety of data scales.
Tyson Neuroth, Franz Sauer, Weixing Wang 0004, Stéphane Ethier, Choong-Seock Chang, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.4
2016 Performance and Portability Studies with OpenACC Accelerated Version of GTC-P
abstract
Accelerator-based heterogeneous computing is of paramount importance to High Performance Computing. The increasing complexity of the cluster architectures requires more generic, high-level programming models. OpenACC is a directive-based parallel programming model, which provides performance on and portability across a wide variety of platforms, including GPU, multicore CPU, and many-core processors. GTC-P is a discovery-science-capable real-world application code based on the Particle-In-Cell (PIC) algorithm that is well-established in the HPC area. Several native versions of GTC-P have been developed for supercomputers on TOP500 with different architectures, including Titan, Mira, etc. Motivated by the state-of-art portability, we implemented the first OpenACC version of GTC-P and evaluated its performance portability across NVIDIA GPUs, Intel x86 and OpenPOWER CPUs. In this paper, we also proposed two key optimization methods for OpenACC implementation of PIC algorithm on multicore CPU and GPU including removing atomic operation and taking advantage of shared memory. OpenACC shows both impressive productivity and performance in a perspective of portability and scalability. The OpenACC version achieves more than 90% performance compared with the native versions with only about 300 LOC.
Yueming Wei, Yichao Wang 0001, Linjin Cai, William Tang 0002, Bei Wang 0002, Stéphane Ethier, Simon See, James Lin 0001
PDCAT6
2016 Extreme scale plasma turbulence simulations on top supercomputers worldwide
abstract
The goal of the extreme scale plasma turbulence studies described in this paper is to expedite the delivery of reliable predictions on confinement physics in large magnetic fusion systems by using world-class supercomputers to carry out simulations with unprecedented resolution and temporal duration. This has involved architecture-dependent optimizations of performance scaling and addressing code portability and energy issues, with the metrics for multi-platform comparisons being “time-to-solution” and “energy-to-solution”. Realistic results addressing how confinement losses caused by plasma turbulence scale from present-day devices to the much larger $25 billion international ITER fusion facility have been enabled by innovative advances in the GTC-P code including (i) implementation of one-sided communication from MPI 3.0 standard; (ii) creative optimization techniques on Xeon Phi processors; and (iii) development of a novel performance model for the key kernels of the PIC code. Results show that modeling data movement is sufficient to predict performance on modern supercomputer platforms.
William Tang 0002, Bei Wang 0002, Stéphane Ethier, Grzegorz Kwasniewski, Torsten Hoefler, Khaled Z. Ibrahim, Kamesh Madduri, Samuel Williams 0001, Leonid Oliker, Carlos Rosales-Fernandez, Timothy J. Williams
SC3
2014 Scibox: Online Sharing of Scientific Data via the Cloud
abstract
Collaborative science demands global sharing of scientific data. But it cannot leverage universally accessible cloud-based infrastructures like Drop Box, as those offer limited interfaces and inadequate levels of access bandwidth. We present the Scibox cloud facility for online sharing scientific data. It uses standard cloud storage solutions, but offers a usage model in which high end codes can write/read data to/from the cloud via the APIs they already use for their I/O actions. With Scibox, data upload/download volumes are controlled via Data Reduction-functions stated by end users and applied at the data source, before data is moved, with further gains in efficiency obtained by combining DR-functions to move exactly what is needed by current data consumers. We evaluate Scibox with science applications and their representative data analytics - the GTS fusion and the combustion image processing - demonstrating the potential for ubiquitous data access with substantial reductions in network traffic.
Jian Huang 0006, Xuechen Zhang 0001, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Stéphane Ethier, Scott Klasky
IPDPS6
2013 Kinetic turbulence simulations at extreme scale on leadership-class systems
abstract
Reliable predictive simulation capability addressing confinement properties in magnetically confined fusion plasmas is critically-important for ITER, a 20 billion dollar international burning plasma device under construction in France. The complex study of kinetic turbulence, which can severely limit the energy confinement and impact the economic viability of fusion systems, requires simulations at extreme scale for such an unprecedented device size. Our newly optimized, global, ab initio particle-in-cell code solving the nonlinear equations underlying gyrokinetic theory achieves excellent performance with respect to "time to solution" at the full capacity of the IBM Blue Gene/Q on 786,432 cores of Mira at ALCF and recently of the 1,572,864 cores of Sequoia at LLNL. Recent multithreading and domain decomposition optimizations in the new GTC-P code represent critically important software advances for modern, low memory per core systems by enabling routine simulations at unprecedented size (130 million grid points ITER-scale) and resolution (65 billion particles).
Bei Wang 0002, Stéphane Ethier, William Tang 0002, Timothy J. Williams, Khaled Z. Ibrahim, Kamesh Madduri, Samuel Williams 0001, Leonid Oliker
SC2
2013 ISABELA for effective in situ compression of scientific data
abstract
SUMMARY Exploding dataset sizes from extreme‐scale scientific simulations necessitates efficient data management and reduction schemes to mitigate I/O costs. With the discrepancy between I/O bandwidth and computational power, scientists are forced to capture data infrequently, thereby making data collection an inherently lossy process. Although data compression can be an effective solution, the random nature of real‐valued scientific datasets renders lossless compression routines ineffective. These techniques also impose significant overhead during decompression, making them unsuitable for data analysis and visualization, which require repeated data access. To address this problem, we propose an effective method for In situ Sort‐And‐B‐spline Error‐bounded Lossy Abatement (ISABELA) of scientific data that is widely regarded as effectively incompressible. With ISABELA, we apply a pre‐conditioner to seemingly random and noisy data along spatial resolution to achieve an accurate fitting model that guarantees a ⩾0.99 correlation with the original data. We further take advantage of temporal patterns in scientific data to compress data by ≈ 85%, while introducing only a negligible overhead on simulations in terms of runtime. ISABELA significantly outperforms existing lossy compression methods, such as wavelet compression, in terms of data reduction and accuracy. We extend upon our previous paper by additionally building a communication‐free, scalable parallel storage framework on top of ISABELA‐compressed data that is ideally suited for extreme‐scale analytical processing. The basis for our storage framework is an inherently local decompression method (it need not decode the entire data), which allows for random access decompression and low‐overhead task division that can be exploited over heterogeneous architectures. Furthermore, analytical operations such as correlation and query processing run quickly and accurately over data in the compressed space. Copyright © 2012 John Wiley & Sons, Ltd.
Sriram Lakshminarasimhan, Neil Shah, Stéphane Ethier, Seung-Hoe Ku, Choong-Seock Chang, Scott Klasky, Robert Latham, Robert B. Ross, Nagiza F. Samatova
Concurr. Comput. Pract. Exp.3
2012 Analytics-Driven Lossless Data Compression for Rapid In-situ Indexing, Storing, and Querying
John Jenkins, Isha Arkatkar, Sriram Lakshminarasimhan, Neil Shah, Eric R. Schendel, Stéphane Ethier, Choong-Seock Chang, Jacqueline Chen, Hemanth Kolla, Scott Klasky, Robert B. Ross, Nagiza F. Samatova
DEXA (2)6
2012 ISOBAR Preconditioner for Effective and High-throughput Lossless Data Compression
abstract
Efficient handling of large volumes of data is a necessity for exascale scientific applications and database systems. To address the growing imbalance between the amount of available storage and the amount of data being produced by high speed (FLOPS) processors on the system, data must be compressed to reduce the total amount of data placed on the file systems. General-purpose loss less compression frameworks, such as zlib and bzlib2, are commonly used on datasets requiring loss less compression. Quite often, however, many scientific data sets compress poorly, referred to as hard-to-compress datasets, due to the negative impact of highly entropic content represented within the data. An important problem in better loss less data compression is to identify the hard-to-compress information and subsequently optimize the compression techniques at the byte-level. To address this challenge, we introduce the In-Situ Orthogonal Byte Aggregate Reduction Compression (ISOBAR-compress) methodology as a preconditioner of loss less compression to identify and optimize the compression efficiency and throughput of hard-to-compress datasets.
Eric R. Schendel, Neil Shah, Jackie Chen, Choong-Seock Chang, Seung-Hoe Ku, Stéphane Ethier, Scott Klasky, Robert Latham, Robert B. Ross, Nagiza F. Samatova
ICDE7
2012 MLOC: Multi-level Layout Optimization Framework for Compressed Scientific Data Exploration with Heterogeneous Access Patterns
abstract
The size and scope of cutting-edge scientific simulations are growing much faster than the I/O and storage capabilities of their runtime environments. The growing gap gets exacerbated by exploratory dataâ"intensive analytics, such as querying simulation data for regions of interest with multivariate, spatio-temporal constraints. Query-driven data exploration induces heterogeneous access patterns that further stress the performance of the underlying storage system. To partially alleviate the problem, data reduction via compression and multi-resolution data extraction are becoming an integral part of I/O systems. While addressing the data size issue, these techniques introduce yet another mix of access patterns to a heterogeneous set of possibilities. Moreover, how extreme-scale datasets are partitioned into multiple files and organized on a parallel file systems augments to an already combinatorial space of possible access patterns. To address this challenge, we present MLOC, a parallel Multilevel Layout Optimization framework for Compressed scientific spatio-temporal data at extreme scale. MLOC proposes multiple fine-grained data layout optimization kernels that form a generic core from which a broader constellation of such kernels can be organically consolidated to enable an effective data exploration with various combinations of access patterns. Specifically, the kernels are optimized for access patterns induced by (a) queryâ"driven multivariate, spatio-temporal constraints, (b) precisionâ"driven data analytics, (c) compressionâ"driven data reduction, (d) multi-resolution data sampling, and (e) multiâ"file data partitioning and organization on a parallel file system. MLOC organizes these optimization kernels within a multiâ"level architecture, on which all the levels can be flexibly re-ordered by userâ"defined priorities. When tested on queryâ"driven exploration of compressed data, MLOC demonstrates a superior performance compared to any state-of-the-art scientific database management technologies.
Zhenhuan Gong, Terry Rogers, John Jenkins, Hemanth Kolla, Stéphane Ethier, Jackie Chen, Robert B. Ross, Scott Klasky, Nagiza F. Samatova
ICPP5
2012 Multi-level Layout Optimization for Efficient Spatio-temporal Queries on ISABELA-compressed Data
abstract
The size and scope of cutting-edge scientific simulations are growing much faster than the I/O subsystems of their runtime environments, not only making I/O the primary bottleneck, but also consuming space that pushes the storage capacities of many computing facilities. These problems are exacerbated by the need to perform data-intensive analytics applications, such as querying the dataset by variable and spatio-temporal constraints, for what current database technologies commonly build query indices of size greater than that of the raw data. To help solve these problems, we present a parallel query-processing engine that can handle both range queries and queries with spatio-temporal constraints, on B-spline compressed data with user-controlled accuracy. Our method adapts to widening gaps between computation and I/O performance by querying on compressed metadata separated into bins by variable values, utilizing Hilbert space-filling curves to optimize for spatial constraints and aggregating data access to improve locality of per-bin stored data, reducing the false positive rate and latency bound I/O operations (such as seek) substantially. We show our method to be efficient with respect to storage, computation, and I/O compared to existing database technologies optimized for query processing on scientific data.
Zhenhuan Gong, Sriram Lakshminarasimhan, John Jenkins, Hemanth Kolla, Stéphane Ethier, Jackie Chen, Robert B. Ross, Scott Klasky, Nagiza F. Samatova
IPDPS5
2012 Byte-precision level of detail processing for variable precision analytics
abstract
I/O bottlenecks in HPC applications are becoming a more pressing problem as compute capabilities continue to outpace I/O capabilities. While double-precision simulation data often must be stored losslessly, the loss of some of the fractional component may introduce acceptably small errors to many types of scientific analyses. Given this observation, we develop a precision level of detail (APLOD) library, which partitions double-precision datasets along user-defined byte boundaries. APLOD parameterizes the analysis accuracy-I/O performance tradeoff, bounds maximum relative error, maintains I/O access patterns compared to full precision, and operates with low overhead. Using ADIOS as an I/O use-case, we show proportional reduction in disk access time to the degree of precision. Finally, we show the effects of partial precision analysis on accuracy for operations such as k-means and Fourier analysis, finding a strong applicability for the use of varying degrees of precision to reduce the cost of analyzing extreme-scale data.
John Jenkins, Eric R. Schendel, Sriram Lakshminarasimhan, David A. Boyuka II, Terry Rogers, Stéphane Ethier, Robert B. Ross, Scott Klasky, Nagiza F. Samatova
SC6
2012 Optimization of Parallel Particle-to-Grid Interpolation on Leading Multicore Platforms
abstract
We are now in the multicore revolution which is witnessing a rapid evolution of architectural designs due to power constraints and correspondingly limited microprocessor clock speeds. Understanding how to efficiently utilize these systems in the context of demanding numerical algorithms is an urgent challenge to meet the ever growing computational needs of high-end computing. In this work, we examine multicore parallel optimization of the particle-to-grid interpolation step in particle-mesh methods, an inherently complex optimization problem due to its low computation intensity, irregular data accesses, and potential fine-grained data hazards. Our evaluated kernels are derived from two important numerical computations: a biological simulation of the heart using the Immersed Boundary (IB) method, and a Gyrokinetic Particle-in-Cell (PIC)-based application for studying fusion plasma microturbulence. We develop several novel synchronization and grid decomposition schemes, as well as low-level optimization techniques to maximize performance on three modern multicore platforms: Intel's Xeon X5550 (Nehalem), AMD's Opteron 2356 (Barcelona), and Sun's UltraSparc T2+ (Niagara). Results show that our optimizations lead to significant performance improvements, achieving up to a 5.6× speedup compared to the reference parallel implementation. Our work also provides valuable insight into the design of future autotuning frameworks for particle-to-grid interpolation on next-generation systems.
Kamesh Madduri, Jimmy Su, Samuel Williams 0001, Leonid Oliker, Stéphane Ethier, Katherine A. Yelick
IEEE Trans. Parallel Distributed Syst.5
2011 Compressing the Incompressible with ISABELA: In-situ Reduction of Spatio-temporal Data
Sriram Lakshminarasimhan, Neil Shah, Stéphane Ethier, Scott Klasky, Robert Latham, Robert B. Ross, Nagiza F. Samatova
Euro-Par (1)3
2011 S-preconditioner for Multi-fold Data Reduction with Guaranteed User-Controlled Accuracy
abstract
The growing gap between the massive amounts of data generated by petascale scientific simulation codes and the capability of system hardware and software to effectively analyze this data necessitates data reduction. Yet, the increasing data complexity challenges most, if not all, of the existing data compression methods. In fact, lossless compression techniques offer no more than 10% reduction on scientific data that we have experience with, which is widely regarded as effectively incompressible. To bridge this gap, in this paper, we advocate a transformative strategy that enables fast, accurate, and multi-fold reduction of double-precision floating-point scientific data. The intuition behind our method is inspired by an effective use of preconditioners for linear algebra solvers optimized for a particular class of computational "dwarfs" (e.g., dense or sparse matrices). Focusing on a commonly used multi-resolution wavelet compression technique as the underlying "solver" for data reduction we propose the S-preconditioner, which transforms scientific data into a form with high global regularity to ensure a significant decrease in the number of wavelet coefficients stored for a segment of data. Combined with the subsequent EQ-calibrator, our resultant method (called S-Preconditioned EQ-Calibrated Wavelets (SPEQC-Wavelets)), robustly achieved a 4- to 5-fold data reduction-while guaranteeing user-defined accuracy of reconstructed data to be within 1% point-by-point relative error, lower than 0.01 Normalized RMSE, and higher than 0.99 Pearson Correlation. In this paper, we show the results we obtained by testing our method on six petascale simulation codes including fusion, combustion, climate, astrophysics, and subsurface groundwater in addition to 13 publicly available scientific datasets. We also demonstrate that application-driven data mining tasks performed on decompressed variables or their derived quantities produce results of comparable quality with the ones for the original data.
Sriram Lakshminarasimhan, Neil Shah, Zhenhuan Gong, Choong-Seock Chang, Jackie Chen, Stéphane Ethier, Hemanth Kolla, Seung-Hoe Ku, Scott Klasky, Robert Latham, Robert B. Ross, Karen Schuchardt, Nagiza F. Samatova
ICDM7
2011 ISABELA-QA: query-driven analytics with ISABELA-compressed extreme-scale scientific data
abstract
Efficient analytics of scientific data from extreme-scale simulations is quickly becoming a top-notch priority. The increasing simulation output data sizes demand for a paradigm shift in how analytics is conducted. In this paper, we argue that query-driven analytics over compressed---rather than original, full-size---data is a promising strategy in order to meet storage-and-I/O-bound application challenges. As a proof-of-principle, we propose a parallel query processing engine, called ISABELA-QA that is designed and optimized for knowledge priors driven analytical processing of spatio-temporal, multivariate scientific data that is initially compressed, in situ, by our ISABELA technology. With ISABELA-QA, the total data storage requirement is less than 23%-30% of the original data, which is upto eight-fold less than what the existing state-of-the-art data management technologies that require storing both the original data and the index could offer. Since ISABELA-QA operates on the metadata generated by our compression technology, its underlying indexing technology for efficient query processing is light-weight; it requires less than 3% of the original data, unlike existing database indexing approaches that require 30%-300% of the original data. Moreover, ISABELA-QA is specifically optimized to retrieve the actual values rather than spatial regions for the variables that satisfy user-specified range queries---a functionality that is critical for high-accuracy data analytics. To the best of our knowledge, this is the first techology that enables query-driven analytics over the compressed spatio-temporal floating-point double-or single-precision data, while offering a light-weight memory and disk storage footprint solution with parallel, scalable, multi-node, multi-core, GPU-based query processing.
Sriram Lakshminarasimhan, John Jenkins, Isha Arkatkar, Zhenhuan Gong, Hemanth Kolla, Seung-Hoe Ku, Stéphane Ethier, Jackie Chen, Choong-Seock Chang, Scott Klasky, Robert Latham, Robert B. Ross, Nagiza F. Samatova
SC7
2011 Gyrokinetic toroidal simulations on leading multi- and manycore HPC systems
abstract
The gyrokinetic Particle-in-Cell (PIC) method is a critical computational tool enabling petascale fusion simulation research. In this work, we present novel multi- and manycore-centric optimizations to enhance performance of GTC, a PIC-based production code for studying plasma microturbulence in tokamak devices. Our optimizations encompass all six GTC sub-routines and include multi-level particle and grid decompositions designed to improve multi-node parallel scaling, particle binning for improved load balance, GPU acceleration of key subroutines, and memory-centric optimizations to improve single-node scaling and reduce memory utilization. The new hybrid MPI-OpenMP and MPI-OpenMP-CUDA GTC versions achieve up to a 2x speedup over the production Fortran code on four parallel systems --- clusters based on the AMD Magny-Cours, Intel Nehalem-EP, IBM BlueGene/P, and NVIDIA Fermi architectures. Finally, strong scaling experiments provide insight into parallel scalability, memory utilization, and programmability trade-offs for large-scale gyrokinetic PIC simulations, while attaining a 1.6× speedup on 49,152 XE6 cores.
Kamesh Madduri, Khaled Z. Ibrahim, Samuel Williams 0001, Eun-Jin Im, Stéphane Ethier, John Shalf, Leonid Oliker
SC5
2011 Multithreaded global address space communication techniques for gyrokinetic fusion applications on ultra-scale platforms
abstract
We present novel parallel language constructs for the communication intensive part of a magnetic fusion simulation code. The focus of this work is the shift phase of charged particles of a tokamak simulation code in toroidal geometry. We introduce new hybrid PGAS/OpenMP implementations of highly optimized hybrid MPI/OpenMP based communication kernels. The hybrid PGAS implementations use an extension of standard hybrid programming techniques, enabling the distribution of high communication work loads of the underlying kernel among OpenMP threads. Building upon lightweight one-sided CAF (Fortran 2008) communication techniques, we also show the benefits of spreading out the communication over a longer period of time, resulting in a reduction of bandwidth requirements and a more sustained communication and computation overlap. Experiments on up to 130560 processors are conducted on the NERSC Hopper system, which is currently the largest HPC platform with hardware support for one-sided communication and show performance improvements of 52 % at highest concurrency.
Robert Preissl, Nathan Wichmann, Bill Long, John Shalf, Stéphane Ethier, Alice E. Koniges
SC5
2011 Gyrokinetic particle-in-cell optimization on emerging multi- and manycore platforms
Kamesh Madduri, Eun-Jin Im, Khaled Z. Ibrahim, Samuel Williams 0001, Stéphane Ethier, Leonid Oliker
Parallel Comput.5
2009 Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors
abstract
We present multicore parallelization strategies for the particle-to-grid interpolation step in the Gyrokinetic Toroidal Code (GTC), a 3D particle-in-cell (PIC) application to study turbulent transport in magnetic-confinement fusion devices. Particle-grid interpolation is a known performance bottleneck in several PIC applications. In GTC, this step involves particles depositing charges to a 3D toroidal mesh, and multiple particles may contribute to the charge at a grid point. We design new parallel algorithms for the GTC charge deposition kernel, and analyze their performance on three leading multicore platforms. We implement thirteen different variants for this kernel and identify the best-performing ones given typical PIC parameters such as the grid size, number of particles per cell, and the GTC-specific particle Larmor radius variation. We find that our best strategies can be 2x faster than the reference optimized MPI implementation, and our analysis provides insight into desirable architectural features for high-performance PIC simulation codes.
Kamesh Madduri, Samuel Williams 0001, Stéphane Ethier, Leonid Oliker, John Shalf, Erich Strohmaier, Katherine A. Yelick
SC3
2007 Scientific Application Performance on Candidate PetaScale Platforms
abstract
After a decade where HEC (high-end computing) capability was dominated by the rapid pace of improvements to CPU clock frequency, the performance of next-generation supercomputers is increasingly differentiated by varying interconnect designs and levels of integration. Understanding the tradeoffs of these system designs, in the context of high-end numerical simulations, is a key step towards making effective petascale computing a reality. This work represents one of the most comprehensive performance evaluation studies to date on modern NEC systems, including the IBM Power5, AMD Opteron, IBM BG/L, and Cray X1E. A novel aspect of our study is the emphasis on full applications, with real input data at the scale desired by computational scientists in their unique domain. We examine six candidate ultra-scale applications, representing a broad range of algorithms and computational structures. Our work includes the highest concurrency experiments to date on five of our six applications, including 32K processor scalability for two of our codes and describe several successful optimizations strategies on BG/L, as well as improved X1E vectorization. Overall results indicate that our evaluated codes have the potential to effectively utilize petascale resources; however, several applications would require reengineering to incorporate the additional levels of parallelism necessary to achieve the vast concurrency of upcoming ultra-scale systems.
Leonid Oliker, Andrew Canning, Jonathan Carter 0002, Costin Iancu, Michael Lijewski, Shoaib Kamil 0001, John Shalf, Hongzhang Shan, Erich Strohmaier, Stéphane Ethier, Tom Goodale
IPDPS10
2005 Leading Computational Methods on Scalar and Vector HEC Platforms
abstract
The last decade has witnessed a rapid proliferation of superscalar cache-based microprocessors to build high-end computing (HEC) platforms, primarily because of their generality, scalability, and cost effectiveness. However, the growing gap between sustained and peak performance for full-scale scientific applications on conventional supercomputers has become a major concern in high performance computing, requiring significantly larger systems and application scalability than implied by peak performance in order to achieve desired performance. The latest generation of custom-built parallel vector systems have the potential to address this issue for numerical algorithms with sufficient regularity in their computational structure. In this work we explore applications drawn from four areas: atmospheric modeling (CAM), magnetic fusion (GTC), plasma physics (LBMHD3D), and material science (PARATEC). We compare performance of the vector-based Cray X1, Earth Simulator, and newly-released NEC SX-8 and Cray X1E, with performance of three leading commodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors. Our work makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new data-decomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES promodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors, with modern parallel vector systems: the Cray X1, Earth Simulator (ES), and the NEC SX-8. Additionally, we examine performance of CAM on the recently-released Cray X1E. Our research team was the first international group to conduct a performance evaluation study at the Earth Simulator Center; remote ES access is not available. Our work builds on our previous efforts [16, 17] and makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new datadecomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES processors; and the largest PARATEC cell size atomistic simulation to date. Overall, results show that the vector architectures attain unprecedented aggregate performance across our application suite, demonstrating the tremendous potential of modern parallel vector systems.
Leonid Oliker, Jonathan Carter 0002, Michael F. Wehner, Andrew Canning, Stéphane Ethier, Arthur A. Mirin, David Parks, Patrick H. Worley, Shigemune Kitawaki, Yoshinori Tsuda
SC5
2005 Performance evaluation of the SX-6 vector architecture for scientific computations
abstract
The growing gap between sustained and peak performance for scientific applications is a well-known problem in high-performance computing. The recent development of parallel vector systems offers the potential to reduce this gap for many computational science codes and deliver a substantial increase in computing capabilities. This paper examines the intranode performance of the NEC SX-6 vector processor, and compares it against the cache-based IBM Power3 and Power4 superscalar architectures, across a number of key scientific computing areas. First, we present the performance of a microbenchmark suite that examines many low-level machine characteristics. Next, we study the behavior of the NAS Parallel Benchmarks. Finally, we evaluate the performance of several scientific computing codes. Overall results demonstrate that the SX-6 achieves high performance on a large fraction of our application suite and often significantly outperforms the cache-based architectures. However, certain classes of applications are not easily amenable to vectorization and would require extensive algorithm and implementation reengineering to utilize the SX-6 effectively. Copyright © 2005 John Wiley & Sons, Ltd.
Leonid Oliker, Andrew Canning, Jonathan Carter 0002, John Shalf, David Skinner, Stéphane Ethier, Rupak Biswas, M. Jahed Djomehri, Rob F. Van der Wijngaart
Concurr. Pract. Exp.6
2004 Scientific Computations on Modern Parallel Vector Systems
abstract
Computational scientists have seen a frustrating trend of stagnating application performance despite dramatic increases in the claimed peak capability of high performance computing systems. This trend has been widely attributed to the use of superscalar-based commodity components who’s architectural designs offer a balance between memory performance, network capability, and execution rate that is poorly matched to the requirements of large-scale numerical computations. Recently, two innovative parallel-vector architectures have become operational: the Japanese Earth Simulator (ES) and the Cray X1. In order to quantify what these modern vector capabilities entail for the scientists that rely on modeling and simulation, it is critical to evaluate this architectural paradigm in the context of demanding computational algorithms. Our evaluation study examines four diverse scientific applications with the potential to run at ultrascale, from the areas of plasma physics, material science, astrophysics, and magnetic fusion. We compare performance between the vector-based ES and X1, with leading superscalar-based platforms: the IBM Power3/4 and the SGI Altix. Our research team was the first international group to conduct a performance evaluation study at the Earth Simulator Center; remote ES access in not available. Results demonstrate that the vector systems achieve excellent performance on our application suite - the highest of any architecture tested to date. However, vectorization of a particle-in-cell code highlights the potential difficulty of expressing irregularly structured algorithms as data-parallel programs.
Leonid Oliker, Andrew Canning, Jonathan Carter 0002, John Shalf, Stéphane Ethier
SC5
2004 Visualizing Gyrokinetic Simulations
abstract
The continuing advancement of plasma science is central to realizing fusion as an inexpensive and safe energy source. Gryokinetic simulations of plasmas are fundamental to the understanding of turbulent transport in fusion plasma. This work discusses the visualization challenges presented by gyrokinetic simulations using magnetic field line following coordinates, and presents an effective solution exploiting programmable graphics hardware to enable interactive volume visualization of 3D plasma flow on a toroidal coordinate system. The new visualization capability can help scientists better understand three-dimensional structures of the modeled phenomena. Both the limitations and future promise of the hardware-accelerated approach are also discussed.
David Crawford, Kwan-Liu Ma, Min-Yu Huang, Scott Klasky, Stéphane Ethier
IEEE Visualization5
2003 Grid -Based Parallel Data Streaming implemented for the Gyrokinetic Toroidal Code
abstract
We have developed a threaded parallel data streaming approach using Globus to transfer multi-terabyte simulation data from a remote supercomputer to the scientistýs home analysis/visualization cluster, as the simulation executes, with negligible overhead. Data transfer experiments show that this concurrent data transfer approach is more favorable compared with writing to local disk and then transferring this data to be post-processed. The present approach is conducive to using the grid to pipeline the simulation with post-processing and visualization. We have applied this method to the Gyrokinetic Toroidal Code (GTC), a 3-dimensional particle-in-cell code used to study micro-turbulence in magnetic confinement fusion from first principles plasma theory.
Scott Klasky, Stéphane Ethier, K. Martins, Douglas McCune, Ravi Samtaney
SC2
2003 Evaluation of Cache-based Superscalar and Cacheless Vector Architectures for Scientific Computations
abstract
The growing gap between sustained and peak performance for scientific applications is a well-known problem in high end computing. The recent development of parallel vector systems offers the potential to bridge this gap for many computational science codes and deliver a substantial increase in comput-ing capabilities. This paper examines the intranode performance of the NEC SX-6 vector processor and the cache-based IBM Power3/4 superscalar architectures across a number of scientific computing areas. First, we present the performance of a microbenchmark suite that examines low-level machine characteristics. Next, we study the behavior of the NAS Parallel Benchmarks. Finally, we evaluate the performance of several scientific computing codes. Results demonstrate that the SX-6 achieves high performance on a large fraction of our applications and often significantly outperforms the cache-based architectures. However, certain applications are not easily amenable to vectorization and would require extensive algorithm and implementation reengineering to utilize the SX-6 effectively.
Leonid Oliker, Andrew Canning, Jonathan Carter 0002, John Shalf, David Skinner, Stéphane Ethier, Rupak Biswas, M. Jahed Djomehri, Rob F. Van der Wijngaart
SC6