Jeremy S. Meredith

dblp:19/4329 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Performance modeling and evaluation · 46% High-performance computing · 40% Emerging computing paradigms · 8%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific visualization
in situ visualization
0.212016
Performance modeling of in situ rendering · SC 2016
Performance modeling and evaluation › workload characterization › parallel workload analysis
communication pattern analysis
0.212015
Automated Characterization of Parallel Application Communication Patterns · HPDC 2015
Performance modeling and evaluation
workload characterization
0.212015
Automated Characterization of Parallel Application Communication Patterns · HPDC 2015
Data mining › visualization
visual data mining
0.112008
High performance multivariate visual data exploration for extremely large data · SC 2008
High-performance computing
performance optimization at scale
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
Emerging computing paradigms › quantum computing › quantum simulation
quantum monte carlo simulation
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
High-performance computing
scientific computing systems
0.112008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
Performance modeling and evaluation › statistical analysis › statistical performance analysis
statistical performance modeling
0.112016
Performance modeling of in situ rendering · SC 2016
Parallel and multicore computing › MPI
MPI program analysis
0.112015
Automated Characterization of Parallel Application Communication Patterns · HPDC 2015
Computational science and engineering
computational physics
0.012008
New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors · SC 2008
High-performance computing
distributed memory systems
0.012008
High performance multivariate visual data exploration for extremely large data · SC 2008

Methods — techniques the papers use, named apart from their topics

statistical modeling · 0.2algorithmic complexity analysis · 0.2parallel coordinates · 0.2index/query technology · 0.2search-based analysis · 0.2mpip profiling · 0.2mixed single-/double precision · 0.2delayed monte carlo updates · 0.2
YearPublicationVenuePosition
2018 Aspen-based performance and energy modeling frameworks
Mariam Umar, Shirley V. Moore, Jeremy S. Meredith, Jeffrey S. Vetter, Kirk W. Cameron
J. Parallel Distributed Comput.3
2017 Durango: Scalable Synthetic Workload Generation for Extreme-Scale Application Performance Modeling and Simulation
abstract
Performance modeling of extreme-scale applications on accurate representations of potential architectures is critical for designing next generation supercomputing systems because it is impractical to construct prototype systems at scale with new network hardware in order to explore designs and policies. However, these simulations often rely on static application traces that can be difficult to work with because of their size and lack of flexibility to extend or scale up without rerunning the original application. To address this problem, we have created a new technique for generating scalable, flexible workloads from real applications, we have implemented a prototype, called Durango, that combines a proven analytical performance modeling language, Aspen, with the massively parallel HPC network modeling capabilities of the CODES framework.
Christopher D. Carothers, Jeremy S. Meredith, Mark P. Blanco, Jeffrey S. Vetter, Misbah Mubarak, Justin M. LaPre, Shirley Moore
SIGSIM-PADS2
2016 A Study of Power-Performance Modeling Using a Domain-Specific Language
abstract
Energy use is now a first-class design constraint in high-performance systems and applications. Improving our understanding of application energy consumption in diverse, heterogeneous systems will be essential to efficient operation. For example, power limits in large scale parallel and distributed systems will require optimizing performance under energy constraints. However, with increased levels of parallelism, complex memory hierarchies, hardware heterogeneity, and diverse programming models and interfaces, improving performance and energy efficiency simultaneously is exceedingly difficult. Our thesis is that estimating energy use, either a priori or as soon as possible at runtime, will be essential to future systems. Such estimates must adapt with changes in applications across hardware configurations. Existing approaches offer insight and detail, but typically are too cumbersome to enable adaptation at runtime or lack portability or accuracy. To overcome these limitations, we propose two energy estimation techniques which use the Aspen domain specific language for performance modeling: ACEE (Algorithmic and Categorical Energy Estimation), a combination of analytical and empirical modeling techniques embedded in a runtime framework that leverages Aspen, and AEEM (Aspen's Embedded Energy Modeling), a system level coarse-grained energy estimation technique that uses performance modeling from Aspen to generate energy estimations at runtime. This paper presents methodology of the models and examines their accuracy as well as their advantages and challenges in several use cases.
Mariam Umar, Jeremy S. Meredith, Jeffrey S. Vetter, Kirk W. Cameron
SBAC-PAD2
2016 Performance modeling of in situ rendering
abstract
With the push to exascale, in situ visualization and analysis will continue to play an important role in high performance computing. Tightly coupling in situ visualization with simulations constrains resources for both, and these constraints force a complex balance of trade-offs. A performance model that provides an a priori answer for the cost of using an in situ approach for a given task would assist in managing the trade-offs between simulation and visualization resources. In this work, we present new statistical performance models, based on algorithmic complexity, that accurately predict the run-time cost of a set of representative rendering algorithms, an essential in situ visualization task. To train and validate the models, we conduct a performance study of an MPI+X rendering infrastructure used in situ with three HPC simulation applications. We then explore feasibility issues using the model for selected in situ rendering questions.
Matthew Larsen, Cyrus Harrison, James Kress, David Pugmire, Jeremy S. Meredith, Hank Childs
SC5
2015 Ray tracing within a data parallel framework
abstract
Current architectural trends on supercomputers have dramatic increases in the number of cores and available computational power per die, but this power is increasingly difficult for programmers to harness effectively. High-level language constructs can simplify programming many-core devices, but this ease comes with a potential loss of processing power, particularly for cross-platform constructs. Recently, scientific visualization packages have embraced language constructs centering around data parallelism, with familiar operators such as map, reduce, gather, and scatter. Complete adoption of data parallelism will require that central visualization algorithms be revisited, and expressed in this new paradigm while preserving both functionality and performance. This investment has a large potential payoff: portable performance in software bases that can span over the many architectures that scientific visualization applications run on. With this work, we present a method for ray tracing consisting of entirely of data parallel primitives. Given the extreme computational power on nodes now prevalent on supercomputers, we believe that ray tracing can supplant rasterization as the work-horse graphics solution for scientific visualization. Our ray tracing method is relatively efficient, and we describe its performance with a series of tests, and also compare to leading-edge ray tracers that are optimized for specific platforms. We find that our data parallel approach leads to results that are acceptable for many scientific visualization use cases, with the key benefit of providing a single code base that can run on many architectures.
Matthew Larsen, Jeremy S. Meredith, Paul A. Navrátil, Hank Childs
PacificVis2
2015 Automated Characterization of Parallel Application Communication Patterns
abstract
A concise description of an application's communication pattern is often useful, for example, as an efficient way to communicate application behavior to a system vendor. Several existing performance analysis tools can capture aspects of an application's communication behavior such as which processes communicated with which others and the communications operations they used. However, a human with a high degree of expertise is still required to recognize and characterize common communication idioms within the performance data collected by those tools. To simplify this characterization for non-experts, we have developed an approach for automatically characterizing a MPI application's communication behavior. We use the mpiP profiling tool to collect information about an application's communication topology and message volume. We then use a post-mortem search-based analysis to compare the application's observed communication pattern against a library of common communication patterns. By comparing the result of the various search paths, our approach identifies the combination of patterns that best matches the observed behavior. To evaluate our approach, we applied it to a synthetic example communication matrix and communication matrices obtained from two scientific applications. We determined that our automated approach was highly effective in characterizing the communication patterns represented in the matrices.
Philip C. Roth, Jeremy S. Meredith, Jeffrey S. Vetter
HPDC2
2015 COMPASS: A Framework for Automated Performance Modeling and Prediction
abstract
Flexible, accurate performance predictions offer numerous benefits such as gaining insight into and optimizing applications and architectures. However, the development and evaluation of such performance predictions has been a major research challenge, due to the architectural complexities. To address this challenge, we have designed and implemented a prototype system, named COMPASS, for automated performance model generation and prediction. COMPASS generates a structured performance model from the target application's source code using automated static analysis, and then, it evaluates this model using various performance prediction techniques. As we demonstrate on several applications, the results of these predictions can be used for a variety of purposes, such as design space exploration, identifying performance tradeoffs for applications, and understanding sensitivities of important parameters. COMPASS can generate these predictions across several types of applications from traditional, sequential CPU applications to GPU-based, heterogeneous, parallel applications. Our empirical evaluation demonstrates a maximum overhead of 4%, flexibility to generate models for 9 applications, speed, ease of creation, and very low relative errors across a diverse set of architectures.
Seyong Lee, Jeremy S. Meredith, Jeffrey S. Vetter
ICS2
2014 Value influence analysis for message passing applications
abstract
People who develop, debug, and optimize applications are most effective when they understand how those applications function. Value influence tracking is an on-line code analysis approach that provides a data-centric perspective on how a value contributes to later computation. Early work on value influence tracking focused on single-process applications. Building upon this early work, we have designed support for performing value influence tracking analyses with applications that use common MPI point-to-point and collective communication operations. In this paper, we describe the design and implementation of an approach for propagating value influence data between the processes of an MPI application that uses these types of operations. To demonstrate and evaluate our approach, we present case studies of using our value influence tracking implementation with the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) and the Model for Prediction Across Scales (MPAS) ocean climate model running on the Keeneland Initial Delivery System (KIDS) Linux cluster. We also discuss how to extend our approach to support MPI one-sided operations and non-blocking collective communication operations.
Philip C. Roth, Jeremy S. Meredith
ICS2
2011 Quartile and Outlier Detection on Heterogeneous Clusters Using Distributed Radix Sort
abstract
In the past few years, performance improvements in CPUs and memory technologies have outpaced those of storage systems. When extrapolated to the exascale, this trend places strict limits on the amount of data that can be written to disk for full analysis, resulting in an increased reliance on characterizing in-memory data. Many of these characterizations are simple, but require sorted data. This paper explores an example of this type of characterization -- the identification of quartiles and statistical outliers -- and presents a performance analysis of a distributed heterogeneous radix sort as well as an assessment of current architectural bottlenecks.
Kyle Spafford, Jeremy S. Meredith, Jeffrey S. Vetter
CLUSTER2
2010 Maestro: Data Orchestration and Tuning for OpenCL Devices
Kyle Spafford, Jeremy S. Meredith, Jeffrey S. Vetter
Euro-Par (2)2
2010 Visualization and Analysis-Oriented Reconstruction of Material Interfaces
abstract
Abstract Reconstructing boundaries along material interfaces from volume fractions is a difficult problem, especially because the under‐resolved nature of the input data allows for many correct interpretations. Worse, algorithms widely accepted as appropriate for simulation are inappropriate for visualization. In this paper, we describe a new algorithm that is specifically intended for reconstructing material interfaces for visualization and analysis requirements. The algorithm performs well with respect to memory footprint and execution time, has desirable properties in various accuracy metrics, and also produces smooth surfaces with few artifacts, even when faced with more than two materials per cell.
Jeremy S. Meredith, Hank Childs
Comput. Graph. Forum1
2009 Accuracy and performance of graphics processors: A Quantum Monte Carlo application case study
Jeremy S. Meredith, Gonzalo Alvarez 0001, Thomas A. Maier, Thomas C. Schulthess, Jeffrey S. Vetter
Parallel Comput.1
2008 New algorithm to enable 400+ TFlop/s sustained performance in simulations of disorder effects in high-Tc superconductors
abstract
Staggering computational and algorithmic advances in recent years now make possible systematic Quantum Monte Carlo (QMC) simulations of high temperature (high-Tc) superconductivity in a microscopic model, the two dimensional (2D) Hubbard model, with parameters relevant to the cuprate materials. Here we report the algorithmic and computational advances that enable us to study the effect of disorder and nano-scale inhomogeneities on the pair-formation and the superconducting transition temperature necessary to understand real materials. The simulation code is written with a generic and extensible approach and is tuned to perform well at scale. Significant algorithmic improvements have been made to make effective use of current supercomputing architectures. By implementing delayed Monte Carlo updates and a mixed single-/double precision mode, we are able to dramatically increase the efficiency of the code. On the Cray XT4 systems of the Oak Ridge National Laboratory (ORNL), for example, we currently run production jobs on 31 thousand processors and thereby routinely achieve a sustained performance that exceeds 200 TFlop/s. On a system with 49 thousand processors we achieved a sustained performance of 409 TFlop/s. We present a study of how random disorder in the effective Coulomb interaction strength affects the superconducting transition temperature in the Hubbard model.
Gonzalo Alvarez 0001, Michael S. Summers, Don E. Maxwell, Markus Eisenbach 0002, Jeremy S. Meredith, Jeffrey M. Larkin, John M. Levesque, Thomas A. Maier, Paul R. C. Kent, Eduardo F. D'Azevedo, Thomas C. Schulthess
SC5
2008 High performance multivariate visual data exploration for extremely large data
abstract
One of the central challenges in modern science is the need to quickly derive knowledge and understanding from large, complex collections of data. We present a new approach that deals with this challenge by combining and extending techniques from high performance visual data analysis and scientific data management. This approach is demonstrated within the context of gaining insight from complex, time-varying datasets produced by a laser wakefield accelerator simulation. Our approach leverages histogram-based parallel coordinates for both visual information display as well as a vehicle for guiding a data mining operation. Data extraction and subsetting are implemented with state-of-the-art index/query technology. This approach, while applied here to accelerator science, is generally applicable to a broad set of science applications, and is implemented in a production-quality visual data analysis infrastructure. We conduct a detailed performance analysis and demonstrate good scalability on a distributed memory Cray XT4 system.
Oliver Rübel, Prabhat, Kesheng Wu, Hank Childs, Jeremy S. Meredith, Cameron G. R. Geddes, Estelle Cormier-Michel, Sean Ahern, Gunther H. Weber, Peter Messmer, Hans Hagen, Bernd Hamann, E. Wes Bethel
SC5
2007 Balancing productivity and performance on the cell broadband engine
abstract
The cell broadband engine (BE) is a heterogeneous multicore processor, combining a general-purpose POWER architecture core with eight independent single-instruction-multiple-data (SIMD) cores. Each core is capable of very high performance; however, users must explicitly manage data movement, scheduling, and synchronization. While these attributes provide some of the cell processorpsilas greatest performance strengths, they also form its greatest weaknesses in terms of developer productivity, code portability, and initial performance efficiencies. In this paper, we evaluate productivity and relative performance improvements of a cell BE system for a diverse set of kernels and applications. Our experimental workload includes algorithms from scientific, cognitive, and imaging problem domains. Our results demonstrate that the cell processor could be several times faster than a SSE-enabled, contemporary dual-core processor, and could sustain a high performance-to-productivity ratio. We outline strategies for transforming applications to exploit the cellpsilas architectural features, and measure productivity by comparing programming effort in terms of lines of code and performance. For instance, our measurements revealed that a covariance matrix creation routine - a common routine in hyperspectral imaging - ran over eight times faster than a 2.66 GHz Intel Woodcrest processor while sustaining a productivity metric of over two by parallelizing across the heterogeneous cores, unrolling loops, and improving instruction level parallelism with SIMD instructions in a high-level language.
Sadaf R. Alam, Jeremy S. Meredith, Jeffrey S. Vetter
CLUSTER2
2007 Analysis of a Computational Biology Simulation Technique on Emerging Processing Architectures
abstract
Multi-paradigm, multi-threaded and multi-core computing devices available today provide several orders of magnitude performance improvement over mainstream microprocessors. These devices include the STI Cell Broadband Engine, graphical processing units (GPU) and the Cray massively-multithreaded processors - available in desktop computing systems as well as proposed for supercomputing platforms. The main challenge in utilizing these powerful devices is their unique programming paradigms. GPUs and the Cell systems require code developers to manage code and data explicitly, while the Cray multithreaded architecture requires them to generate a very large number of threads or independent tasks concurrently. In this paper, we explain strategies for optimizing a molecular dynamics (MD) calculation that is used in biomolecular simulations on three devices: Cell, GPU and MTA-2. We show that the Cray MTA-2 system requires minimal code modification and does not outperform the microprocessor runs; but it demonstrates an improved workload scaling behavior over the microprocessor implementation. On the other hand, substantial porting and optimization efforts on the Cell and the GPU systems result in a 5times to 6times improvement, respectively, over a 2.2 GHz Opteron system.
Jeremy S. Meredith, Sadaf R. Alam, Jeffrey S. Vetter
IPDPS1
2005 A Contract Based System For Large Data Visualization
abstract
VisIt is a richly featured visualization tool that is used to visualize some of the largest simulations ever run. The scale of these simulations requires that optimizations are incorporated into every operation VisIt performs. But the set of applicable optimizations that VisIt can perform is dependent on the types of operations being done. Complicating the issue, VisIt has a plugin capability that allows new, unforeseen components to be added, making it even harder to determine which optimizations can be applied. We introduce the concept of a contract to the standard data flow network design. This contract enables each component of the data flow network to modify the set of optimizations used. In addition, the contract allows for new components to be accommodated gracefully within VisIt's data flow network system.
Hank Childs, Eric Brugger, Kathleen S. Bonnell, Jeremy S. Meredith, Mark C. Miller, Brad Whitlock, Nelson L. Max
IEEE Visualization4