James P. Ahrens

dblp:88/2598 · also James Paul Ahrens, Jim Ahrens · DBLP profile ↗
← Back
51ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-9378-282XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2025 STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data
abstract
Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.
Daoce Wang, Pascal Grosset, Jesus Pulido, Jiannan Tian, Tushar M. Athawale, Jinda Jia, Baixi Sun, Boyuan Zhang 0002, Sian Jin, Kai Zhao 0008, James P. Ahrens, Fengguang Song
SC11
2024 A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization
abstract
Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to $3.3 \times$ under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.
Daoce Wang, Pascal Grosset, Jesus Pulido, Tushar M. Athawale, Jiannan Tian, Kai Zhao 0008, Zarija Lukic, Axel Huebl, Zhe Wang 0059, James P. Ahrens, Dingwen Tao
SC10
2024 TAC+: Optimizing Error-Bounded Lossy Compression for 3D AMR Simulations
abstract
Today's scientific simulations require significant data volume reduction because of the enormous amounts of data produced and the limited I/O bandwidth and storage space. Error-bounded lossy compression has been considered one of the most effective solutions to the above problem. However, little work has been done to improve error-bounded lossy compression for Adaptive Mesh Refinement (AMR) simulation data. Unlike the previous work that only leverages 1D compression, in this work, we propose an approach (TAC) to leverage high-dimensional SZ compression for each refinement level of AMR data. To remove the data redundancy across different levels, we propose several pre-process strategies and adaptively use them based on the data features. We further optimizeTACtoTAC+by improving the lossless encoding stage of SZ compression to handle many small AMR data blocks after the pre-processing efficiently. Experiments on 10 AMR datasets from three real-world large-scale AMR simulations demonstrate thatTAC+can improve the compression ratio by up to 4.9× under the same data distortion, compared to the state-of-the-art method. In addition, we leverage the flexibility of our approach to tune the error bound for each level, which achieves much lower data distortion on two application-specific metrics.
Daoce Wang, Jesus Pulido, Pascal Grosset, Sian Jin, Jiannan Tian, Kai Zhao 0008, James P. Ahrens, Dingwen Tao
IEEE Trans. Parallel Distributed Syst.7
2024 MolSieve: A Progressive Visual Analytics System for Molecular Dynamics Simulations
abstract
Molecular Dynamics (MD) simulations are ubiquitous in cutting-edge physio-chemical research. They provide critical insights into how a physical system evolves over time given a model of interatomic interactions. Understanding a system's evolution is key to selecting the best candidates for new drugs, materials for manufacturing, and countless other practical applications. With today's technology, these simulations can encompass millions of unit transitions between discrete molecular structures, spanning up to several milliseconds of real time. Attempting to perform a brute-force analysis with data-sets of this size is not only computationally impractical, but would not shed light on the physically-relevant features of the data. Moreover, there is a need to analyze simulation ensembles in order to compare similar processes in differing environments. These problems call for an approach that is analytically transparent, computationally efficient, and flexible enough to handle the variety found in materials-based research. In order to address these problems, we introduce MolSieve, a progressive visual analytics system that enables the comparison of multiple long-duration simulations. Using MolSieve, analysts are able to quickly identify and compare regions of interest within immense simulations through its combination of control charts, data-reduction techniques, and highly informative visual components. A simple programming interface is provided which allows experts to fit MolSieve to their needs. To demonstrate the efficacy of our approach, we present two case studies of MolSieve and report on findings from domain collaborators.
Rostyslav Hnatyshyn, Jieqiong Zhao, Danny Perez, James P. Ahrens, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.4
2023 AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications
abstract
As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences exponential growth. Adaptive Mesh Refinement (AMR) has emerged as an effective solution to address these two challenges. Concurrently, error-bounded lossy compression is recognized as one of the most efficient approaches to tackle the latter issue. Despite their respective advantages, few attempts have been made to investigate how AMR and error-bounded lossy compression can function together. To this end, this study presents a novel in-situ lossy compression framework that employs the HDF5 filter to improve both I/O costs and boost compression quality for AMR applications. We implement our solution into the AMReX framework and evaluate on two real-world AMR applications, Nyx and WarpX, on the Summit supercomputer. Experiments with 4096 CPU cores demonstrate that AMRIC improves the compression ratio by up to 81× and the I/O performance by up to 39× over AMReX's original compression solution.
Daoce Wang, Jesus Pulido, Pascal Grosset, Jiannan Tian, Sian Jin, Houjun Tang, Jean M. Sexton, Sheng Di, Kai Zhao 0008, Bo Fang 0002, Zarija Lukic, Franck Cappello, James P. Ahrens, Dingwen Tao
SC13
2022 TAC: Optimizing Error-Bounded Lossy Compression for Three-Dimensional Adaptive Mesh Refinement Simulations
abstract
Today's scientific simulations require a significant reduction of data volume because of extremely large amounts of data they produce and the limited I/O bandwidth and storage space. Error-bounded lossy compression has been considered one of the most effective solutions to the above problem. However, little work has been done to improve error-bounded lossy compression for Adaptive Mesh Refinement (AMR) simulation data. Unlike the previous work that only leverages 1D compression, in this work, we propose to leverage high-dimensional (e.g., 3D) compression for each refinement level of AMR data. To remove the data redundancy across different levels, we propose three pre-process strategies and adaptively use them based on the data characteristics. Experiments on seven AMR datasets from a real-world large-scale AMR simulation demonstrate that our proposed approach can improve the compression ratio by up to 3.3X under the same data distortion, compared to the state-of-the-art method. In addition, we leverage the flexibility of our approach to tune the error bound for each level, which achieves much lower data distortion on two application-specific metrics.
Daoce Wang, Jesus Pulido, Pascal Grosset, Sian Jin, Jiannan Tian, James P. Ahrens, Dingwen Tao
HPDC6
2022 Optimization and Augmentation for Data Parallel Contour Trees
abstract
Contour trees are used for topological data analysis in scientific visualization. While originally computed with serial algorithms, recent work has introduced a vector-parallel algorithm. However, this algorithm is relatively slow for fully augmented contour trees which are needed for many practical data analysis tasks. We therefore introduce a representation called the hyperstructure that enables efficient searches through the contour tree and use it to construct a fully augmented contour tree in data parallel, with performance on average 6 times faster than the state-of-the-art parallel algorithm in the TTK topological toolkit.
Hamish A. Carr, Oliver Rübel, Gunther H. Weber, James P. Ahrens
IEEE Trans. Vis. Comput. Graph.4
2021 In Situ Adaptive Spatio-Temporal Data Summarization
abstract
Scientists nowadays use data sets generated from large-scale scientific computational simulations to understand the intricate details of various physical phenomena. These simulations produce large volumes of data at a rapid pace, containing thousands of time steps so that the spatiotemporal dynamics of the modeled phenomenon and its associated features can be captured with sufficient detail. Storing all the time steps into disks to perform traditional offline analysis will soon become prohibitive as the gap between the data generation speed and disk I/O speed continues to increase. In situ analysis, i.e., in-place analysis of data when it is being produced, has emerged as a solution to this problem. In this work, we present an information-theoretic approach for in situ reduction of large-scale time-varying data sets via a combination of key and fused time steps. We show that this approach can greatly minimize the output data storage footprint while preserving the temporal evolution of data features. A detailed in situ application study is carried out to demonstrate the in situ viability of our technique for efficiently summarizing thousands of time steps generated from a large-scale real-life computational simulation code.
Soumya Dutta, Humayra Tasnim, Terece L. Turton, James P. Ahrens
IEEE BigData4
2021 Adaptive Configuration of In Situ Lossy Compression for Cosmology Simulations via Fine-Grained Rate-Quality Modeling
abstract
Extreme-scale cosmological simulations have been widely used by today's researchers and scientists on leadership supercomputers. A new generation of error-bounded lossy compressors has been used in workflows to reduce storage requirements and minimize the impact of throughput limitations while saving large snapshots of high-fidelity data for post-hoc analysis. In this paper, we propose to adaptively provide compression configurations to compute partitions of cosmological simulations with newly designed post-analysis aware rate-quality modeling. The contribution is fourfold: (1) We propose a novel adaptive approach to select feasible error bounds for different partitions, showing the possibility and efficiency of adaptively configuring lossy compression for each partition individually. (2) We build models to estimate the overall loss of post-analysis result due to lossy compression and to estimate compression ratio, based on the property of each partition. (3) We develop an efficient optimization guideline to determine the best-fit configuration of error bounds combination in order to maximize the compression ratio under acceptable post-analysis quality loss. (4) Our approach introduces negligible overheads for feature extraction and error-bound optimization for each partition, enabling post-analysis-aware in situ lossy compression for cosmological simulations. Experiments show that our proposed models are highly accurate and reliable. Our fine-grained adaptive configuration approach improves the compression ratio of up to 73% on the tested datasets with the same post-analysis distortion with only 1% performance overhead.
Sian Jin, Jesus Pulido, Pascal Grosset, Jiannan Tian, Dingwen Tao, James P. Ahrens
HPDC6
2021 Probabilistic Data-Driven Sampling via Multi-Criteria Importance Analysis
abstract
Although supercomputers are becoming increasingly powerful, their components have thus far not scaled proportionately. Compute power is growing enormously and is enabling finely resolved simulations that produce never-before-seen features. However, I/O capabilities lag by orders of magnitude, which means only a fraction of the simulation data can be stored for post hoc analysis. Prespecified plans for saving features and quantities of interest do not work for features that have not been seen before. Data-driven intelligent sampling schemes are needed to detect and save important parts of the simulation while it is running. Here, we propose a novel sampling scheme that reduces the size of the data by orders-of-magnitude while still preserving important regions. The approach we develop selects points with unusual data values and high gradients. We demonstrate that our approach outperforms traditional sampling schemes on a number of tasks.
Ayan Biswas 0001, Soumya Dutta, Earl Lawrence, John Patchett, Jon Calhoun 0001, James P. Ahrens
IEEE Trans. Vis. Comput. Graph.6
2021 Scalable Contour Tree Computation by Data Parallel Peak Pruning
abstract
As data sets grow to exascale, automated data analysis and visualization are increasingly important, to intermediate human understanding and to reduce demands on disk storage via in situ analysis. Trends in architecture of high performance computing systems necessitate analysis algorithms to make effective use of combinations of massively multicore and distributed systems. One of the principal analytic tools is the contour tree, which analyses relationships between contours to identify features of more than local importance. Unfortunately, the predominant algorithms for computing the contour tree are explicitly serial, and founded on serial metaphors, which has limited the scalability of this form of analysis. While there is some work on distributed contour tree computation, and separately on hybrid GPU-CPU computation, there is no efficient algorithm with strong formal guarantees on performance allied with fast practical performance. We report the first shared SMP algorithm for fully parallel contour tree computation, with formal guarantees of O(lg V lg t) parallel steps and O(V lg V) work for data with V samples and t contour tree supernodes, and implementations with more than 30× parallel speed up on both CPU using TBB and GPU using Thrust and up 70× speed up compared to the serial sweep and merge algorithm.
Hamish A. Carr, Gunther H. Weber, Christopher M. Sewell, Oliver Rübel, Patricia K. Fasel, James P. Ahrens
IEEE Trans. Vis. Comput. Graph.6
2020 ETH: An Architecture for Exploring the Design Space of In-situ Scientific Visualization
abstract
As high-performance computing (HPC) moves towards the exascale era, large-scale scientific simulations are generating enormous datasets. Many techniques (e.g., in-situ methods, data sampling, and compression) have been proposed to help visualize these large datasets under various constraints such as storage, power, and energy. However, evaluating these techniques and understanding the trade-offs (e.g., performance, efficiency, and quality) remains a challenging task.To enable exploration of the design space across such trade-offs, we propose the Exploration Test Harness (ETH), an architecture for the early-stage exploration of visualization and rendering approaches, job layout, and visualization pipelines. ETH covers a broader parameter space than current large-scale visualization applications such as ParaView and VisIt. It also promotes the study of simulation-visualization coupling strategies through a data-centric approach, rather than requiring coupling with a specific scientific simulation code. Furthermore, with experimentation on an extensively instrumented supercomputer, we study more metrics of interest than was previously possible. Importantly, ETH will help to answer important what-if scenarios and trade-off questions in the early stages of pipeline development, helping scientists to make informed choices about how to best couple a simulation code with visualization at extreme scale.
Greg Abram, Vignesh Adhinarayanan, Wu-chun Feng, David H. Rogers 0001, James P. Ahrens
IPDPS5
2020 Understanding GPU-Based Lossy Compression for Extreme-Scale Cosmological Simulations
abstract
To help understand our universe better, researchers and scientists currently run extreme-scale cosmology simulations on leadership supercomputers. However, such simulations can generate large amounts of scientific data, which often result in expensive costs in data associated with data movement and storage. Lossy compression techniques have become attractive because they significantly reduce data size and can maintain high data fidelity for post-analysis. In this paper, we propose to use GPU-based lossy compression for extreme-scale cosmological simulations. Our contributions are threefold: (1) we implement multiple GPU-based lossy compressors to our open-source compression benchmark and analysis framework named Foresight; (2) we use Foresight to comprehensively evaluate the practicality of using GPU-based lossy compression on two real-world extreme-scale cosmology simulations, namely HACC and Nyx, based on a series of assessment metrics; and (3) we develop a general optimization guideline on how to determine the best-fit configurations for different lossy compressors and cosmological simulations. Experiments show that GPU-based lossy compression can provide necessary accuracy on post-analysis for cosmological simulations and high compression ratio of 5 ~ 15× on the tested datasets, as well as much higher compression and decompression throughput than CPU-based compressors.
Sian Jin, Pascal Grosset, Christopher M. Biwer, Jesus Pulido, Jiannan Tian, Dingwen Tao, James P. Ahrens
IPDPS7
2020 Foresight: analysis that matters for data reduction
abstract
As the computation power of supercomputers increases, so does simulation size, which in turn produces orders-of-magnitude more data. Because generated data often exceed the simulation's disk quota, many simulations would stand to benefit from data-reduction techniques to reduce storage requirements. Such techniques include autoencoders, data compression algorithms, and sampling. Lossy compression techniques can significantly reduce data size, but such techniques come at the expense of losing information that could result in incorrect post hoc analysis results. To help scientists determine the best compression they can get while keeping their analyses accurate, we have developed Foresight, an analysis framework that enables users to evaluate how different data-reduction techniques will impact their analyses. We use particle data from a cosmology simulation, turbulence data from Direct Numerical Simulation, and asteroid impact data from xRage to demonstrate how Foresight can help scientists determine the best data-reduction technique for their simulations.
Pascal Grosset, Christopher M. Biwer, Jesus Pulido, Arvind T. Mohan, Ayan Biswas 0001, John Patchett, Terece L. Turton, David H. Rogers 0001, Daniel Livescu, James P. Ahrens
SC10
2019 IEEE Visualization and Graphics Technical Committee (VGTC)
abstract
The IEEE Visualization and Graphics Technical Committee (VGTC) is a formal subcommittee of the Technical Activities Board (TAB) of the IEEE Computer Society.The VGTC provides technical leadership and organizes technical activities in the areas of visualization and visual analytics, computer graphics, virtual and augmented reality, and interaction.VGTC-sponsored conferences and symposia include VIS, VR,
James P. Ahrens
IEEE Trans. Vis. Comput. Graph.1
2019 Drag and Track: A Direct Manipulation Interface for Contextualizing Data Instances within a Continuous Parameter Space
abstract
We present a direct manipulation technique that allows material scientists to interactively highlight relevant parameterized simulation instances located in dimensionally reduced spaces, enabling a user-defined understanding of a continuous parameter space. Our goals are two-fold: first, to build a user-directed intuition of dimensionally reduced data, and second, to provide a mechanism for creatively exploring parameter relationships in parameterized simulation sets, called ensembles. We start by visualizing ensemble data instances in dimensionally reduced scatter plots. To understand these abstract views, we employ user-defined virtual data instances that, through direct manipulation, search an ensemble for similar instances. Users can create multiple of these direct manipulation queries to visually annotate the spaces with sets of highlighted ensemble data instances. User-defined goals are therefore translated into custom illustrations that are projected onto the dimensionally reduced spaces. Combined forward and inverse searches of the parameter space follow naturally allowing for continuous parameter space prediction and visual query comparison in the context of an ensemble. The potential for this visualization technique is confirmed via expert user feedback for a shock physics application and synthetic model analysis.
Daniel Orban, Daniel F. Keefe, Ayan Biswas 0001, James P. Ahrens, David H. Rogers 0001
IEEE Trans. Vis. Comput. Graph.4
2018 Modeling and Visualization of Uncertainty-Aware Geometry Using Multi-variate Normal Distributions
abstract
Many applications are dealing with geometric data that are affected by uncertainty. This uncertainty is important to analyze, visualize, and understand. We present a methodology to model uncertain geometry based on multi-variate normal distributions. In addition, we propose a visualization technique to represent a hull for uncertain geometry capturing a user-defined percentage of the underlying uncertain geometry. To show the effectiveness of our approach, we have modeled and visualized uncertain datasets from different applications.
Christina Gillmann, Thomas Wischgoll, Bernd Hamann, James P. Ahrens
PacificVis4
2018 Build and Execution Environment (BEE): an Encapsulated Environment Enabling HPC Applications Running Everywhere
abstract
Variations in High Performance Computing (HPC) system software configurations mean that applications are typically configured and built for specific HPC environments. Building applications can require a significant investment of time and effort for application users and requires application users to have additional technical knowledge. Linux container technologies such as Docker and Charliecloud bring great benefits to the application development, build and deployment processes. While cloud platforms already widely support containers, HPC systems still have non-uniform support of container technologies. In this work, we propose a unified runtime framework - Build and Execution Environment (BEE) across both HPC and cloud platforms that allows users to run their containerized HPC applications across all supported platforms without modification. We design four BEE backends for four different classes of HPC or cloud platform so that together they cover the majority of mainstream computing platforms for HPC users. Evaluations show that BEE provides an easy-to-use unified user interface, execution environment, and comparable performance.
Jieyang Chen, Qiang Guan, Xin Liang 0001, Paul Bryant, Patricia Grubel, Allen McPherson, Li-Ta Lo, Tim Randles, Zizhong Chen, James P. Ahrens
IEEE BigData10
2018 In situ TensorView: In situ Visualization of Convolutional Neural Networks
abstract
Convolutional Neural Networks(CNNs) are complex systems trained to recognize images, texts and more. However, once trained, they are regarded as black-boxes that are not easy to analyze and understand. Visualizing the dynamics within such deep artificial neural networks can provide a better understanding of how they are learning and making predictions. In the field of scientific simulations, visualization tools like Paraview have long been utilized to provide insights. We present in situ TensorView to visualize the training and functioning of CNNs as if they are systems of scientific simulations. In situ TensorView is a loosely coupled in situ visualization open framework that provides multiple viewers with the ability to visualize and understand their networks. It leverages the capability of co-processing from Paraview to provide real-time visualization during training and predicting phases, and avoids heavy I/O overhead. Tensorview is easily coupled with Tensorflow, as it only requires the insertion of a few lines of code into a TensorFlow framework. In this work, we showcase visualizing LeNet-5 and VGG16 using in situ TensorView. With the insight provided by Tensorview, users can adjust network architectures, or compress pre-trained networks guided by visualization results.
Qiang Guan, Li-Ta Lo, Simon Su, Zhengyong Ren, James P. Ahrens, Trilce Estrada
IEEE BigData6
2018 BeeFlow: A Workflow Management System for In Situ Processing across HPC and Cloud Systems
abstract
In this paper, we propose BeeFlow - an in situ analysis enabled workflow management system across multiple platforms using Docker containers. BeeFlow can support both traditional workflows as well as workflows with in situ analysis. BeeFlow leverages Docker containers to provide a portable, flexible, and reproducible workflow management system across HPC and cloud platforms. We showcase how current in situ visualization workflows can apply BeeFlow with DOE production codes VPIC and Flecsale.
Jieyang Chen, Qiang Guan, Zhao Zhang 0007, Xin Liang 0001, Louis James Vernon, Allen McPherson, Li-Ta Lo, Patricia Grubel, Tim Randles, Zizhong Chen, James P. Ahrens
ICDCS11
2018 Remote visual analysis of large turbulence databases at multiple scales
Jesus Pulido, Daniel Livescu, Kalin Kanov, Randal C. Burns, Curtis Canada, James P. Ahrens, Bernd Hamann
J. Parallel Distributed Comput.6
2018 The Good, the Bad, and the Ugly: A Theoretical Framework for the Assessment of Continuous Colormaps
abstract
A myriad of design rules for what constitutes a "good" colormap can be found in the literature. Some common rules include order, uniformity, and high discriminative power. However, the meaning of many of these terms is often ambiguous or open to interpretation. At times, different authors may use the same term to describe different concepts or the same rule is described by varying nomenclature. These ambiguities stand in the way of collaborative work, the design of experiments to assess the characteristics of colormaps, and automated colormap generation. In this paper, we review current and historical guidelines for colormap design. We propose a specified taxonomy and provide unambiguous mathematical definitions for the most common design rules.
Roxana Bujack, Terece L. Turton, Francesca Samsel, Colin Ware, David H. Rogers 0001, James P. Ahrens
IEEE Trans. Vis. Comput. Graph.6
2017 Homogeneity guided probabilistic data summaries for analysis and visualization of large-scale data sets
abstract
High-resolution simulation data sets provide plethora of information, which needs to be explored by application scientists to gain enhanced understanding about various phenomena. Visual-analytics techniques using raw data sets are often expensive due to the data sets' extreme sizes. But, interactive analysis and visualization is crucial for big data analytics, because scientists can then focus on the important data and make critical decisions quickly. To assist efficient exploration and visualization, we propose a new region-based statistical data summarization scheme. Our method is superior in quality, as compared to the existing statistical summarization techniques, with a more compact representation, reducing the overall storage cost. The quantitative and visual efficacy of our proposed method is demonstrated using several data sets along with an in situ application study for an extreme-scale flow simulation.
Soumya Dutta, Jonathan Woodring, Han-Wei Shen, Jen-Ping Chen, James P. Ahrens
PacificVis5
2017 Characterizing and Modeling Power and Energy for Extreme-Scale In-Situ Visualization
abstract
Plans for exascale computing have identified power and energy as looming problems for simulations running at that scale. In particular, writing to disk all the data generated by these simulations is becoming prohibitively expensive due to the energy consumption of the supercomputer while it idles waiting for data to be written to permanent storage. In addition, the power cost of data movement is also steadily increasing. A solution to this problem is to write only a small fraction of the data generated while still maintaining the cognitive fidelity of the visualization. With domain scientists increasingly amenable towards adopting an in-situ framework that can identify and extract valuable data from extremely large simulation results and write them to permanent storage as compact images, a large-scale simulation will commit to disk a reduced dataset of data extracts that will be much smaller than the raw results, resulting in a savings in both power and energy. The goal of this paper is two-fold: (i) to understand the role of in-situ techniques in combating power and energy issues of extreme-scale visualization and (ii) to create a model for performance, power, energy, and storage to facilitate what-if analysis. Our experiments on a specially instrumented, dedicated 150-node cluster show that while it is difficult to achieve power savings in practice using in-situ techniques, applications can achieve significant energy savings due to shorter write times for in-situ visualization. We present a characterization of power and energy for in-situ visualization; an application-aware, architecture-specific methodology for modeling and analysis of such in-situ workflows; and results that uncover indirect power savings in visualization workflows for high-performance computing (HPC).
Vignesh Adhinarayanan, Wu-chun Feng, David H. Rogers 0001, James P. Ahrens, Scott Pakin
IPDPS4
2017 Preface
abstract
The papers in this special issue were presented at IEEE VIS 2016, held during October 23-28, 2016 in Baltimore, MD. VIS contains three conferences, held concurrently: the IEEE Visual Analytics Science and Technology Conference (IEEE VAST 2016), the IEEE Information Visualization Conference (IEEE InfoVis 2016), and the IEEE Scientific Visualization Conference (IEEE SciVis2016).
Gennady L. Andrienko, Shixia Liu, John T. Stasko, Niklas Elmqvist, Bongshin Lee, Kwan-Liu Ma, James P. Ahrens, Robert M. Kirby, Jos B. T. M. Roerdink
IEEE Trans. Vis. Comput. Graph.7
2016 Animated versus static views of steady flow patterns
abstract
Two experiments were conducted to test the hypothesis that animated representations of vector fields are more effective than common static representations even for steady flow. We compared four flow visualization methods: animated streamlets, animated orthogonal line segments (where short lines were elongated orthogonal to the flow direction but animated in the direction of flow), static equally spaced streamlines, and static arrow grids. The first experiment involved a pattern detection task in which the participant searched for an anomalous flow pattern in a field of similar patterns. The results showed that both the animation methods produced more accurate and faster responses. The second experiment involved mentally tracing an advection path from a central dot in the flow field and marking where the path would cross the boundary of a surrounding circle. For this task the animated streamlets resulted in better performance than the other methods, but the animated orthogonal particles resulted in the worst performance. We conclude with recommendations for the representation of steady flow patterns.
Colin Ware, Daniel Bolan, Ricky Miller, David H. Rogers 0001, James P. Ahrens
SAP5
2016 In Situ Methods, Infrastructures, and Applications on High Performance Computing Platforms
abstract
Abstract The considerable interest in the high performance computing (HPC) community regarding analyzing and visualization data without first writing to disk, i. e., in situ processing, is due to several factors. First is an I/O cost savings, where data is analyzed/visualized while being generated, without first storing to a filesystem. Second is the potential for increased accuracy, where fine temporal sampling of transient analysis might expose some complex behavior missed in coarse temporal sampling. Third is the ability to use all available resources, CPU's and accelerators, in the computation of analysis products. This STAR paper brings together researchers, developers and practitioners using in situ methods in extreme‐scale HPC with the goal to present existing methods, infrastructures, and a range of computational science and engineering applications using in situ analysis and visualization.
Andrew C. Bauer, Hasan Abbasi, James P. Ahrens, Hank Childs, Berk Geveci, Scott Klasky, Kenneth Moreland, Patrick O'Leary, Venkatram Vishwanath, Brad Whitlock, E. Wes Bethel
Comput. Graph. Forum3
2016 Cinema image-based in situ analysis and visualization of MPAS-ocean simulations
Patrick O'Leary, James P. Ahrens, Sébastien Jourdain, Scott Wittenburg, David H. Rogers 0001, Mark R. Petersen
Parallel Comput.2
2016 In Situ Eddy Analysis in a High-Resolution Ocean Climate Model
abstract
An eddy is a feature associated with a rotating body of fluid, surrounded by a ring of shearing fluid. In the ocean, eddies are 10 to 150 km in diameter, are spawned by boundary currents and baroclinic instabilities, may live for hundreds of days, and travel for hundreds of kilometers. Eddies are important in climate studies because they transport heat, salt, and nutrients through the world's oceans and are vessels of biological productivity. The study of eddies in global ocean-climate models requires large-scale, high-resolution simulations. This poses a problem for feasible (timely) eddy analysis, as ocean simulations generate massive amounts of data, causing a bottleneck for traditional analysis workflows. To enable eddy studies, we have developed an in situ workflow for the quantitative and qualitative analysis of MPAS-Ocean, a high-resolution ocean climate model, in collaboration with the ocean model research and development process. Planned eddy analysis at high spatial and temporal resolutions will not be possible with a postprocessing workflow due to various constraints, such as storage size and I/O time, but the in situ workflow enables it and scales well to ten-thousand processing elements.
Jonathan Woodring, Mark R. Petersen, Andre Schmeißer, John Patchett, James P. Ahrens, Hans Hagen
IEEE Trans. Vis. Comput. Graph.5
2015 Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approach
abstract
Large-scale simulations can produce hundreds of terabytes to petabytes of data, complicating and limiting the efficiency of workflows. Traditionally, outputs are stored on the file system and analyzed in post-processing. With the rapidly increasing size and complexity of simulations, this approach faces an uncertain future. Trending techniques consist of performing the analysis in-situ, utilizing the same resources as the simulation, and/or off-loading subsets of the data to a compute-intensive analysis system. We introduce an analysis framework developed for HACC, a cosmological N-body code, that uses both in-situ and co-scheduling approaches for handling petabyte-scale outputs. We compare different analysis set-ups ranging from purely off-line, to purely in-situ to in-situ/co-scheduling. The analysis routines are implemented using the PISTON/VTK-m framework, allowing a single implementation of an algorithm that simultaneously targets a variety of GPU, multi-core, and many-core architectures.
Christopher M. Sewell, Katrin Heitmann, Hal Finkel, George Zagaris, Suzanne Parete-Koon, Patricia K. Fasel, Adrian Pope, Nicholas Frontiere, Li-Ta Lo, O. E. Bronson Messer, Salman Habib 0002, James P. Ahrens
SC12
2014 An Image-Based Approach to Extreme Scale in Situ Visualization and Analysis
abstract
Extreme scale scientific simulations are leading a charge to exascale computation, and data analytics runs the risk of being a bottleneck to scientific discovery. Due to power and I/O constraints, we expect in situ visualization and analysis will be a critical component of these workflows. Options for extreme scale data analysis are often presented as a stark contrast: write large files to disk for interactive, exploratory analysis, or perform in situ analysis to save detailed data about phenomena that a scientists knows about in advance. We present a novel framework for a third option - a highly interactive, image-based approach that promotes exploration of simulation results, and is easily accessed through extensions to widely used open source tools. This in situ approach supports interactive exploration of a wide range of results, while still significantly reducing data movement and storage.
James P. Ahrens, Sébastien Jourdain, Patrick O'Leary, John Patchett, David H. Rogers 0001, Mark R. Petersen
SC1
2013 Taming massive distributed datasets: data sampling using bitmap indices
Yu Su 0011, Gagan Agrawal, Jonathan Woodring, Kary L. Myers, Joanne Wendelberger, James P. Ahrens
HPDC6
2012 Analyzing the evolution of large scale structures in the universe with velocity based methods
abstract
The formation of cosmic structure results from the action of gravity on matter in an expanding Universe. As the evolution proceeds, the velocity field changes from being single-valued almost everywhere in space to being multi-valued over a complex web of `multistreaming' regions associated with the formation of large-scale structure (LSS) such as halos (or clumps), filaments, and sheets. Until recently, these structures have been investigated primarily via the (scalar) mass density field. In this application paper we apply data analysis and visualization techniques to cosmological simulations with the aim of studying multistreaming regions using velocity-based probes. Compared to the current practice of using density information (e.g., morphology estimators, locating overdense regions with halo finders), we show that velocity-based methods can provide useful supporting, as well as complementary, information. Because the density field and multistreaming are correlated but do not contain the same information, new and interesting information about the properties of the large-scale structure may be extracted, e.g., capturing dynamical behavior not possible with density-based estimators. Incorporating a novel method for setting thresholds for the velocity-based estimators, we study the relationships between the density field as represented by compact overdense halos and the different properties of multistreaming regions as represented by different velocity-based estimators.
Uliana Popov, Eddy Chandra, Katrin Heitmann, Salman Habib 0002, James P. Ahrens, Alex T. Pang
PacificVis5
2012 Jitter-free co-processing on a prototype exascale storage stack
abstract
In the petascale era, the storage stack used by the extreme scale high performance computing community is fairly homogeneous across sites. On the compute edge of the stack, file system clients or IO forwarding services direct IO over an interconnect network to a relatively small set of IO nodes. These nodes forward the requests over a secondary storage network to a spindle-based parallel file system. Unfortunately, this architecture will become unviable in the exascale era. As the density growth of disks continues to outpace increases in their rotational speeds, disks are becoming increasingly cost-effective for capacity but decreasingly so for bandwidth. Fortunately, new storage media such as solid state devices are filling this gap; although not cost-effective for capacity, they are so for performance. This suggests that the storage stack at exascale will incorporate solid state storage between the compute nodes and the parallel file systems. There are three natural places into which to position this new storage layer: within the compute nodes, the IO nodes, or the parallel file system. In this paper, we argue that the IO nodes are the appropriate location for HPC workloads and show results from a prototype system that we have built accordingly. Running a pipeline of computational simulation and visualization, we show that our prototype system reduces total time to completion by up to 30%.
John Bent, Sorin Faibish, James P. Ahrens, Gary Grider, John Patchett, Percy Tzelnic, Jonathan Woodring
MSST3
2012 Interface Exchange as an Indicator for Eddy Heat Transport
abstract
Abstract The ocean contains many large‐scale, long‐lived vortices, called mesoscale eddies, that are believed to have a role in the transport and redistribution of salt, heat, and nutrients throughout the ocean. Determining this role, however, has proven to be a challenge, since the mechanics of eddies are only partly understood; a standard definition for these ocean eddies does not exist and, therefore, scientifically meaningful, robust methods for eddy extraction, characterization, tracking and visualization remain a challenge. To shed light on the nature and potential roles of eddies, we extend our previous work on eddy identification and tracking to construct a new metric to characterize the transfer of water into and out of eddies across their boundary, and produce several visualizations of this new metric to provide clues about the role eddies play in the global ocean.
Sean Williams, Mark R. Petersen, Matthew Hecht, Mathew Maltrud, John Patchett, James P. Ahrens, Bernd Hamann
Comput. Graph. Forum6
2012 Guest Editor's Introduction: Special Section on the Eurographics Symposium on Parallel Graphics and Visualization (EGPGV)
abstract
The articles in this special section contain selected papers from the Eurographics Symposium on Parallel Graphics and Visualization (EGPGV).
James P. Ahrens, Kurt Debattista
IEEE Trans. Vis. Comput. Graph.1
2011 VisIO: Enabling Interactive Visualization of Ultra-Scale, Time Series Data via High-Bandwidth Distributed I/O Systems
abstract
Petascale simulations compute at resolutions ranging into billions of cells and write terabytes of data for visualization and analysis. Interactive visualization of this time series is a desired step before starting a new run. The I/O subsystem and associated network often are a significant impediment to interactive visualization of time-varying data, as they are not configured or provisioned to provide necessary I/O read rates. In this paper, we propose a new I/O library for visualization applications: VisIO. Visualization applications commonly use N-to-N reads within their parallel enabled readers which provides an incentive for a shared-nothing approach to I/O, similar to other data-intensive approaches such as Hadoop. However, unlike other data-intensive applications, visualization requires: (1) interactive performance for large data volumes, (2) compatibility with MPI and POSIX file system semantics for compatibility with existing infrastructure, and (3) use of existing file formats and their stipulated data partitioning rules. VisIO, provides a mechanism for using a non-POSIX distributed file system to provide linear scaling of I/O bandwidth. In addition, we introduce a novel scheduling algorithm that helps to co-locate visualization processes on nodes with the requested data. Testing using VisIO integrated into Para View was conducted using the Hadoop Distributed File System (HDFS) on TACC's Longhorn cluster. A representative dataset, VPIC, across 128 nodes showed a 64.4% read performance improvement compared to the provided Lustre installation. Also tested, was a dataset representing a global ocean salinity simulation that showed a 51.4% improvement in read performance over Lustre when using our VisIO system. VisIO, provides powerful high-performance I/O services to visualization applications, allowing for interactive performance with ultra-scale, time-series data.
Christopher Mitchell, James P. Ahrens, Jun Wang 0001
IPDPS2
2011 Visualization and Analysis of Eddies in a Global Ocean Simulation
abstract
Abstract We present analysis and visualization of flow data from a high‐resolution simulation of the dynamical behavior of the global ocean. Of particular scientific interest are coherent vortical features called mesoscale eddies. We first extract high‐vorticity features using a metric from the oceanography community called the Okubo‐Weiss parameter. We then use a new circularity criterion to differentiate eddies from other non‐eddy features like meanders in strong background currents. From these data, we generate visualizations showing the three‐dimensional structure and distribution of ocean eddies. Additionally, the characteristics of each eddy are recorded to form an eddy census that can be used to investigate correlations among variables such as eddy thickness, depth, and location. From these analyses, we gain insight into the role eddies play in large‐scale ocean circulation.
Sean Williams, Matthew Hecht, Mark R. Petersen, Richard Strelitz, Mathew Maltrud, James P. Ahrens, Mario Hlawitschka, Bernd Hamann
Comput. Graph. Forum6
2011 In-situ Sampling of a Large-Scale Particle Simulation for Interactive Visualization and Analysis
abstract
Abstract We describe a simulation‐time random sampling of a large‐scale particle simulation, the RoadRunner Universe MC3cosmological simulation, for interactive post‐analysis and visualization. Simulation data generation rates will continue to be far greater than storage bandwidth rates by many orders of magnitude. This implies that only a very small fraction of data generated by a simulation can ever be stored and subsequently post‐analyzed. The limiting factors in this situation are similar to the problem in many population surveys: there aren't enough human resources to query a large population. To cope with the lack of resources, statistical sampling techniques are used to create a representative data set of a large population. Following this analogy, we propose to store a simulation‐time random sampling of the particle data for post‐analysis, with level‐of‐detail organization, to cope with the bottlenecks. A sample is stored directly from the simulation in a level‐of‐detail format for post‐visualization and analysis, which amortizes the cost of post‐processing and reduces workflow time. Additionally by sampling during the simulation, we are able to analyze the entire particle population to record full population statistics and quantify sample error.
Jonathan Woodring, James P. Ahrens, J. Figg, Joanne Wendelberger, Salman Habib 0002, Katrin Heitmann
Comput. Graph. Forum2
2011 Adaptive Extraction and Quantification of Geophysical Vortices
abstract
We consider the problem of extracting discrete two-dimensional vortices from a turbulent flow. In our approach we use a reference model describing the expected physics and geometry of an idealized vortex. The model allows us to derive a novel correlation between the size of the vortex and its strength, measured as the square of its strain minus the square of its vorticity. For vortex detection in real models we use the strength parameter to locate potential vortex cores, then measure the similarity of our ideal analytical vortex and the real vortex core for different strength thresholds. This approach provides a metric for how well a vortex core is modeled by an ideal vortex. Moreover, this provides insight into the problem of choosing the thresholds that identify a vortex. By selecting a target coefficient of determination (i.e., statistical confidence), we determine on a per-vortex basis what threshold of the strength parameter would be required to extract that vortex at the chosen confidence. We validate our approach on real data from a global ocean simulation and derive from it a map of expected vortex strengths over the global ocean.
Sean Williams, Mark R. Petersen, Peer-Timo Bremer, Matthew Hecht, Valerio Pascucci, James P. Ahrens, Mario Hlawitschka, Bernd Hamann
IEEE Trans. Vis. Comput. Graph.6
2010 Verification of the time evolution of cosmological simulations via hypothesis-driven comparative and quantitative visualization
abstract
We describe a visualization-assisted process for the verification of cosmological simulation codes. The need for code verification stems from the requirement for very accurate predictions in order to interpret observational data confidently. We compare different simulation algorithms in order to reliably predict differences in simulation results and understand their dependence on input parameter settings. Our verification process consists of the integration of iterative hypothesis-verification with comparative, feature and quantitative visualization. We validate this process by verifying the time evolution results of three different cosmology simulation codes. The purpose of this verification is to study the accuracy of AMR methods versus other N-body simulation methods for cosmological simulations.
Chung-Hsing Hsu, James P. Ahrens, Katrin Heitmann
PacificVis2
2009 VisMashup: Streamlining the Creation of Custom Visualization Applications
abstract
Visualization is essential for understanding the increasing volumes of digital data. However, the process required to create insightful visualizations is involved and time consuming. Although several visualization tools are available, including tools with sophisticated visual interfaces, they are out of reach for users who have little or no knowledge of visualization techniques and/or who do not have programming expertise. In this paper, we propose VisMashup, a new framework for streamlining the creation of customized visualization applications. Because these applications can be customized for very specific tasks, they can hide much of the complexity in a visualization specification and make it easier for users to explore visualizations by manipulating a small set of parameters. We describe the framework and how it supports the various tasks a designer needs to carry out to develop an application, from mining and exploring a set of visualization specifications (pipelines), to the creation of simplified views of the pipelines, and the automatic generation of the application and its interface. We also describe the implementation of the system and demonstrate its use in two real application scenarios.
Emanuele Santos, Lauro Didier Lins, James P. Ahrens, Juliana Freire, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.3
2007 Scout: a data-parallel programming language for graphics processors
Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Jamaludin Mohd-Yusof, Greg Roth, Sharen J. Cummins
Parallel Comput.3
2006 Ultra-scale visualization - Workshop on ultra-scale visualization
abstract
The output from the massively parallel scientific simulations is so voluminous and complex that advanced visualization technologies are necessary to interpret the calculated results. Even though visualization technology has progressed significantly in recent years, we are barely capable of visualizing and analyzing terascale data to its full extent, and petascale datasets are on the horizon. This workshop aims at addressing this pressing issue by fostering communication between visualization researchers and practitioners. The workshop attendees will be introduced to the latest and greatest research innovations in large data visualization and also help direct further research direction through an open discussion session.
James P. Ahrens, Hank Childs, John P. Clyne, E. Wes Bethel, Jian Huang 0007, Scott Klasky, Kwan-Liu Ma, Kenneth Moreland, Michael E. Papka, Valerio Pascucci, Han-Wei Shen, Deborah Silver
SC1
2004 Scout: A Hardware-Accelerated System for Quantitatively Driven Visualization and Analysis
abstract
Quantitative techniques for visualization are critical to the successful analysis of both acquired and simulated scientific data. Many visualization techniques rely on indirect mappings, such as transfer functions, to produce the final imagery. In many situations, it is preferable and more powerful to express these mappings as mathematical expressions, or queries, that can then be directly applied to the data. We present a hardware-accelerated system that provides such capabilities and exploits current graphics hardware for portions of the computational tasks that would otherwise be executed on the CPU. In our approach, the direct programming of the graphics processor using a concise data parallel language, gives scientists the capability to efficiently explore and visualize data sets.
Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Charles D. Hansen, Greg Roth
IEEE Visualization3
2003 Interoperability of Visualization Software and Data Models is NOT an Achievable Goal
abstract
The scientific visualization community faces a crisis: there exist many individual tools that can be used to perform visualization, but there is little, if any, hope of being able to use tools from different sources as part of a single application. As a result, our community is fractured, and can be characterized as "islands of capability." The purpose of this panel is to probe the issues that prevent such interoperability, and engage in frank discussion about how our community can rectify these maladies. The issues to be discussed include but are not limited to: (1)lack of "standards" for data storage and modelling of N-dimensional scientific data, similar to those used for raster image files; (2)lack of "standard" interfaces for common visualization tools; (3)the visualization needs of the computational science research community, who are the primary consumers of technology from the visualization community; (4)lack of organization within our community to push for definition and adoption of such "standards;" (5)lack of organization within our community to serve as a "broker" and "promoter" for tools that might conform to even the weakest of standards. The panelist lineup represents a diverse cross-section of expertise and opinions about the panel topic. The panelists themselves are in disagreement about the severity of the problem, and potential solutions. The topic of this panel is highly germane to future growth of visualization as a science, and promises to be highly engaging for panelists and audience members alike.
E. Wes Bethel, Greg Abram, John Shalf, Randy Frank, James P. Ahrens, Steven G. Parker, Nagiza F. Samatova, Mark C. Miller
IEEE Visualization5
2000 Next-generation visualization displays: the research challenges of building tiled displays (panel session)
James P. Ahrens, Kai Li 0001, Daniel A. Reed
IEEE Visualization1
1998 Multi-source data analysis challenges
Samuel P. Uselton, Lloyd Treinish, James P. Ahrens, E. Wes Bethel, Andrei State
IEEE Visualization3
1997 Wildfire visualization (case study)
abstract
The ability to forecast the progress of crisis events would significantly reduce human suffering and loss of life, the destruction of property and expenditures for assessment and recovery. Los Alamos National Laboratory has established a scientific thrust in crisis forecasting to address this national challenge. In the initial phase of this project, scientists at Los Alamos are developing computer models to predict the spread of a wildfire. Visualization of the results of the wildfire simulation will be used by scientists to assess the quality of the simulation and eventually by fire personnel as a visual forecast of the wildfire's evolution. The fire personnel and scientists want the visualization to look as realistic as possible without compromising scientific accuracy. This paper describes how the visualization was created, analyzes the tools and approach that were used, and suggests directions for future work and research.
James P. Ahrens, Patrick S. McCormick, James Bossert, Jon Reisner, Judith Winterkamp
IEEE Visualization1
1995 Cost-Effective Data-Parallel Load Balancing
James P. Ahrens, Charles D. Hansen
ICPP (2)1
1993 Fast data parallel polygon rendering
abstract
This paper describes a data parallel method forpolygon rendering on a massively parallel machine. This method, based on a simple shading model, is targeted for applications which require very fast rendering for extremely large sets of polygons. Such sets are found in many scienti c visualization applications. The renderer can handle arbitrarily complex polygons which need notbe meshed. Issues involving load balancing are addressed andadataparallel load balancing algorithm is presented. The rendering and load balancing algorithms are implemented onboth the CM-200 and the CM-5. Experimental results are presented. This rendering toolkit enables a scientist to display 3D shaded polygons directly from aparallel machine avoiding the transmission of huge amounts of data to a post-processing rendering system. 1
Frank A. Ortega, Charles D. Hansen, James P. Ahrens
SC3