EDBT 2026 Demo / reviewers in the wild / expert
Tom Peterka
dblp:26/2029 · also Thomas Peterka
· DBLP profile ↗
67ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0002-0525-3205ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | F-Hash: Feature-Based Hash Design for Time-Varying Volume Visualization via Multi-Resolution Tesseract EncodingabstractInteractive time-varying volume visualization is challenging due to its complex spatiotemporal features and sheer size of the dataset. Recent works transform the original discrete time-varying volumetric data into continuous Implicit Neural Representations (INR) to address the issues of compression, rendering, and super-resolution in both spatial and temporal domains. However, training the INR takes a long time to converge, especially when handling large-scale time-varying volumetric datasets. In this work, we proposed F-Hash, a novel feature-based multi-resolution Tesseract encoding architecture to greatly enhance the convergence speed compared with existing input encoding methods for modeling time-varying volumetric data. The proposed design incorporates multi-level collision-free hash functions that map dynamic 4D multi-resolution embedding grids without bucket waste, achieving high encoding capacity with compact encoding parameters. Our encoding method is agnostic to time-varying feature detection methods, making it a unified encoding solution for feature tracking and evolution visualization. Experiments show the F-Hash achieves state-of-the-art convergence speed in training various time-varying volumetric datasets for diverse features. We also proposed an adaptive ray marching algorithm to optimize the sample streaming for faster rendering of the time-varying neural representation. Jianxin Sun 0001, David Lenz 0002, Hongfeng Yu 0001, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Make the Fastest Faster: Importance Mask Synthesis for Interactive Volume Visualization Using Reconstruction Neural NetworksabstractVisualizing a large-scale volumetric dataset with high resolution is challenging due to the substantial computational time and space complexity. Recent deep learning-based image inpainting methods significantly improve rendering latency by reconstructing a high-resolution image for visualization in constant time on GPU from a partially rendered image where only a portion of pixels go through the expensive rendering pipeline. However, existing solutions need to render every pixel of either a predefined regular sampling pattern or an irregular sample pattern predicted from a low-resolution image rendering. Both methods require a significant amount of expensive pixel-level rendering. In this work, we provide Importance Mask Learning (IML) and Synthesis (IMS) networks, which are the first attempts to directly synthesize important regions of the regular sampling pattern from the user's view parameters, to further minimize the number of pixels to render by jointly considering the dataset, user behavior, and the downstream reconstruction neural network. Our solution is a unified framework to handle various types of inpainting methods through the proposed differentiable compaction/decompaction layers. Experiments show our method can further improve the overall rendering latency of state-of-the-art volume visualization methods using reconstruction neural network for free when rendering scientific volumetric datasets. Our method can also directly optimize the off-the-shelf pre-trained reconstruction neural networks without elongated retraining. Jianxin Sun 0001, David Lenz 0002, Hongfeng Yu 0001, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Regularized Multi-Decoder Ensemble for an Error-Aware Scene Representation NetworkabstractFeature grid Scene Representation Networks (SRNs) have been applied to scientific data as compact functional surrogates for analysis and visualization. As SRNs are black-box lossy data representations, assessing the prediction quality is critical for scientific visualization applications to ensure that scientists can trust the information being visualized. Currently, existing architectures do not support inference time reconstruction quality assessment, as coordinate-level errors cannot be evaluated in the absence of ground truth data. By employing the uncertain neural network architecture in feature grid SRNs, we obtain prediction variances during inference time to facilitate confidence-aware data reconstruction. Specifically, we propose a parameter-efficient multi-decoder SRN (MDSRN) architecture consisting of a shared feature grid with multiple lightweight multilayer perceptron decoders. MDSRN can generate a set of plausible predictions for a given input coordinate to compute the mean as the prediction of the multi-decoder ensemble and the variance as a confidence score. The coordinate-level variance can be rendered along with the data to inform the reconstruction quality, or be integrated into uncertainty-aware volume visualization algorithms. To prevent the misalignment between the quantified variance and the prediction quality, we propose a novel variance regularization loss for ensemble learning that promotes the Regularized multi-decoder SRN (RMDSRN) to obtain a more reliable variance that correlates closely to the true model error. We comprehensively evaluate the quality of variance quantification and data reconstruction of Monte Carlo Dropout (MCD), Mean Field Variational Inference (MFVI), Deep Ensemble (DE), and Predicting Variance (PV) in comparison with our proposed MDSRN and RMDSRN applied to state-of-the-art feature grid SRNs across diverse scalar field datasets. We demonstrate that RMDSRN attains the most accurate data reconstruction and competitive variance-error correlation among uncertain SRNs under the same neural network parameter budgets. Furthermore, we present an adaptation of uncertainty-aware volume rendering and shed light on the potential of incorporating uncertain predictions in improving the quality of volume rendering for uncertain SRNs. Through ablation studies on the regularization strength and decoder count, we show that MDSRN and RMDSRN are expected to perform sufficiently well with a default configuration without requiring customized hyperparameter settings for different datasets. Tianyu Xiong, Skylar W. Wurster, Hanqi Guo 0001, Tom Peterka, Han-Wei Shen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Viper: A High-Performance I/O Framework for Transparently Updating, Storing, and Transferring Deep Neural Network ModelsabstractScientific workflows increasingly need to train a DNN model in real-time during an experiment (e.g. using ground truth from a simulation), while using it at the same time for inferences. Instead of sharing the same model instance, the training (producer) and inference server (consumer) often use different model replicas that are kept synchronized. In addition to efficient I/O techniques to keep the model replica of the producer and consumer synchronized, there is another important trade-off: frequent model updates enhance inference quality but may slow down training; infrequent updates may lead to less precise inference results. To address these challenges, we introduce Viper: a new I/O framework designed to determine a near-optimal checkpoint schedule and accelerate the delivery of the latest model updates. Viper builds an inference performance predictor to identify the optimal checkpoint schedule to balance the trade-off between training slowdown and inference quality improvement. It also creates a memory-first model transfer engine to accelerate model delivery through direct memory-to-memory communication. Our experiments show that Viper can reduce the model update latency by ≈ 9x using the GPU-to-GPU data transfer engine and ≈ 3x using the DRAM-to-DRAM host data transfer. The checkpoint schedule obtained from Viper’s predictor also demonstrates improved cumulative inference accuracy compared to the baseline of epoch-based solutions. Jaime Cernuda, Neeraj Rajesh, Keith Bateman, Orcun Yildiz, Tom Peterka, Arnur Nigmetov, Dmitriy Morozov, Xian-He Sun, Antonios Kougkas, Bogdan Nicolae |
ICPP | 6 |
| 2024 | Extreme-scale workflows: A perspective from the JLESC international community
Orcun Yildiz, Amal Gueroudji, Julien Bigot, Bruno Raffin, Rosa M. Badia, Tom Peterka |
Future Gener. Comput. Syst. | 6 |
| 2024 | Automated defect identification in coherent diffraction imaging with smart continual learning
Orcun Yildiz, Raghavan Krishnan, Henry Chan, Mathew J. Cherukara, Prasanna Balaprakash, Subramanian Sankaranarayanan, Tom Peterka |
Neural Comput. Appl. | 7 |
| 2024 | Adaptively Placed Multi-Grid Scene Representation Networks for Large-Scale Data VisualizationabstractScene representation networks (SRNs) have been recently proposed for compression and visualization of scientific data. However, state-of-the-art SRNs do not adapt the allocation of available network parameters to the complex features found in scientific data, leading to a loss in reconstruction quality. We address this shortcoming with an adaptively placed multi-grid SRN (APMGSRN) and propose a domain decomposition training and inference technique for accelerated parallel training on multi-GPU systems. We also release an open-source neural volume rendering application that allows plug-and-play rendering with any PyTorch-based SRN. Our proposed APMGSRN architecture uses multiple spatially adaptive feature grids that learn where to be placed within the domain to dynamically allocate more neural network resources where error is high in the volume, improving state-of-the-art reconstruction accuracy of SRNs for scientific data without requiring expensive octree refining, pruning, and traversal like previous adaptive models. In our domain decomposition approach for representing large-scale data, we train an set of APMGSRNs in parallel on separate bricks of the volume to reduce training time while avoiding overhead necessary for an out-of-core solution for volumes too large to fit in GPU memory. After training, the lightweight SRNs are used for realtime neural volume rendering in our open-source renderer, where arbitrary view angles and transfer functions can be explored. A copy of this paper, all code, all models used in our experiments, and all supplemental materials and videos are available at https://github.com/skywolf829/APMGSRN. Skylar W. Wurster, Tianyu Xiong, Han-Wei Shen, Hanqi Guo 0001, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | TROPHY: A Topologically Robust Physics-Informed Tracking Framework for Tropical CyclonesabstractTropical cyclones (TCs) are among the most destructive weather systems. Realistically and efficiently detecting and tracking TCs are critical for assessing their impacts and risks. In particular, the eye is a signature feature of a mature TC. Therefore, knowing the eyes' locations and movements is crucial for both operational weather forecasts and climate risk assessments. Recently, a multilevel robustness framework has been introduced to study the critical points of time-varying vector fields. The framework quantifies the robustness (i.e., structural stability) of critical points across varying neighborhoods. By relating the multilevel robustness with critical point tracking, the framework has demonstrated its potential in cyclone tracking. An advantage is that it identifies cyclonic features using only 2D wind vector fields, which is encouraging as most tracking algorithms require multiple dynamic and thermodynamic variables at different altitudes. A disadvantage is that the framework does not scale well computationally for datasets containing a large number of cyclones. This paper introduces a topologically robust physics-informed tracking framework (TROPHY) for TC tracking. The main idea is to integrate physical knowledge of TC to drastically improve the computational efficiency of multilevel robustness framework for large-scale climate datasets. First, during preprocessing, we propose a physics-informed feature selection strategy to filter 90% of critical points that are short-lived and have low stability, thus preserving good candidates for TC tracking. Second, during in-processing, we impose constraints during the multilevel robustness computation to focus only on physics-informed neighborhoods of TCs. We apply TROPHY to 30 years of 2D wind fields from reanalysis data in ERA5 and generate a number of TC tracks. In comparison with the observed tracks, we demonstrate that TROPHY can capture TC characteristics (e.g., frequency, intensity, duration, latitudes with maximum intensity, and genesis) that are comparable to and sometimes even better than a well-validated TC tracking algorithm that requires multiple dynamic and thermodynamic scalar fields. Lin Yan 0003, Hanqi Guo 0001, Tom Peterka, Bei Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Scalable Volume Visualization for Big Scientific Data Modeled by Functional ApproximationabstractConsidering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches. Jianxin Sun 0001, David Lenz 0002, Hongfeng Yu 0001, Tom Peterka |
IEEE Big Data | 4 |
| 2023 | LowFive: In Situ Data Transport for High-Performance WorkflowsabstractWe describe LowFive, a new data transport layer based on the HDF5 data model, for in situ workflows. Executables using LowFive can communicate in situ (using in-memory data and MPI message passing), reading and writing traditional HDF5 files to physical storage, and combining the two modes. Minimal and often no source-code modification is needed for programs that already use HDF5. LowFive maintains deep copies or shallow references of datasets, configurable by the user. More than one task can produce (write) data, and more than one task can consume (read) data, accommodating fan-in and fan-out in the workflow task graph. LowFive supports data redistribution from n producer processes to m consumer processes. We demonstrate the above features in a series of experiments featuring both synthetic benchmarks as well as a representative use case from a scientific workflow, and we also compare with other data transport solutions in the literature. Tom Peterka, Dmitriy Morozov, Arnur Nigmetov, Orcun Yildiz, Bogdan Nicolae, Philip E. Davis |
IPDPS | 1 |
| 2023 | Neural Stream FunctionsabstractWe present a neural network approach to compute stream functions, which are scalar functions with gradients orthogonal to a given vector field. As a result, isosurfaces of the stream function extract stream surfaces, which can be visualized to analyze flow features. Our approach takes a vector field as input and trains an implicit neural representation to learn a stream function for that vector field. The network learns to map input coordinates to a stream function value by minimizing the inner product of the gradient of the neural network’s output and the vector field. Since stream function solutions may not be unique, we give optional constraints for the network to learn particular stream functions of interest. Specifically, we introduce regularizing loss functions that can optionally be used to generate stream function solutions whose stream surfaces follow the flow field’s curvature, or that can learn a stream function that includes a stream surface passing through a seeding rake. We also discuss considerations for properly visualizing the trained implicit network and extracting artifact-free surfaces. We compare our results with other implicit solutions and present qualitative and quantitative results for several synthetic and simulated vector fields. Skylar W. Wurster, Hanqi Guo 0001, Tom Peterka, Han-Wei Shen |
PacificVis | 3 |
| 2023 | Toward Feature-Preserving Vector Field CompressionabstractThe objective of this work is to develop error-bounded lossy compression methods to preserve topological features in 2D and 3D vector fields. Specifically, we explore the preservation of critical points in piecewise linear and bilinear vector fields. We define the preservation of critical points as, without any false positive, false negative, or false type in the decompressed data, (1) keeping each critical point in its original cell and (2) retaining the type of each critical point (e.g., saddle and attracting node). The key to our method is to adapt a vertex-wise error bound for each grid point and to compress input data together with the error bound field using a modified lossy compressor. Our compression algorithm can be also embarrassingly parallelized for large data handling and in situ processing. We benchmark our method by comparing it with existing lossy compressors in terms of false positive/negative/type rates, compression ratio, and various vector field visualizations with several scientific applications. Xin Liang 0001, Sheng Di, Franck Cappello, Mukund Raj, Kenji Ono, Zizhong Chen, Tom Peterka, Hanqi Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | Deep Hierarchical Super Resolution for Scientific DataabstractWe present a novel technique for hierarchical super resolution (SR) with neural networks (NNs), which upscales volumetric data represented with an octree data structure to a high-resolution uniform gridwith minimal seam artifacts on octree node boundaries. Our method uses existing state-of-the-art SR models and adds flexibility to upscale input data with varying levels of detail across the domain, instead of only uniform grid data that are supported in previous approaches.The key is to use a hierarchy of SR NNs, each trained to perform 2× SR between two levels of detail, with a hierarchical SR algorithm that minimizes seam artifacts by starting from the coarsest level of detail and working up.We show that our hierarchical approach outperforms baseline interpolation and hierarchical upscaling methods, and demonstrate the usefulness of our proposed approach across three use cases including data reduction using hierarchical downsampling+SR instead of uniform downsampling+SR, computation savings for hierarchical finite-time Lyapunov exponent field calculation, and super-resolving low-resolution simulation results for a high-resolution approximation visualization. Skylar W. Wurster, Hanqi Guo 0001, Han-Wei Shen, Tom Peterka, Jiayi Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Reinforcement Learning for Load-Balanced Parallel Particle TracingabstractWe explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors. Jiayi Xu 0001, Hanqi Guo 0001, Han-Wei Shen, Mukund Raj, Skylar W. Wurster, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Accelerating Scientific Workflows on HPC Platforms with In Situ ProcessingabstractScientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks. Tu Mai Anh Do, Loïc Pottier, Orcun Yildiz, Karan Vahi, Patrycja Krawczuk, Tom Peterka, Ewa Deelman |
CCGRID | 6 |
| 2022 | Towards Low-Overhead Resilience for Data Parallel Deep LearningabstractData parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead. Bogdan Nicolae, Tanner Hobson, Orcun Yildiz, Tom Peterka, Dmitriy Morozov |
CCGRID | 4 |
| 2022 | A Multi-Branch Decoder Network Approach to Adaptive Temporal Data Selection and Reconstruction for Big Scientific Simulation DataabstractA key challenge in scientific simulation is that the simulation outputs often require intensive I/O and storage space to store the results for effective post hoc analysis. This article focuses on aquality-aware adaptive temporal data selection and reconstructionproblem where the goal is to adaptively select simulation data samples at certain key timesteps in situ and reconstruct the discarded samples with quality assurance during post hoc analysis. This problem is motivated by the limitation of current solutions that a significant amount of simulation data samples are either discarded or aggregated during the sampling process, leading to inaccurate modeling of the simulated phenomena. Two unique challenges exist: 1) the sampling decisions have to be made in situ and adapted to the dynamics of the complex scientific simulation data; 2) the reconstruction error must be strictly bounded to meet the application requirement. To address the above challenges, we developDeepSample, an error-controlled convolutional neural network framework, that jointly integrates a set of coherent multi-branch deep decoders to effectively reconstruct the simulation data with rigorous quality assurance. The results on two real-world scientific simulation applications show that DeepSample significantly outperforms other state-of-the-art methods on both sampling efficiency and reconstructed simulation data quality. Yang Zhang 0031, Hanqi Guo 0001, Lanyu Shang, Dong Wang 0002, Tom Peterka |
IEEE Trans. Big Data | 5 |
| 2021 | Shared-Memory Communication for Containerized WorkflowsabstractScientific computation increasingly consists of a workflow of interrelated tasks. Containerization can make workflow systems more manageable, reproducible, and portable, but containers can impede communication due to their focus on encapsulation. In some circumstances, shared-memory regions are an effective way to improve performance of workflows; however sharing memory between containerized workflow tasks is difficult. In this work, we have created a software library called Dhmem that manages shared memory between workflow tasks in separate containers, with minimal code change and performance overhead. Instead of all code being in the same container, Dhmem allows a separate container for each workflow task to be constructed completely independently. Dhmem enables additional functionality: easy integration in existing workflow systems, communication configuration at runtime based on the environment, and scalable performance. Tanner Hobson, Orcun Yildiz, Bogdan Nicolae, Jian Huang 0007, Tom Peterka |
CCGRID | 5 |
| 2021 | FTK: A Simplicial Spacetime Meshing Framework for Robust and Scalable Feature TrackingabstractWe present the Feature Tracking Kit (FTK), a framework that simplifies, scales, and delivers various feature-tracking algorithms for scientific data. The key of FTK is our simplicial spacetime meshing scheme that generalizes both regular and unstructured spatial meshes to spacetime while tessellating spacetime mesh elements into simplices. The benefits of using simplicial spacetime meshes include (1) reducing ambiguity cases for feature extraction and tracking, (2) simplifying the handling of degeneracies using symbolic perturbations, and (3) enabling scalable and parallel processing. The use of simplicial spacetime meshing simplifies and improves the implementation of several feature-tracking algorithms for critical points, quantum vortices, and isosurfaces. As a software framework, FTK provides end users with VTK/ParaView filters, Python bindings, a command line interface, and programming interfaces for feature-tracking applications. We demonstrate use cases as well as scalability studies through both synthetic data and scientific applications including tokamak, fluid dynamics, and superconductivity simulations. We also conduct end-to-end performance studies on the Summit supercomputer. FTK is open sourced under the MIT license: https://github.com/hguo/ftk. Hanqi Guo 0001, David Lenz 0002, Jiayi Xu 0001, Xin Liang 0001, Iulian R. Grindeanu, Han-Wei Shen, Tom Peterka, Todd S. Munson, Ian T. Foster |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2021 | Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and AnalysisabstractWe present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. We demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively. Jiayi Xu 0001, Hanqi Guo 0001, Han-Wei Shen, Mukund Raj, Xueyun Wang, Xueqiao Xu, Zhehui Wang, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2020 | Toward Feature-Preserving 2D and 3D Vector Field CompressionabstractThe objective of this work is to develop error-bounded lossy compression methods to preserve topological features in 2D and 3D vector fields. Specifically, we explore the preservation of critical points in piecewise linear vector fields. We define the preservation of critical points as, without any false positive, false negative, or false type change in the decompressed data, (1) keeping each critical point in its original cell and (2) retaining the type of each critical point (e.g., saddle and attracting node). The key to our method is to adapt a vertex-wise error bound for each grid point and to compress input data together with the error bound field using a modified lossy compressor. Our compression algorithm can be also embarrassingly parallelized for large data handling and in situ processing. We benchmark our method by comparing it with existing lossy compressors in terms of false positive/negative/type rates, compression ratio, and various vector field visualizations with several scientific applications. Xin Liang 0001, Hanqi Guo 0001, Sheng Di, Franck Cappello, Mukund Raj, Kenji Ono, Zizhong Chen, Tom Peterka |
PacificVis | 9 |
| 2020 | Fast Automatic Knot Placement Method for Accurate B-spline Curve Fitting
Raine Yeh, Youssef S. G. Nashed, Tom Peterka, Xavier Tricoche |
Comput. Aided Des. | 3 |
| 2020 | eFESTA: Ensemble Feature Exploration with Surface Density EstimatesabstractWe propose surface density estimate (SDE) to model the spatial distribution of surface features-isosurfaces, ridge surfaces, and streamsurfaces-in 3D ensemble simulation data. The inputs of SDE computation are surface features represented as polygon meshes, and no field datasets are required (e.g., scalar fields or vector fields). The SDE is defined as the kernel density estimate of the infinite set of points on the input surfaces and is approximated by accumulating the surface densities of triangular patches. We also propose an algorithm to guide the selection of a proper kernel bandwidth for SDE computation. An ensemble Feature Exploration method based on Surface densiTy EstimAtes (eFESTA) is then proposed to extract and visualize the major trends of ensemble surface features. For an ensemble of surface features, each surface is first transformed into a density field based on its contribution to the SDE, and the resulting density fields are organized into a hierarchical representation based on the pairwise distances between them. The hierarchical representation is then used to guide visual exploration of the density fields as well as the underlying surface features. We demonstrate the application of our method using isosurface in ensemble scalar fields, Lagrangian coherent structures in uncertain unsteady flows, and streamsurfaces in ensemble fluid flows. Hanqi Guo 0001, Han-Wei Shen, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | InSituNet: Deep Image Synthesis for Parameter Space Exploration of Ensemble SimulationsabstractWe propose InSituNet, a deep learning based surrogate model to support parameter space exploration for ensemble simulations that are visualized in situ. In situ visualization, generating visualizations at simulation time, is becoming prevalent in handling large-scale simulations because of the I/O and storage constraints. However, in situ visualization approaches limit the flexibility of post-hoc exploration because the raw simulation data are no longer available. Although multiple image-based approaches have been proposed to mitigate this limitation, those approaches lack the ability to explore the simulation parameters. Our approach allows flexible exploration of parameter space for large-scale ensemble simulations by taking advantage of the recent advances in deep learning. Specifically, we design InSituNet as a convolutional regression model to learn the mapping from the simulation and visualization parameters to the visualization results. With the trained model, users can generate new images for different simulation parameters under various visualization settings, which enables in-depth analysis of the underlying ensemble simulations. We demonstrate the effectiveness of InSituNet in combustion, cosmology, and ocean simulations through quantitative and qualitative evaluations. Junpeng Wang 0001, Hanqi Guo 0001, Ko-Chih Wang, Han-Wei Shen, Mukund Raj, Youssef S. G. Nashed, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2020 | CECAV-DNN: Collective Ensemble Comparison and Visualization using Deep Neural NetworksabstractWe propose a deep learning approach to collectively compare two or multiple ensembles, each of which is a collection of simulation outputs. The purpose of collective comparison is to help scientists understand differences between simulation models by comparing their ensemble simulation outputs. However, the collective comparison is non-trivial because the spatiotemporal distributions of ensemble simulation outputs reside in a very high dimensional space. To this end, we choose to train a deep discriminative neural network to measure the dissimilarity between two given ensembles, and to identify when and where the two ensembles are different. We also design and develop a visualization system to help users understand the collective comparison results based on the discriminative network. We demonstrate the effectiveness of our approach with two real-world applications, including the ensemble comparison of the community atmosphere model (CAM) and the rapid radiative transfer model for general circulation models (RRTMG) for climate research, and the comparison of computational fluid dynamics (CFD) ensembles with different spatial resolutions. Junpeng Wang 0001, Hanqi Guo 0001, Han-Wei Shen, Tom Peterka |
Vis. Informatics | 5 |
| 2019 | Scalable, High-Order Continuity Across Block Boundaries of Functional Approximations Computed in ParallelabstractWe investigate the representation of discrete scientific data with a Ckfunctional model, where Ckdenotes k-th order continuity, in a distributed-memory parallel setting. The Multivariate Functional Approximation (MFA) model is a piecewise-continuous functional approximation based on multi-variate high-dimensional B-splines. When computing an MFA approximation in parallel over multiple blocks in a spatial domain decomposition, the interior of each block will be Ck, k being the B-spline polynomial degree, but discontinuities exist across neighboring block boundaries. We present an efficient and scalable solution that involves blending neighboring approximations to ensure Ckcontinuity across block boundaries. We show that after decomposing the domain in structured, overlapping blocks and approximating blocks independently to high degrees of accuracy, we can extend the local solution, in a postprocessing step, to the global domain by using compact, multidimensional smoothstep functions. We prove that this approach, which can be viewed as an extended partition of unity approximation method, is scalable on high-performance computing architectures. Iulian R. Grindeanu, Tom Peterka, Vijay Mahadevan, Youssef S. G. Nashed |
CLUSTER | 2 |
| 2019 | A Scalable Hybrid Scheme for Ray-Casting of Unstructured Volume DataabstractWe present an algorithm for parallel volume rendering that is a hybrid between classical object order and image order techniques. The algorithm operates on unstructured grids (and structured ones), and thus can deal with block boundaries interleaving in complex ways. It also deals effectively with cases that are prone to load imbalance, i.e., cases where cell sizes differ dramatically, either because of the nature of the input data, or because of the effects of the camera transformation. The algorithm divides work over resources such that each phase of its processing is bounded in the amount of computation it can perform. We demonstrate its efficacy through a series of studies, varying over camera position, data set size, transfer function, image size, and processor count. At its biggest, our experiments scaled up to 8,192 processors and operated on data sets with more than one billion cells. In total, we find that our hybrid algorithm performs well in all cases. This is because our algorithm naturally adapts its computation based on workload, and can operate like either an object order technique or an image order technique in scenarios where those techniques are efficient. Roba Binyahib, Tom Peterka, Matthew Larsen, Kwan-Liu Ma, Hank Childs |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Extreme-Scale Stochastic Particle Tracing for Uncertain Unsteady Flow Visualization and AnalysisabstractWe present an efficient and scalable solution to estimate uncertain transport behaviors-stochastic flow maps (SFMs)-for visualizing and analyzing uncertain unsteady flows. Computing flow maps from uncertain flow fields is extremely expensive because it requires many Monte Carlo runs to trace densely seeded particles in the flow. We reduce the computational cost by decoupling the time dependencies in SFMs so that we can process shorter sub time intervals independently and then compose them together for longer time periods. Adaptive refinement is also used to reduce the number of runs for each location. We parallelize over tasks-packets of particles in our design-to achieve high efficiency in MPI/thread hybrid programming. Such a task model also enables CPU/GPU coprocessing. We show the scalability on two supercomputers, Mira (up to 256K Blue Gene/Q cores) and Titan (up to 128K Opteron cores and 8K GPUs), that can trace billions of particles in seconds. Hanqi Guo 0001, Han-Wei Shen, Emil M. Constantinescu, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | Dynamic Data Repartitioning for Load-Balanced Parallel Particle TracingabstractWe present a novel dynamic load-balancing algorithm based on data repartitioning for parallel particle tracing in flow visualization. Instead of static data assignment, we dynamically repartition the data into blocks and reassign the blocks to processes to balance the workload distribution among the processes. Block repartitioning is performed based on a dynamic workload estimation method that predicts the workload in the flow field on the fly as the input. In our approach, we allow data duplication in the repartitioning, enabling the same data blocks to be assigned to multiple processes. Load balance is achieved by regularly exchanging the blocks (together with the particles in the blocks) among processes according to the output of the data repartitioning. Compared with other load-balancing algorithms, our approach does not need any preprocessing on the raw data and does not require any dedicated process for work scheduling, while it has the capability to balance uneven workload efficiently. Results show improved load balance and high efficiency of our method on tracing particles in both steady and unsteady flow. Jiang Zhang 0002, Hanqi Guo 0001, Xiaoru Yuan, Tom Peterka |
PacificVis | 4 |
| 2018 | Spark-DIY: A Framework for Interoperable Spark Operations with High Performance Block-Based Data ModelsabstractToday's scientific applications are increasingly relying on a variety of data sources, storage facilities, and computing infrastructures, and there is a growing demand for data analysis and visualization for these applications. In this context, exploiting Big Data frameworks for scientific computing is an opportunity to incorporate high-level libraries, platforms, and algorithms for machine learning, graph processing, and streaming; inherit their data awareness and fault-tolerance; and increase productivity. Nevertheless, limitations exist when Big Data platforms are integrated with an HPC environment, namely poor scalability, severe memory overhead, and huge development effort. This paper focuses on a popular Big Data framework -Apache Spark- and proposes an architecture to support the integration of highly scalable MPI block-based data models and communication patterns with a map-reduce-based programming model. The resulting platform preserves the data abstraction and programming interface of Spark, without conducting any changes in the framework, but allows the user to delegate operations to the MPI layer. The evaluation of our prototype shows that our approach integrates Spark and MPI efficiently at scale, so end users can take advantage of the productivity facilitated by the rich ecosystem of high-level Big Data tools and libraries based on Spark, without compromising efficiency and scalability. Silvina Caíno-Lores, Jesús Carretero 0001, Bogdan Nicolae, Orcun Yildiz, Tom Peterka |
BDCAT | 5 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 30 |
| 2018 | Dynamic Load Balancing Based on Constrained K-D Tree Decomposition for Parallel Particle TracingabstractWe propose a dynamically load-balanced algorithm for parallel particle tracing, which periodically attempts to evenly redistribute particles across processes based on k-d tree decomposition. Each process is assigned with (1) a statically partitioned, axis-aligned data block that partially overlaps with neighboring blocks in other processes and (2) a dynamically determined k-d tree leaf node that bounds the active particles for computation; the bounds of the k-d tree nodes are constrained by the geometries of data blocks. Given a certain degree of overlap between blocks, our method can balance the number of particles as much as possible. Compared with other load-balancing algorithms for parallel particle tracing, the proposed method does not require any preanalysis, does not use any heuristics based on flow features, does not make any assumptions about seed distribution, does not move any data blocks during the run, and does not need any master process for work redistribution. Based on a comprehensive performance study up to 8K processes on a Blue Gene/Q system, the proposed algorithm outperforms baseline approaches in both load balance and scalability on various flow visualization and analysis problems. Jiang Zhang 0002, Hanqi Guo 0001, Fan Hong, Xiaoru Yuan, Tom Peterka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | In situ magnetic flux vortex visualization in time-dependent Ginzburg-Landau superconductor simulationsabstractWe present an in situ visualization framework to capture comprehensive details of vortex dynamics in superconductor simulations. Vortices, which determine all electromagnetic properties of type-II superconductors, are extracted and tracked at the same time with GPU-based time-dependent Ginzburg-Landau superconductor simulations. The in situ workflow involves three parts: (1) a tightly coupled GPU-accelerated algorithm that detects primitives for ambiguity-free vortex tracking, (2) a loosely coupled task-parallel feature-tracking method, and (3) a web-based remote visualization tool for vortex dynamics analysis. Our design minimizes the data movement and storage, maximizes the resource utilization, and reduces the slowdown of the simulation. Our solution captures all vortex dynamics in the simulation, previously impossible with traditional post hoc methods. We also demonstrate in situ visualization cases that help scientists understand how vortices cut each other and recombine into new vortices, which are directly related to energy dissipation of superconducting materials. Hanqi Guo 0001, Tom Peterka, Andreas Glatz |
PacificVis | 2 |
| 2017 | Manala: A Flexible Flow Control Library for Asynchronous Task CommunicationabstractTasks coupled in an in situ workflow may not process data at the same speed, potentially causing overflows in the communication channel between them. To prevent this problem, software infrastructures for in situ workflows usually impose a strict FIFO policy that has the side-effect of slowing down faster tasks to the speed of the slower ones. This may not be the desired behavior; for example, a scientist may prefer to drop older data in the communication channel in order to visualize the latest snapshot of a simulation. In this paper, we present Manala, a flexible flow control library designed to manage the flow of messages between a producer and a consumer in an in situ workflow. Manala intercepts messages from the producer, stores them, and selects the message to forward to the consumer depending on the flow control policy. The library is designed to ease the creation of new flow control policies and buffering mechanisms. We demonstrate with three examples how changing the flow control policy between tasks can influence the performance and results of scientific workflows. The first example focuses on materials science with LAMMPS and a synthetic diffraction analysis code. The second example is an interactive visualization scenario with Gromacs as the producer and Damaris/Viz as consumer. Our third example studies different strategies to perform an asynchronous checkpoint with Gromacs. Matthieu Dreher, Kiran Sasikumar, Subramanian Sankaranarayanan, Tom Peterka |
CLUSTER | 4 |
| 2017 | Detection of Silent Data Corruption in Adaptive Numerical Integration SolversabstractScientific computing requires trust in results. In high-performance computing, trust is impeded by silent data corruption (SDC), in other words corruption that remains unnoticed. Numerical integration solvers are especially sensitive to SDCs because an SDC introduced in a certain step affects all the following steps. SDCs can even cause the solver to become unstable. Adaptive solvers can change the step size, by comparing an estimation of the approximation error with an user-defined tolerance. If the estimation exceeds the tolerance, the step is rejected and recomputed. Adaptive solvers have an inherent resilience, because some SDCs might have no consequences on the accuracy of the results, and some SDCs might push the approximation error beyond the tolerance. Our first contribution shows that the rejection mechanism is not reliable enough to reject all SDCs that affect the results' accuracy, because the estimation is also corrupted. We therefore provide another protection mechanism: at the end of each step, a second error estimation is employed to increase the redundancy. Because of the complex dynamics, the choice of the second estimate is difficult: two methods are explored. We evaluated them in HyPar and PETSc, on a cluster of 4,096 cores. We injected SDCs that are large enough to affect the trust or the convergence of the solvers. The new approach can detect 99% of the SDCs, reducing by more than 10 times the number of undetected SDCs. Compared with replication, a classic SDC detector, our protection mechanism reduces the memory overhead by more than 2 times and the computational overhead by more than 20 times in our experiments. Pierre-Louis Guhur, Emil M. Constantinescu, Debojyoti Ghosh, Tom Peterka, Franck Cappello |
CLUSTER | 4 |
| 2017 | Automatic Data Filtering for In Situ WorkflowsabstractIn situ workflows contain tasks that exchange messages composed of several data fields. However, a consumer task may not necessarily need all the data fields from its producer. For example, a molecular dynamics simulation can produce atom positions, velocities, and forces; but some analyses require only atom positions. The user should decide whether to specialize the output of a producer task for a particular consumer and get better performance or to send more data than required by the consumer. The first option limits task portability, while the second wastes resources. In this paper, we introduce contracts for in situ tasks. A contract specifies for a producer each data field available for output and for a consumer the data fields needed as input. Comparing a producer and consumer contract allows automatic selection of the data fields a producer has to send for that consumer. We integrated our contracts mechanism within Decaf, a middleware for building and executing in situ workflows. Contracts enable to automatically extract at the producer the data the consumer needs. We evaluate the cost and performance of message extraction at runtime with both synthetic examples and a real scientific workflow coupling a molecular dynamics simulation with three different data analytics codes. Our contract-based automatic data extraction removes the need to specialize producers while entailing small overheads. Clément Mommessin, Matthieu Dreher, Bruno Raffin, Tom Peterka |
CLUSTER | 4 |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo |
Euro-Par | 22 |
| 2016 | Adaptive Performance-Constrained In Situ Visualization of Atmospheric SimulationsabstractWhile many parallel visualization tools now provide in situ visualization capabilities, the trend has been to feed such tools with large amounts of unprocessed output data and let them render everything at the highest possible resolution. This leads to an increased run time of simulations that still have to complete within a fixed-length job allocation. In this paper, we tackle the challenge of enabling in situ visualization under performance constraints. Our approach shuffles data across processes according to its content and filters out part of it in order to feed a visualization pipeline with only a reorganized subset of the data produced by the simulation. Our framework leverages fast, generic evaluation procedures to score blocks of data, using information theory, statistics, and linear algebra. It monitors its own performance and adapts dynamically to achieve appropriate visual fidelity within predefined performance constraints. Experiments on the Blue Waters supercomputer with the CM1 simulation show that our approach enables a 5x speedup with respect to the initial visualization pipeline and is able to meet performance constraints. Matthieu Dorier, Robert Sisneros, Leonardo Arturo Bautista-Gomez, Tom Peterka, Leigh Orf, Lokman Rahmani, Gabriel Antoniu, Luc Bougé |
CLUSTER | 4 |
| 2016 | Bredala: Semantic Data Redistribution for In Situ ApplicationsabstractIn situ processing is a promising solution to the problem of imbalance between computational capabilities and I/O bandwidth in current and future supercomputers. Initially designed for staging I/O, in situ middleware now can support a wide range of domains such as visualization, machine learning, filtering, and feature tracking. Doing so requires in situ middleware to manage complex heterogeneous codes using different data structures. Data need to be transformed and reorganized along the data path to fit the analysis needs. However, redistributing complex data structures is difficult. In many cases, arbitrarily splitting the arrays of a data structure destroys the semantic integrity of the data. We present Bredala, a lightweight library to annotate a data model with enough information to preserve the semantic integrity of the data during a redistribution. Bredala allows developers to describe how to split and merge a data model safely, operations usually done by in situ middleware. We evaluate the cost and performance of our library in a molecular dynamics application. We show that our data model can simplify the workflow graph of large-scale applications, improve the reusability of tasks, and offer an efficient alternative to redistribute the data. Matthieu Dreher, Tom Peterka |
CLUSTER | 2 |
| 2016 | Parallel DTFE Surface Density Field ReconstructionabstractWe improve the interpolation accuracy and efficiency of the Delaunay tessellation field estimator (DTFE) for surface density field reconstruction by proposing an algorithm that takes advantage of the adaptive triangular mesh for line-of-sight integration. The costly computation of an intermediate 3D grid is completely avoided by our method and only optimally chosen interpolation points are computed, thus, the overall computational cost is significantly reduced. The algorithm is implemented as a parallel shared-memory kernel for large-scale grid rendered field reconstructions in our distributed-memory framework designed for N-body gravitational lensing simulations in large volumes. We also introduce a load balancing scheme to optimize the efficiency of processing a large number of field reconstructions. Our results show our kernel outperforms existing software packages for volume weighted density field reconstruction, achieving~10x speedup, and our load balancing algorithm gains an additional~3.6x speedup at scales with~16k processes. Esteban Rangel, Nan Li 0022, Salman Habib 0002, Tom Peterka, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
CLUSTER | 4 |
| 2016 | Lightweight and Accurate Silent Data Corruption Detection in Ordinary Differential Equation Solvers
Pierre-Louis Guhur, Hong Zhang 0006, Tom Peterka, Emil M. Constantinescu, Franck Cappello |
Euro-Par | 3 |
| 2016 | Efficient delaunay tessellation through K-D tree decompositionabstractDelaunay tessellations are fundamental data structures in computational geometry. They are important in data analysis, where they can represent the geometry of a point set or approximate its density. The algorithms for computing these tessellations at scale perform poorly when the input data is unbalanced. We investigate the use of k-d trees to evenly distribute points among processes and compare two strategies for picking split points between domain regions. Because resulting point distributions no longer satisfy the assumptions of existing parallel Delaunay algorithms, we develop a new parallel algorithm that adapts to its input and prove its correctness. We evaluate the new algorithm using two late-stage cosmology datasets. The new running times are up to 50 times faster using k-d tree compared with regular grid decomposition. Moreover, in the unbalanced data sets, decomposing the domain into a k-d tree is up to five times faster than decomposing it into a regular grid. Dmitriy Morozov, Tom Peterka |
SC | 2 |
| 2016 | Mining Graphs for Understanding Time-Varying Volumetric DataabstractA notable recent trend in time-varying volumetric data analysis and visualization is to extract data relationships and represent them in a low-dimensional abstract graph view for visual understanding and making connections to the underlying data. Nevertheless, the ever-growing size and complexity of data demands novel techniques that go beyond standard brushing and linking to allow significant reduction of cognition overhead and interaction cost. In this paper, we present a mining approach that automatically extracts meaningful features from a graph-based representation for exploring time-varying volumetric data. This is achieved through the utilization of a series of graph analysis techniques including graph simplification, community detection, and visual recommendation. We investigate the most important transition relationships for time-varying data and evaluate our solution with several time-varying data sets of different sizes and characteristics. For gaining insights from the data, we show that our solution is more efficient and effective than simply asking users to extract relationships via standard interaction techniques, especially when the data set is large and the relationships are complex. We also collect expert feedback to confirm the usefulness of our approach. Chaoli Wang 0001, Tom Peterka, Robert L. Jacob |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2016 | Finite-Time Lyapunov Exponents and Lagrangian Coherent Structures in Uncertain Unsteady FlowsabstractThe objective of this paper is to understand transport behavior in uncertain time-varying flow fields by redefining the finite-time Lyapunov exponent (FTLE) and Lagrangian coherent structure (LCS) as stochastic counterparts of their traditional deterministic definitions. Three new concepts are introduced: the distribution of the FTLE (D-FTLE), the FTLE of distributions (FTLE-D), and uncertain LCS (U-LCS). The D-FTLE is the probability density function of FTLE values for every spatiotemporal location, which can be visualized with different statistical measurements. The FTLE-D extends the deterministic FTLE by measuring the divergence of particle distributions. It gives a statistical overview of how transport behaviors vary in neighborhood locations. The U-LCS, the probabilities of finding LCSs over the domain, can be extracted with stochastic ridge finding and density estimation algorithms. We show that our approach produces better results than existing variance-based methods do. Our experiments also show that the combination of D-FTLE, FTLE-D, and U-LCS can help users understand transport behaviors and find separatrices in ensemble simulations of atmospheric processes. Hanqi Guo 0001, Tom Peterka, Han-Wei Shen, Scott M. Collis, Jonathan J. Helmus |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2016 | Extracting, Tracking, and Visualizing Magnetic Flux Vortices in 3D Complex-Valued Superconductor Simulation DataabstractWe propose a method for the vortex extraction and tracking of superconducting magnetic flux vortices for both structured and unstructured mesh data. In the Ginzburg-Landau theory, magnetic flux vortices are well-defined features in a complex-valued order parameter field, and their dynamics determine electromagnetic properties in type-II superconductors. Our method represents each vortex line (a 1D curve embedded in 3D space) as a connected graph extracted from the discretized field in both space and time. For a time-varying discrete dataset, our vortex extraction and tracking method is as accurate as the data discretization. We then apply 3D visualization and 2D event diagrams to the extraction and tracking results to help scientists understand vortex dynamics and macroscale superconductor behavior in greater detail than previously possible. Hanqi Guo 0001, Carolyn L. Phillips, Tom Peterka, Dmitry A. Karpeyev, Andreas Glatz |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Efficient Range Distribution Query for Visualizing Scientific DataabstractVisualization applications implicitly run queries on the data to retrieve distributions and statistical measures derivable from distributions. Distribution based data summaries can substitute for the raw data to answer statistical queries of different kinds. However, frequent access to the raw data is no longer practical, if possible at all, for answering large number of queries on large-scale data. Our work addresses the issue by accelerating range distribution query, which returns the distribution of an axis-aligned query region. Maintaining the interactivity of such query is a challenging task because the workload and the response time of such queries scale up with the data and the query size. In this paper, we present a framework for answering range distribution queries for any arbitrary region in near constant time, regardless of data and query size. We adapt an integral histogram based data structure to bound the workload which is a combination of computation, I/O and communication cost. We propose two novel transformations of this data structure -- a decomposition and a similarity-driven indexing -- to reduce the huge storage cost associated with it. In addition to studying the performance of range distribution query, we also demonstrate the benefits that our technique offers to visualization applications which directly or indirectly require distributions. Abon Chaudhuri, Tzu-Hsuan Wei, Teng-Yok Lee, Han-Wei Shen, Tom Peterka |
PacificVis | 5 |
| 2014 | Scalable Computation of Stream Surfaces on Large Scale Vector FieldsabstractStream surfaces and streamlines are two popular methods for visualizing three-dimensional flow fields. While several parallel streamline computation algorithms exist, relatively little research has been done to parallelize stream surface generation. This is because load-balanced parallel stream surface computation is nontrivial, due to the strong dependency in computing the positions of the particles forming the stream surface front. In this paper, we present a new algorithm that computes stream surfaces efficiently. In our algorithm, seeding curves are divided into segments, which are then assigned to the processes. Each process is responsible for integrating the segments assigned to it. To ensure a balanced computational workload, work stealing and dynamic refinement of seeding curve segments are employed to improve the overall performance. We demonstrate the effectiveness of our parallel stream surface algorithm using several large scale flow field data sets, and show the performance and scalability on HPC systems. Kewei Lu, Han-Wei Shen, Tom Peterka |
SC | 3 |
| 2014 | High-Performance Computation of Distributed-Memory Parallel 3D Voronoi and Delaunay TessellationabstractComputing a Voronoi or Delaunay tessellation from a set of points is a core part of the analysis of many simulated and measured datasets: N-body simulations, molecular dynamics codes, and LIDAR point clouds are just a few examples. Such computational geometry methods are common in data analysis and visualization, but as the scale of simulations and observations surpasses billions of particles, the existing serial and shared memory algorithms no longer suffice. A distributed-memory scalable parallel algorithm is the only feasible approach. The primary contribution of this paper is a new parallel Delaunay and Voronoi tessellation algorithm that automatically determines which neighbor points need to be exchanged among the sub domains of a spatial decomposition. Other contributions include periodic and wall boundary conditions, comparison of our method using two popular serial libraries, and application to numerous science datasets. Tom Peterka, Dmitriy Morozov, Carolyn L. Phillips |
SC | 1 |
| 2014 | Processing MPI Derived Datatypes on Noncontiguous GPU-Resident DataabstractDriven by the goals of efficient and generic communication of noncontiguous data layouts in GPU memory, for which solutions do not currently exist, we present a parallel, noncontiguous data-processing methodology through the MPI datatypes specification. Our processing algorithm utilizes a kernel on the GPU to pack arbitrary noncontiguous GPU data by enriching the datatypes encoding to expose a fine-grained, data-point level of parallelism. Additionally, the typically tree-based datatype encoding is preprocessed to enable efficient, cached access across GPU threads. Using CUDA, we show that the computational method outperforms DMA-based alternatives for several common data layouts as well as more complex data layouts for which reasonable DMA-based processing does not exist. Our method incurs low overhead for data layouts that closely match best-case DMA usage or that can be processed by layout-specific implementations. We additionally investigate usage scenarios for data packing that incur resource contention, identifying potential pitfalls for various packing strategies. We also demonstrate the efficacy of kernel-based packing in various communication scenarios, showing multifold improvement in point-to-point communication and evaluating packing within the context of the SHOC stencil benchmark and HACC mesh analysis. John Jenkins, James Dinan, Pavan Balaji, Tom Peterka, Nagiza F. Samatova, Rajeev Thakur |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Dataflow coordination of data-parallel tasks via MPI 3.0abstractScientific applications are often complex collections of many large-scale tasks. Mature tools exist for describing task-parallel workflows consisting of serial tasks, and a variety of tools exist for programming a single data-parallel operation. However, few tools cover the intersection of these two models. In this work, we extend the load balancing library ADLB to support parallel tasks. We demonstrate how applications can easily be composed of parallel tasks using Swift dataflow scripts, which are compiled to ADLB programs with performance comparable to hand-coded equivalents. By combining this framework with data-parallel analysis libraries, we are able to dynamically execute many instances of a parallel data analysis application in support of a parameter exploration workload. Justin M. Wozniak, Tom Peterka, Timothy G. Armstrong, James Dinan, Ewing L. Lusk, Michael Wilde, Ian T. Foster |
EuroMPI | 2 |
| 2013 | Scalable Resolution Display WallsabstractThis article will describe the progress since 2000 on research and development in 2-D and 3-D scalable resolution display walls that are built from tiling individual lower resolution flat panel displays. The article will describe approaches and trends in display hardware construction, middleware architecture, and user-interaction design. The article will also highlight examples of use cases and the benefits the technology has brought to their respective disciplines. Jason Leigh, Andrew E. Johnson 0001, Luc Renambot, Tom Peterka, Byungil Jeong, Dan Sandin, Jonas Talandis, Ratko Jagodic, Sungwon Nam, Hyejung Hur |
Proc. IEEE | 4 |
| 2012 | The Parallel Computation of Morse-Smale ComplexesabstractTopology-based techniques are useful for multiscale exploration of the feature space of scalar-valued functions, such as those derived from the output of large-scale simulations. The Morse-Smale (MS) complex, in particular, allows robust identification of gradient-based features, and therefore is suitable for analysis tasks in a wide range of application domains. In this paper, we develop a two-stage algorithm to construct the 1-skeleton of the Morse-Smale complex in parallel, the first stage independently computing local features per block and the second stage merging to resolve global features. Our implementation is based on MPI and a distributed-memory architecture. Through a set of scalability studies on the IBM Blue Gene/P supercomputer, we characterize the performance of the algorithm as block sizes, process counts, merging strategy, and levels of topological simplification are varied, for datasets that vary in feature composition and size. We conclude with a strong scaling study using scientific datasets computed by combustion and hydrodynamics simulations. Attila Gyulassy, Valerio Pascucci, Tom Peterka, Robert B. Ross |
IPDPS | 3 |
| 2012 | Versatile Communication Algorithms for Data Analysis
Tom Peterka, Robert B. Ross |
EuroMPI | 1 |
| 2012 | The universe at extreme scale: multi-petaflop sky simulation on the BG/QabstractRemarkable observational advances have established a compelling cross-validated model of the Universe. Yet, two key pillars of this model -- dark matter and dark energy -- remain mysterious. Next-generation sky surveys will map billions of galaxies to explore the physics of the 'Dark Universe'. Science requirements for these surveys demand simulations at extreme scales; these will be delivered by the HACC (Hybrid/Hardware Accelerated Cosmology Code) framework. HACC's novel algorithmic structure allows tuning across diverse architectures, including accelerated and multi-core systems. On the IBM BG/Q, HACC attains unprecedented scalable performance - currently 6.23 PFlops at 62% of peak and 92% parallel efficiency on 786,432 cores (48 racks) - at extreme problem sizes with up to almost two trillion particles, larger than any cosmological simulation yet performed. HACC simulations at these scales will for the first time enable tracking individual galaxies over the entire volume of a cosmological survey. Salman Habib 0002, Vitali A. Morozov, Hal Finkel, Adrian Pope, Katrin Heitmann, Kalyan Kumaran, Tom Peterka, Joseph A. Insley, David Daniel, Patricia K. Fasel, Nicholas Frontiere, Zarija Lukic |
SC | 7 |
| 2012 | Parallel particle advection and FTLE computation for time-varying flow fieldsabstractFlow fields are an important product of scientific simulations. One popular flow visualization technique is particle advection, in which seeds are traced through the flow field. One use of these traces is to compute a powerful analysis tool called the Finite-Time Lyapunov Exponent (FTLE) field, but no existing particle tracing algorithms scale to the particle injection frequency required for high-resolution FTLE analysis. In this paper, a framework to trace the massive number of particles necessary for FTLE computation is presented. A new approach is explored, in which processes are divided into groups, and are responsible for mutually exclusive spans of time. This pipelining over time intervals reduces overall idle time of processes and decreases I/O overhead. Our parallel FTLE framework is capable of advecting hundreds of millions of particles at once, with performance scaling up to tens of thousands of processes. Boonthanome Nouanesengsy, Teng-Yok Lee, Kewei Lu, Han-Wei Shen, Tom Peterka |
SC | 5 |
| 2011 | A Study of Parallel Particle Tracing for Steady-State and Time-Varying Flow FieldsabstractParticle tracing for streamline and path line generation is a common method of visualizing vector fields in scientific data, but it is difficult to parallelize efficiently because of demanding and widely varying computational and communication loads. In this paper we scale parallel particle tracing for visualizing steady and unsteady flow fields well beyond previously published results. We configure the 4D domain decomposition into spatial and temporal blocks that combine in-core and out-of-core execution in a flexible way that favors faster run time or smaller memory. We also compare static and dynamic partitioning approaches. Strong and weak scaling curves are presented for tests conducted on an IBM Blue Gene/P machine at up to 32 K processes using a parallel flow visualization library that we are developing. Datasets are derived from computational fluid dynamics simulations of thermal hydraulics, liquid mixing, and combustion. Tom Peterka, Robert B. Ross, Boonthanome Nouanesengsy, Teng-Yok Lee, Han-Wei Shen, Wesley Kendall, Jian Huang 0007 |
IPDPS | 1 |
| 2011 | Simplified parallel domain traversalabstractMany data-intensive scientific analysis techniques require global domain traversal, which over the years has been a bottleneck for efficient parallelization across distributed-memory architectures. Inspired by MapReduce and other simplified parallel programming approaches, we have designed DStep, a flexible system that greatly simplifies efficient parallelization of domain traversal techniques at scale. In order to deliver both simplicity to users as well as scalability on HPC platforms, we introduce a novel two-tiered communication architecture for managing and exploiting asynchronous communication loads. We also integrate our design with advanced parallel I/O techniques that operate directly on native simulation output. We demonstrate DStep by performing teleconnection analysis across ensemble runs of terascale atmospheric CO2 and climate data, and we show scalability results on up to 65,536 IBM BlueGene/P cores. Wesley Kendall, Melissa R. Dumas, Tom Peterka, Jian Huang 0007, David Erickson |
SC | 4 |
| 2011 | An image compositing solution at scaleabstractThe only proven method for performing distributed-memory parallel rendering at large scales, tens of thousands of nodes, is a class of algorithms called sort last. The fundamental operation of sort-last parallel rendering is an image composite, which combines a collection of images generated independently on each node into a single blended image. Over the years numerous image compositing algorithms have been proposed as well as several enhancements and rendering modes to these core algorithms. However, the testing of these image compositing algorithms has been with an arbitrary set of enhancements, if any are applied at all. In this paper we take a leading production-quality image-compositing framework, IceT, and use it as a testing framework for the leading image compositing algorithms of today. As we scale IceT to ever increasing job sizes, we consider the image compositing systems holistically, incorporate numerous optimizations, and discover several improvements to the process never considered before. We conclude by demonstrating our solution on 64K cores of the Intrepid Blue-Gene/P at Argonne National Laboratories. Kenneth Moreland, Wesley Kendall, Tom Peterka, Jian Huang 0007 |
SC | 3 |
| 2009 | End-to-End Study of Parallel Volume Rendering on the IBM Blue Gene/PabstractIn addition to their role as simulation engines, modern supercomputers can be harnessed for scientific visualization. Their extensive concurrency, parallel storage systems, and high-performance interconnects can mitigate the expanding size and complexity of scientific datasets and prepare for in situ visualization of these data. In ongoing research into testing parallel volume rendering on the IBM Blue Gene/P (BG/P), we measure performance of disk I/O, rendering, and compositing on large datasets, and evaluate bottlenecks with respect to system-specific I/O and communication patterns. To extend the scalability of the direct-send image compositing stage of the volume rendering algorithm, we limit the number of compositing cores when many small messages are exchanged. To improve the data-loading stage of the volume renderer, we study the I/O signatures of the algorithm in detail. The results of this research affirm that a distributed-memory computing architecture such as BG/P is a scalable platform for large visualization problems. Tom Peterka, Hongfeng Yu 0001, Robert B. Ross, Kwan-Liu Ma, Robert Latham |
ICPP | 1 |
| 2009 | Terascale data organization for discovering multivariate climatic trendsabstractCurrent visualization tools lack the ability to perform full-range spatial and temporal analysis on terascale scientific datasets. Two key reasons exist for this shortcoming: I/O and postprocessing on these datasets are being performed in suboptimal manners, and the subsequent data extraction and analysis routines have not been studied in depth at large scales. We resolved these issues through advanced I/O techniques and improvements to current query-driven visualization methods. We show the efficiency of our approach by analyzing over a terabyte of multivariate satellite data and addressing two key issues in climate science: time-lag analysis and drought assessment. Our methods allowed us to reduce the end-to-end execution times on these problems to one minute on a Cray XT4 machine. Wesley Kendall, Markus Glatter, Jian Huang 0007, Tom Peterka, Robert Latham, Robert B. Ross |
SC | 4 |
| 2009 | A configurable algorithm for parallel image-compositing applicationsabstractCollective communication operations can dominate the cost of large-scale parallel algorithms. Image compositing in parallel scientific visualization is a reduction operation where this is the case. We present a new algorithm called Radix-k that in many cases performs better than existing compositing algorithms. It does so through a set of configurable parameters, the radices, that determine the number of communication partners in each message round. The algorithm embodies and unifies binary swap and direct-send, two of the best-known compositing methods, and enables numerous other configurations through appropriate choices of radices. While the algorithm is not tied to a particular computing architecture or network topology, the selection of radices allows Radix-k to take advantage of new supercomputer interconnect features such as multiporting. We show scalability across image size and system size, including both powers of two and nonpowers-of-two process counts. Tom Peterka, David Goodell, Robert B. Ross, Han-Wei Shen, Rajeev Thakur |
SC | 1 |
| 2008 | Advances in the Dynallax Solid-State Dynamic Parallax Barrier Autostereoscopic Visualization Display SystemabstractA solid-state dynamic parallax barrier autostereoscopic display mitigates some of the restrictions present in static barrier systems, such as fixed view-distance range, slow response to head movements, and fixed stereo operating mode. By dynamically varying barrier parameters in real time, viewers may move closer to the display and move faster laterally than with a static barrier system, and the display can switch between 3D and 2D modes by disabling the barrier on a per-pixel basis. Moreover, Dynallax can output four independent eye channels when two viewers are present, and both head-tracked viewers receive an independent pair of left-eye and right-eye perspective views based on their position in 3D space. The display device is constructed by using a dual-stacked LCD monitor where a dynamic barrier is rendered on the front display and a modulated virtual environment composed of two or four channels is rendered on the rear display. Dynallax was recently demonstrated in a small-scale head-tracked prototype system. This paper summarizes the concepts presented earlier, extends the discussion of various topics, and presents recent improvements to the system. Tom Peterka, Robert Kooima, Dan Sandin, Andrew E. Johnson 0001, Jason Leigh, Thomas A. DeFanti |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2007 | A GPU Sub-pixel Algorithm for Autostereoscopic Virtual RealityabstractAutostereoscopic displays enable unencumbered immersive virtual reality, but at a significant computational expense. This expense impacts the feasibility of autostereo displays in high-performance real-time interactive applications. A new autostereo rendering algorithm named autostereo combiner addresses this problem using the programmable vertex and fragment pipelines of modern graphics processing units (GPUs). This algorithm is applied to the Varrier, a large-scale, head-tracked, parallax barrier autostereo virtual reality platform. In this capacity, the Combiner algorithm has shown performance gains of 4x over traditional parallax barrier rendering algorithms. It has enabled high-performance rendering at sub-pixel scales, affording a 2x increase in resolution and showing a 1.4x improvement in visual acuity Robert Kooima, Tom Peterka, Javier Girado, Jinghua Ge, Dan Sandin, Thomas A. DeFanti |
VR | 2 |
| 2007 | Dynallax: Solid State Dynamic Parallax Barrier Autostereoscopic VR DisplayabstractA novel barrier strip autostereoscopic (AS) display is demonstrated using a solid-state dynamic parallax barrier. A dynamic barrier mitigates restrictions inherent in static barrier systems such as fixed view distance range, slow response to head movements, and fixed stereo operating mode. By dynamically varying barrier parameters in real time, viewers may move closer to the display and move faster laterally than with a static barrier system. Furthermore, users can switch between 3D and 2D modes by disabling the barrier. Dynallax is head-tracked, directing view channels to positions in space reported by a tracking system in real time. Such head-tracked parallax barrier systems have traditionally supported only a single viewer, but by varying the barrier period to eliminate conflicts between viewers, Dynallax presents four independent eye channels when two viewers are present. Each viewer receives an independent pair of left and right eye perspective views based on their position in 3D space. The display device is constructed using a dual-stacked LCD monitor where a dynamic barrier is rendered on the front display and the rear display produces a modulated VR scene composed of two or four channels. A small-scale head-tracked prototype VR system is demonstrated. Tom Peterka, Robert Kooima, Javier Girado, Jinghua Ge, Dan Sandin, Andrew E. Johnson 0001, Jason Leigh, Jürgen P. Schulze, Thomas A. DeFanti |
VR | 1 |
| 2006 | The global lambda visualization facility: An international ultra-high-definition wide-area visualization collaboratory
Jason Leigh, Luc Renambot, Andrew E. Johnson 0001, Byungil Jeong, Ratko Jagodic, Nicholas Schwarz, Dmitry Svistula, Rajvikram Singh, Julieta Aguilera, Venkatram Vishwanath, Brenda Lopez, Dan Sandin, Tom Peterka, Javier Girado, Robert Kooima, Jinghua Ge, Lance Long, Alan Verlo, Thomas A. DeFanti, Maxine D. Brown, Donna J. Cox, Robert Patterson, Patrick Dorn, Paul Wefel, Stuart Levy, Jonas Talandis, Joe Reitzer, Tom Prudhomme, Tom Coffin, Paul Wielinga, Bram Stolk, Gee Bum Koo, Jaeyoun Kim, Sangwoo Han, Jongwon Kim 0001, Brian Corrie II, Todd Zimmerman, Pierre Boulanger, Manuel Garcia |
Future Gener. Comput. Syst. | 14 |
| 2006 | Personal Varrier: Autostereoscopic virtual reality display for distributed scientific visualization
Tom Peterka, Dan Sandin, Jinghua Ge, Javier Girado, Robert Kooima, Jason Leigh, Andrew E. Johnson 0001, Marcus Thiébaux, Thomas A. DeFanti |
Future Gener. Comput. Syst. | 1 |
| 2005 | The VarrierTM autostereoscopic virtual reality displayabstractVirtual reality (VR) has long been hampered by the gear needed to make the experience possible; specifically, stereo glasses and tracking devices. Autostereoscopic display devices are gaining popularity by freeing the user from stereo glasses, however few qualify as VR displays. The Electronic Visualization Laboratory (EVL) at the University of Illinois at Chicago (UIC) has designed and produced a large scale, high resolution head-tracked barrier-strip autostereoscopic display system that produces a VR immersive experience without requiring the user to wear any encumbrances. The resulting system, called Varrier, is a passive parallax barrier 35-panel tiled display that produces a wide field of view, head-tracked VR experience. This paper presents background material related to parallax barrier autostereoscopy, provides system configuration and construction details, examines Varrier interleaving algorithms used to produce the stereo images, introduces calibration and testing, and discusses the camera-based tracking subsystem. Dan Sandin, Todd Margolis, Jinghua Ge, Javier Girado, Tom Peterka, Thomas A. DeFanti |
ACM Trans. Graph. | 5 |