VLDB 2026 Research / reviewers in the wild / expert
Peer-Timo Bremer
dblp:20/3591 · also Timo Bremer
· DBLP profile ↗
121ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0003-4107-3831ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 85 · 4 first-author · 23 since 2021Systems, architecture and hardware · 26 · 5 since 2021Artificial intelligence and machine learning · 11 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Alignment and Comparison of Hierarchical Morse Complexes with ApplicationsabstractAbstract Scalar fields derived from 3D X‐ray CT scans of samples undergoing ex situ processes, such as thermal aging, chemical etching, or mechanical stress, pose unique challenges for characterizing similarities and differences across acquisitions. Typically, a sample A (source) is imaged, removed, and subjected to experimental conditions that alter its microstructure, and then re‐imaged as sample B (target) to study the resulting changes. Direct comparison between A and B is rendered impractical if not impossible for current techniques because the challenges of physical and morphological changes are compounded by the effects of geometric misalignment, differences in reconstruction parameters, discretization artifacts, and changes in acquisition settings such as position, beam intensity, or exposure time (the acquisition for sample B often happens at a much later time and the device may have been upgraded or changed). To overcome these challenges, we introduce a geometry‐rich topological representation that uses the hierarchical Morse complex to capture the structural relationships among regions segmented within each sample and shape descriptors to characterize their metric properties. With this data structure, we cast the similarity problem as a sequence of optimizations, each minimizing differences in structure and geometry at a given resolution. The sequence of optimizations begins by aligning the fine‐scale segmentations of the two samples. Following optimization minimizes differences across incrementally coarser levels, producing a fully synchronized hierarchical representation of the two samples. In addition, we introduce a visualization framework that enables interactive exploration and manual editing of the matched hierarchies, thereby allowing an expert user to further improve the quality of the comparison. We apply our workflow to characterize changes in grain structure for energetic materials undergoing aging, match segmentations for materials under different stress conditions, and perform image registration that outperforms state‐of‐the‐art techniques. Aniketh Venkat, Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
Comput. Graph. Forum | 3 |
| 2026 | Re-Evaluating Virtual Reality Manipulation Techniques for Precise Alignment of Complex 3D ObjectsabstractPrior research has developed a number of manipulation techniques that can achieve precise object placement in virtual reality, but studies of these techniques typically use simple objects. We conducted a study comparing two existing techniques, (AMP-IT and WISDOM), during alignment of objects with complex geometry to evaluate the potential influence of geometric complexity on performance, usability, workload and preference. Our findings indicate that participants had faster completion times and higher trial completion rates with AMP-IT on high-precision alignment tasks, contrary to earlier findings that used simple objects. Yet WISDOM is still preferred and considered more usable, despite increased workload and poorer performance, exposing participants' willingness to trade objective performance for comfort during use. Cherelle Connor, Alexander Giovannelli, Leonardo Pavanatto, Francielly Rodrigues, Haichao Miao, Vuthea Chheang, Brian Giera, Peer-Timo Bremer, Doug A. Bowman |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | Investigating the Influence of Playback Interactivity during Guided Tours for Asynchronous Collaboration in Virtual RealityabstractCollaborative virtual environments allow workers to contribute to team projects across space and time. While much research has closely examined the problem of working in different spaces at the same time, few have investigated the best practices for collaborating in those spaces at different times aside from textual and auditory annotations. We designed a system that allows experts to record a tour inside a virtual inspection space, preserving knowledge and providing later observers with insights through a 3D playback of the expert’s inspection. We also created several interactions to ensure that observers are tracking the tour and remaining engaged. We conducted a user study to evaluate the influence of these interactions on an observing user’s information recall and user experience. Findings indicate that independent viewpoint control during a tour enhances the user experience compared to fully passive playback and that additional interactivity can improve auditory and spatial recall of key information conveyed during the tour. Alexander Giovannelli, Leonardo Pavanatto, Shakiba Davari, Haichao Miao, Vuthea Chheang, Brian Giera, Peer-Timo Bremer, Doug A. Bowman |
VR | 7 |
| 2025 | Exploring Multiscale Navigation of Homogeneous and Dense Objects with Progressive Refinement in Virtual RealityabstractLocating small features in a large, dense object in virtual reality (VR) poses a significant interaction challenge. While existing multiscale techniques support transitions between various levels of scale, they are not focused on handling dense, homogeneous objects with hidden features. We propose a novel approach that applies the concept of progressive refinement to VR navigation, enabling focused inspections. We conducted a user study where we varied two independent variables in our design, navigation style (STRUCTURED vs. UNSTRUCTURED) and display mode (SELECTION vs. EVERYTHING), to better understand their effects on efficiency and awareness during multiscale navigation. Our results showed that unstructured navigation can be faster than structured and that displaying only the selection can be faster than displaying the entire object. However, using an everything display mode can support better location awareness and object understanding. Leonardo Pavanatto, Alexander Giovannelli, Brian Giera, Peer-Timo Bremer, Haichao Miao, Doug A. Bowman |
VR | 4 |
| 2025 | Exploring Bichronous Collaboration in Virtual EnvironmentsabstractVirtual environments (VEs) empower geographically distributed teams to collaborate on a shared project regardless of time. Existing research has separately investigated collaborations within these VEs at the same time (i.e., synchronous) or different times (i.e., asynchronous). In this work, we highlight the often-overlooked concept of bichronous collaboration and define it as the seamless integration of archived information during a real-time collaborative session. We revisit the time-space matrix of computer-supported cooperative work (CSCW) and reclassify the time dimension as a continuum. We describe a system that empowers collaboration across the temporal states of the time continuum within a VE during remote work. We conducted a user study using the system to discover how the bichronous temporal state impacts the user experience during a collaborative inspection. Findings indicate that the bichronous temporal state is beneficial to collaborative activities for information processing, but has drawbacks such as changed interaction and positioning behaviors in the VE. Alexander Giovannelli, Shakiba Davari, Cherelle Connor, Fionn Murphy, Trey Davis, Haichao Miao, Vuthea Chheang, Brian Giera, Peer-Timo Bremer, Doug A. Bowman |
VRST | 9 |
| 2025 | Bimodal Visualization of Industrial X-Ray and Neutron Computed Tomography DataabstractAdvanced manufacturing creates increasingly complex objects with material compositions that are often difficult to characterize by a single modality. Our collaborating domain scientists are going beyond traditional methods by employing both X-ray and neutron computed tomography to obtain complementary representations expected to better resolve material boundaries. However, the use of two modalities creates its own challenges for visualization, requiring either complex adjustments of bimodal transfer functions or the need for multiple views. Together with experts in nondestructive evaluation, we designed a novel interactive bimodal visualization approach to create a combined view of the co-registered X-ray and neutron acquisitions of industrial objects. Using an automatic topological segmentation of the bivariate histogram of X-ray and neutron values as a starting point, the system provides a simple yet effective interface to easily create, explore, and adjust a bimodal visualization. We propose a widget with simple brushing interactions that enables the user to quickly correct the segmented histogram results. Our semiautomated system enables domain experts to intuitively explore large bimodal datasets without the need for either advanced segmentation algorithms or knowledge of visualization techniques. We demonstrate our approach using synthetic examples, industrial phantom objects created to stress bimodal scanning techniques, and real-world objects, and we discuss expert feedback. Xuan Huang 0007, Haichao Miao, Hyojin Kim 0001, Andrew Townsend, Kyle Champley, Joseph W. Tringe, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | "Understanding Robustness Lottery": A Geometric Visual Comparative Analysis of Neural Network Pruning ApproachesabstractDeep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. This work aims to shed light on how different pruning methods alter the network's internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption. Shusen Liu 0001, Xin Yu 0002, Bhavya Kailkhura, Jie Cao 0010, James Diffenderfer, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | LatticeAnalytics: Strut-Level Visualization and Inspection of Additively Manufactured Lattice StructuresabstractAdditive manufacturing (AM) is revolutionizing the production of custom components with complex internal geometries, essential for high-performance applications in diverse fields such as medicine and defense. These AM parts optimize strength while minimizing weight by utilizing internal lattice structures consisting of large quantities of small interconnected struts. However, the complexity of these structures, combined with the challenges of using X-ray Computed Tomography (XCT) data, makes validation of part reliability difficult. This ultimately inhibits the development of novel parts for our collaborating material scientists. We introduce LatticeAnalytics, a novel framework specifically designed for visual inspection of defects in these lattice structures. Our framework offers an end-to-end solution that includes the data management of XCT scans, enables remote access for geographically dispersed teams through a web-based dashboard, and incorporates novel visualizations. Our analysis is facilitated by a coarse alignment between the lattice's nominal model, a spatial graph, and the XCT data. We employ a simple VR-based approach for fast and rough alignment, followed by an offline registration and identification of the struts. With the nodes and struts aligned and identified in the volume, our framework allows querying of subvolumes containing a single strut at multiple resolutions. This avoids computation over the entire lattice and also allow for easy parallelization of down-stream computations, such as strut-specific metrics. To depict a fast overview of the strut quality, we introduce two innovative visual encodings, crucial for our collaborators' research in creating novel AM parts: the Contour View and the Roughness Map, which depict critical geometrical and surface features of individual struts in standardized two 2D views. We evaluated the integrated system through expert interviews. The feedback confirms the framework's practicality and its effectiveness in enhancing current inspection workflows. It solves major bottlenecks for our collaborators, ultimately helping them create novel parts with advanced properties. Haichao Miao, Saurabh Narain, Vuthea Chheang, Garrett Hooten, Raiyan Seede, Pavol Klacansky, Kaila Morgen Bertsch, Gabe Guss, Brian Giera, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2024 | HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific ApplicationsabstractRecent advancements in Machine Learning (ML) have substantially improved its predictive and computational abilities, offering promising opportunities for surrogate modeling in scientific applications. By accurately approximating complex functions with low computational cost, ML-based surrogates can accelerate scientific applications by replacing computationally intensive components with faster model inference. However, integrating ML models into these applications remains a significant challenge, hindering the widespread adoption of ML surrogates as an approximation technique in modern scientific computing. We propose an easy-to-use directive-based programming model that enables developers to seamlessly describe the use of ML models in scientific applications. The runtime support, as instructed by the programming model, performs data assimilation using the original algorithm and can replace the algorithm with model inference. Our evaluation across five benchmarks, testing over 5000 ML models, shows up to $83.6 \times$ speed improvements with minimal accuracy loss (as low as 0.01 RMSE). Zane Fink, Konstantinos Parasyris, Praneet Rathi, Giorgis Georgakoudis, Harshitha Menon, Peer-Timo Bremer |
SC | 6 |
| 2024 | AVA: Towards Autonomous Visualization Agents through Visual Perception-Driven Decision-MakingabstractAbstract With recent advances in multi‐modal foundation models, the previously text‐only large language models (LLM) have evolved to incorporate visual input, opening up unprecedented opportunities for various applications in visualization. Compared to existing work on LLM‐based visualization works that generate and control visualization with textual input and output only, the proposed approach explores the utilization of the visual processing ability of multi‐modal LLMs to develop Autonomous Visualization Agents (AVAs) that can evaluate the generated visualization and iterate on the result to accomplish user‐defined objectives defined through natural language. We propose the first framework for the design of AVAs and present several usage scenarios intended to demonstrate the general applicability of the proposed paradigm. Our preliminary exploration and proof‐of‐concept agents suggest that this approach can be widely applicable whenever the choices of appropriate visualization parameters require the interpretation of previous visual output. Our study indicates that AVAs represent a general paradigm for designing intelligent visualization systems that can achieve high‐level visualization goals, which pave the way for developing expert‐level visualization agents in the future. Shusen Liu 0001, Haichao Miao, Matthew L. Olson, Valerio Pascucci, Peer-Timo Bremer |
Comput. Graph. Forum | 6 |
| 2024 | A Visual Comparison of Silent Error PropagationabstractHigh-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. To accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution. Harshitha Menon, Kathryn Mohror, Shusen Liu 0001, Luanzheng Guo, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Cross-GAN Auditing: Unsupervised Identification of Attribute Level Similarities and Differences Between Pretrained Generative ModelsabstractGenerative Adversarial Networks (GANs) are notoriously difficult to train especially for complex distributions and with limited data. This has driven the need for tools to audit trained networks in human intelligible format, for example, to identify biases or ensure fairness. Existing GAN audit tools are restricted to coarse-grained, modeldata comparisons based on summary statistics such as FID or recall. In this paper, we propose an alternative approach that compares a newly developed GAN against a prior baseline. To this end, we introduce Cross-GAN Auditing (xGA) that, given an established “reference” GAN and a newly proposed “client” GAN, jointly identifies intelligible attributes that are either common across both GANs, novel to the client GAN, or missing from the client GAN. This provides both users and model developers an intuitive assessment of similarity and differences between GANs. We introduce novel metrics to evaluate attribute-based GAN auditing approaches and use these metrics to demonstrate quantitatively that xGA outperforms baseline approaches. We also include qualitative results that illustrate the common, novel and missing attributes identified by xGA from GANs trained on a variety of image datasets1 Matthew L. Olson, Shusen Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Weng-Keen Wong |
CVPR | 5 |
| 2023 | Scalable Comparative Visualization of Ensembles of Call GraphsabstractOptimizing the performance of large-scale parallel codes is critical for efficient utilization of computing resources. Code developers often explore various execution parameters, such as hardware configurations, system software choices, and application parameters, and are interested in detecting and understanding bottlenecks in different executions. They often collect hierarchical performance profiles represented as call graphs, which combine performance metrics with their execution contexts. The crucial task of exploring multiple call graphs together is tedious and challenging because of the many structural differences in the execution contexts and significant variability in the collected performance metrics (e.g., execution runtime). In this paper, we present Ensemble CallFlow to support the exploration of ensembles of call graphs using new types of visualizations, analysis, graph operations, and features. We introduce ensemble-Sankey, a new visual design that combines the strengths of resource-flow (Sankey) and box-plot visualization techniques. Whereas the resource-flow visualization can easily and intuitively describe the graphical nature of the call graph, the box plots overlaid on the nodes of Sankey convey the performance variability within the ensemble. Our interactive visual interface provides linked views to help explore ensembles of call graphs, e.g., by facilitating the analysis of structural differences, and identifying similar or distinct call graphs. We demonstrate the effectiveness and usefulness of our design through case studies on large-scale parallel codes. Suraj P. Kesavan, Harsh Bhatia, Abhinav Bhatele, Stephanie Brink, Olga Pearce, Todd Gamblin, Peer-Timo Bremer, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | A Study of the Locality of Persistence-Based Queries and Its Implications for the Efficiency of Localized Data StructuresabstractScientific datasets are often analyzed and visualized using isosurfaces. The connected components at or above the isovalue defining these isosurfaces are called superlevel-set components. The vertex set of these superlevel-set components can be used to compute local statistics, such as mean temperature or histogram per component, or to segment the data. However, in datasets produced by acquisition devices or simulations, noise induces many spurious components that clutter the visualization and analysis results. Many of these spurious components would disappear if the data values were slightly adjusted. The notion of persistence captures the stability of a component with respect to function value changes, and so we are interested in computing persistence quickly. Locality of computation is critical for parallel scalability, minimization of communication in a distributed environment, or an out-of-core processing. The recently introduced merge forest attained high performance by exploiting locality, thereby avoiding communication until needed to resolve a feature query. We extend the merge forest to support persistence-based queries and study the locality of these queries by evaluating the traversals of regions of data during a query. We confirm that the majority of evaluated datasets have the property that the noise is mostly local, and thus can be efficiently eliminated without performing a global analysis. Finally, we compare the query running times with those of a triplet merge tree because a triplet merge tree answers all proposed queries in constant time and can be constructed from a merge tree in linear time. Pavol Klacansky, Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
PacificVis | 3 |
| 2022 | Virtual Inspection of Additively Manufactured PartsabstractAdvanced manufacturing techniques, such as additive manufacturing, enable the design of increasingly complex components for a wide range of industrial applications. However, this complexity makes qualification of the parts, determining whether a part is within some margin of error from the initial design, difficult. To inspect and qualify complex internal geometries that are not accessible with an external probe, parts are typically scanned with computed tomography (CT), and manually compared to the computer-aided design (CAD) model using visual inspections. Matching the CAD model to the 3D reconstructed object is challenging in a traditional desktop environment due to the lack of depth perception and 3D interaction. An additional challenge comes from the geometric complexity of CAD meshes and large-scale CT scans. We present a virtual reality (VR) system for manual qualification, providing a novel defect visualization method. First, we describe a semiautomatic CAD-to-Scan Registration approach in VR using a finite element mesh. Second, we introduce the Defect Box, which enables full-resolution inspection for massive scans and CAD-CT comparison of local defect regions. Finally, our system includes intuitive 3D Metrology methods that enable natural interactions for the measurement of features and defects in VR. We demonstrate our approach on both real and synthetic data and discuss feedback from four expert users in nondestructive qualification. Pavol Klacansky, Haichao Miao, Attila Gyulassy, Andrew Townsend, Kyle Champley, Joseph W. Tringe, Valerio Pascucci, Peer-Timo Bremer |
PacificVis | 8 |
| 2022 | Sparsity Improves Unsupervised Attribute Discovery in StyleganabstractRich semantics exist in latent spaces inferred using deep generative models. The ability to extract and interpret them is not only essential for understanding the underlying factors of variation in the data distribution, but also crucial for con-trolled image generation. Several methods have been proposed to identify semantically meaningful linear directions, either through existing annotations, or relying on identifying directions of large variation that arise from the data representation of the network. In this paper, we identify a new criterion, representation sparsity, that allows us to produce extremely efficient yet diverse semantic directions in GAN (generative adversarial network) latent spaces. The observation also reveals a potential deeper connection between representation sparsity and semantics in deep neural networks that worth further exploration. Shusen Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICASSP | 4 |
| 2022 | Models Out of Line: A Fourier Lens on Distribution Shift RobustnessabstractImproving the accuracy of deep neural networks on out-of-distribution (OOD) data is critical to an acceptance of deep learning in real world applications. It has been observed that accuracies on in-distribution (ID) versus OOD data follow a linear trend and models that outperform this baseline are exceptionally rare (and referred to as ``effectively robust”). Recently, some promising approaches have been developed to improve OOD robustness: model pruning, data augmentation, and ensembling or zero-shot evaluating large pretrained models. However, there still is no clear understanding of the conditions on OOD data and model properties that are required to observe effective robustness. We approach this issue by conducting a comprehensive empirical study of diverse approaches that are known to impact OOD robustness on a broad range of natural and synthetic distribution shifts of CIFAR-10 and ImageNet. In particular, we view the "effective robustness puzzle" through a Fourier lens and ask how spectral properties of both models and OOD data correlate with OOD robustness. We find this Fourier lens offers some insight into why certain robust models, particularly those from the CLIP family, achieve OOD robustness. However, our analysis also makes clear that no known metric is consistently the best explanation of OOD robustness. Thus, to aid future research into the OOD puzzle, we address the gap in publicly-available models with effective robustness by introducing a set of pretrained CIFAR-10 models---$RobustNets$---with varying levels of OOD robustness. Sara Fridovich-Keil, Brian R. Bartoldson, James Diffenderfer, Bhavya Kailkhura, Peer-Timo Bremer |
NeurIPS | 5 |
| 2022 | Single Model Uncertainty Estimation via Stochastic Data CenteringabstractWe are interested in estimating the uncertainties of deep neural networks, which play an important role in many scientific and engineering problems. In this paper, we present a striking new finding that an ensemble of neural networks with the same weight initialization, trained on datasets that are shifted by a constant bias gives rise to slightly inconsistent trained models, where the differences in predictions are a strong indicator of epistemic uncertainties. Using the neural tangent kernel (NTK), we demonstrate that this phenomena occurs in part because the NTK is not shift-invariant. Since this is achieved via a trivial input transformation, we show that this behavior can therefore be approximated by training a single neural network -- using a technique that we call $\Delta-$UQ -- that estimates uncertainty around prediction by marginalizing out the effect of the biases during inference. We show that $\Delta-$UQ's uncertainty estimates are superior to many of the current methods on a variety of benchmarks-- outlier rejection, calibration under distribution shift, and sequential design optimization of black box functions. Code for $\Delta-$UQ can be accessed at github.com/LLNL/DeltaUQ Jayaraman J. Thiagarajan, Rushil Anirudh, Vivek Sivaraman Narayanaswamy, Peer-Timo Bremer |
NeurIPS | 4 |
| 2022 | Enabling machine learning-ready HPC ensembles with Merlin
Jayson Luc Peterson, Benjamin Bay, Joe Koning, Peter B. Robinson, Jessica Semler, Jeremy White, Rushil Anirudh, Kevin Athey, Peer-Timo Bremer, Francesco Di Natale, Jim Gaffney, Sam Ade Jacobs, Bhavya Kailkhura, Bogdan Kustowski, Steve H. Langer, Brian K. Spears, Jayaraman J. Thiagarajan, Brian Van Essen, Jae-Seung Yeom |
Future Gener. Comput. Syst. | 9 |
| 2022 | AMM: Adaptive Multilinear MeshesabstractAdaptive representations are increasingly indispensable for reducing the in-memory and on-disk footprints of large-scale data. Usual solutions are designed broadly along two themes: reducing data precision, e.g., through compression, or adapting data resolution, e.g., using spatial hierarchies. Recent research suggests that combining the two approaches, i.e., adapting both resolution and precision simultaneously, can offer significant gains over using them individually. However, there currently exist no practical solutions to creating and evaluating such representations at scale. In this work, we present a new resolution-precision-adaptive representation to support hybrid data reduction schemes and offer an interface to existing tools and algorithms. Through novelties in spatial hierarchy, our representation, Adaptive Multilinear Meshes (AMM), provides considerable reduction in the mesh size. AMM creates a piecewise multilinear representation of uniformly sampled scalar data and can selectively relax or enforce constraints on conformity, continuity, and coverage, delivering a flexible adaptive representation. AMM also supports representing the function using mixed-precision values to further the achievable gains in data reduction. We describe a practical approach to creating AMM incrementally using arbitrary orderings of data and demonstrate AMM on six types of resolution and precision datastreams. By interfacing with state-of-the-art rendering tools through VTK, we demonstrate the practical and computational advantages of our representation for visualization techniques. With an open-source release of our tool to create AMM, we make such evaluation of data reduction accessible to the community, which we hope will foster new opportunities and future data reduction schemes. Harsh Bhatia, Duong Hoang, Nathan Morrical, Valerio Pascucci, Peer-Timo Bremer, Peter Lindstrom 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Towards replacing physical testing of granular materials with a Topology-based ModelabstractIn the study of packed granular materials, the performance of a sample (e.g., the detonation of a high-energy explosive) often correlates to measurements of a fluid flowing through it. The "effective surface area," the surface area accessible to the airflow, is typically measured using a permeametry apparatus that relates the flow conductance to the permeable surface area via the Carman-Kozeny equation. This equation allows calculating the flow rate of a fluid flowing through the granules packed in the sample for a given pressure drop. However, Carman-Kozeny makes inherent assumptions about tunnel shapes and flow paths that may not accurately hold in situations where the particles possess a wide distribution in shapes, sizes, and aspect ratios, as is true with many powdered systems of technological and commercial interest. To address this challenge, we replicate these measurements virtually on micro-CT images of the powdered material, introducing a new Pore Network Model based on the skeleton of the Morse-Smale complex. Pores are identified as basins of the complex, their incidence encodes adjacency, and the conductivity of the capillary between them is computed from the cross-section at their interface. We build and solve a resistive network to compute an approximate laminar fluid flow through the pore structure. We provide two means of estimating flow-permeable surface area: (i) by direct computation of conductivity, and (ii) by identifying dead-ends in the flow coupled with isosurface extraction and the application of the Carman-Kozeny equation, with the aim of establishing consistency over a range of particle shapes, sizes, porosity levels, and void distribution patterns. Aniketh Venkat, Attila Gyulassy, Graham Kosiba, Amitesh Maiti, Henry Reinstein, Richard Gee, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Accurate and Robust Feature Importance Estimation under Distribution ShiftsabstractWith increasing reliance on the outcomes of black-box models in critical applications, post-hoc explainability tools that do not require access to the model internals are often used to enable humans understand and trust these models. In particular, we focus on the class of methods that can reveal the influence of input features on the predicted outputs. Despite their wide-spread adoption, existing methods are known to suffer from one or more of the following challenges: computational complexities, large uncertainties and most importantly, inability to handle real-world domain shifts. In this paper, we propose PRoFILE (Producing Robust Feature Importances using Loss Estimates), a novel feature importance estimation method that addresses all these challenges. Through the use of a loss estimator jointly trained with the predictive model and a causal objective, PRoFILE can accurately estimate the feature importance scores even under complex distribution shifts, without any additional re-training. To this end, we also develop learning strategies for training the loss estimator, namely contrastive and dropout calibration, and find that it can effectively detect distribution shifts. Using empirical studies on several benchmark image and non-image data, we show significant improvements over state-of-the-art approaches, both in terms of fidelity and robustness. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer, Andreas Spanias |
AAAI | 4 |
| 2021 | Distributed merge forest: a new fast and scalable approach for topological analysis at scaleabstractTopological analysis is used in several domains to identify and characterize important features in scientific data, and is now one of the established classes of techniques of proven practical use in scientific computing. The growth in parallelism and problem size tackled by modern simulations poses a particular challenge for these approaches. Fundamentally, the global encoding of topological features necessitates interprocess communication that limits their scaling. In this paper, we extend a new topological paradigm to the case of distributed computing, where the construction of a global merge tree is replaced by a distributed data structure, the merge forest, trading slower individual queries on the structure for faster end-to-end performance and scaling. Empirically, the queries that are most negatively affected also tend to have limited practical use. Our experimental results demonstrate the scalability of both the merge forest construction and the parallel queries needed in scientific workflows, and contrast this scalability with the two established alternatives that construct variations of a global tree. Xuan Huang 0007, Pavol Klacansky, Steve Petruzza, Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
ICS | 5 |
| 2021 | Understanding a program's resiliency through error propagationabstractAggressive technology scaling trends have worsened the transient fault problem in high-performance computing (HPC) systems. Some faults are benign, but others can lead to silent data corruption (SDC), which represents a serious problem; a fault introducing an error that is not readily detected nto an HPC simulation. Due to the insidious nature of SDCs, researchers have worked to understand their impact on applications. Previous studies have relied on expensive fault injection campaigns with uniform sampling to provide overall SDC rates, but this solution does not provide any feedback on the code regions without samples. Harshitha Menon, Kathryn Mohror, Peer-Timo Bremer, Yarden Livnat, Valerio Pascucci |
PPoPP | 4 |
| 2021 | Generalizable coordination of large multiscale workflows: challenges and learnings at scaleabstractThe advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine-Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications. Harsh Bhatia, Francesco Di Natale, Joseph Y. Moon, Joseph R. Chavez, Fikret Aydin, Christopher B. Stanley, Tomas Oppelstrup, Chris Neale, Sara Kokkila Schumacher, Dong H. Ahn, Stephen Herbein, Timothy S. Carpenter, Sandrasegaram Gnanakaran, Peer-Timo Bremer, James N. Glosli, Felice C. Lightstone, Helgi I. Ingólfsson |
SC | 15 |
| 2021 | Leveraging Topological Events in Tracking Graphs for Understanding Particle DiffusionabstractAbstract Single particle tracking (SPT) of fluorescent molecules provides significant insights into the diffusion and relative motion of tagged proteins and other structures of interest in biology. However, despite the latest advances in high‐resolution microscopy, individual particles are typically not distinguished from clusters of particles. This lack of resolution obscures potential evidence for how merging and splitting of particles affect their diffusion and any implications on the biological environment. The particle tracks are typically decomposed into individual segments at observed merge and split events, and analysis is performed without knowing the true count of particles in the resulting segments. Here, we address the challenges in analyzing particle tracks in the context of cancer biology. In particular, we study the tracks of KRAS protein, which is implicated in nearly 20% of all human cancers, and whose clustering and aggregation have been linked to the signaling pathway leading to uncontrolled cell growth. We present a new analysis approach for particle tracks by representing them as tracking graphs and using topological events – merging and splitting, to disambiguate the tracks. Using this analysis, we infer a lower bound on the count of particles as they cluster and create conditional distributions of diffusion speeds before and after merge and split events. Using thousands of time‐steps of simulated and in‐vitro SPT data, we demonstrate the efficacy of our method, as it offers the biologists a new, detailed look into the relationship between KRAS clustering and diffusion speeds. Torin McDonald, Rebika Shrestha, Xiyu Yi, Harsh Bhatia, De Chen, Debanjan Goswami, Valerio Pascucci, Thomas Turbyville, Peer-Timo Bremer |
Comput. Graph. Forum | 9 |
| 2021 | Coverage-Based Designs Improve Sample Mining and Hyperparameter Optimization
Gowtham Muniraju, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Cihan Tepedelenlioglu, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Vector Field Decompositions Using Multiscale Poisson KernelabstractExtraction of multiscale features using scale-space is one of the fundamental approaches to analyze scalar fields. However, similar techniques for vector fields are much less common, even though it is well known that, for example, turbulent flows contain cascades of nested vortices at different scales. The challenge is that the ideas related to scale-space are based upon iteratively smoothing the data to extract features at progressively larger scale, making it difficult to extract overlapping features. Instead, we consider spatial regions of influence in vector fields as scale, and introduce a new approach for the multiscale analysis of vector fields. Rather than smoothing the flow, we use the natural Helmholtz-Hodge decomposition to split it into small-scale and large-scale components using progressively larger neighborhoods. Our approach creates a natural separation of features by extracting local flow behavior, for example, a small vortex, from large-scale effects, for example, a background flow. We demonstrate our technique on large-scale, turbulent flows, and show multiscale features that cannot be extracted using state-of-the-art techniques. Harsh Bhatia, Robert M. Kirby, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Efficient and Flexible Hierarchical Data Layouts for a Unified Encoding of Scalar Field Precision and ResolutionabstractTo address the problem of ever-growing scientific data sizes making data movement a major hindrance to analysis, we introduce a novel encoding for scalar fields: a unified tree of resolution and precision, specifically constructed so that valid cuts correspond to sensible approximations of the original field in the precision-resolution space. Furthermore, we introduce a highly flexible encoding of such trees that forms a parameterized family of data hierarchies. We discuss how different parameter choices lead to different trade-offs in practice, and show how specific choices result in known data representation schemes such as zfp [52], idx [58], and jpeg2000 [76]. Finally, we provide system-level details and empirical evidence on how such hierarchies facilitate common approximate queries with minimal data movement and time, using real-world data sets ranging from a few gigabytes to nearly a terabyte in size. Experiments suggest that our new strategy of combining reductions in resolution and precision is competitive with state-of-the-art compression techniques with respect to data quality, while being significantly more flexible and orders of magnitude faster, and requiring significantly reduced resources. Duong Hoang, Brian Summa, Harsh Bhatia, Peter Lindstrom 0001, Pavol Klacansky, Will Usher 0001, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | SpotSDC: Revealing the Silent Data Corruption Propagation in High-Performance Computing SystemsabstractThe trend of rapid technology scaling is expected to make the hardware of high-performance computing (HPC) systems more susceptible to computational errors due to random bit flips. Some bit flips may cause a program to crash or have a minimal effect on the output, but others may lead to silent data corruption (SDC), i.e., undetected yet significant output errors. Classical fault injection analysis methods employ uniform sampling of random bit flips during program execution to derive a statistical resiliency profile. However, summarizing such fault injection result with sufficient detail is difficult, and understanding the behavior of the fault-corrupted program is still a challenge. In this article, we introduce SpotSDC, a visualization system to facilitate the analysis of a program's resilience to SDC. SpotSDC provides multiple perspectives at various levels of detail of the impact on the output relative to where in the source code the flipped bit occurs, which bit is flipped, and when during the execution it happens. SpotSDC also enables users to study the code protection and provide new insights to understand the behavior of a fault-injected program. Based on lessons learned, we demonstrate how what we found can improve the fault injection campaign method. Harshitha Menon, Dan Maljovec, Yarden Livnat, Shusen Liu 0001, Kathryn Mohror, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Visualizing Hierarchical Performance Profiles of Parallel Codes Using CallFlowabstractCalling context trees (CCTs) couple performance metrics with call paths, helping understand the execution and performance of parallel programs. To identify performance bottlenecks, programmers and performance analysts visually explore CCTs to form and validate hypotheses regarding degraded performance. However, due to the complexity of parallel programs, existing visual representations do not scale to applications running on a large number of processors. We present CallFlow, an interactive visual analysis tool that provides a high-level overview of CCTs together with semantic refinement operations to progressively explore CCTs. Using a flow-based metaphor, we visualize a CCT by treating execution time as a resource spent during the call chain, and demonstrate the effectiveness of our design with case studies on large-scale, production simulation codes. Huu Tan Nguyen, Abhinav Bhatele, Suraj P. Kesavan, Harsh Bhatia, Todd Gamblin, Kwan-Liu Ma, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2020 | Building Calibrated Deep Models via Uncertainty Matching with Auxiliary Interval PredictorsabstractWith rapid adoption of deep learning in critical applications, the question of when and how much to trust these models often arises, which drives the need to quantify the inherent uncertainties. While identifying all sources that account for the stochasticity of models is challenging, it is common to augment predictions with confidence intervals to convey the expected variations in a model's behavior. We require prediction intervals to be well-calibrated, reflect the true uncertainties, and to be sharp. However, existing techniques for obtaining prediction intervals are known to produce unsatisfactory results in at least one of these criteria. To address this challenge, we develop a novel approach for building calibrated estimators. More specifically, we use separate models for prediction and interval estimation, and pose a bi-level optimization problem that allows the former to leverage estimates from the latter through an uncertainty matching strategy. Using experiments in regression, time-series forecasting, and object localization, we show that our approach achieves significant improvements over existing uncertainty quantification methods, both in terms of model fidelity and calibration error. Jayaraman J. Thiagarajan, Bindya Venkatesh, Prasanna Sattigeri, Peer-Timo Bremer |
AAAI | 4 |
| 2020 | A Statistical Mechanics Framework for Task-Agnostic Sample Design in Machine LearningabstractIn this paper, we present a statistical mechanics framework to understand the effect of sampling properties of training data on the generalization gap of machine learning (ML) algorithms. We connect the generalization gap to the spatial properties of a sample design characterized by the pair correlation function (PCF). In particular, we express generalization gap in terms of the power spectra of the sample design and that of the function to be learned. Using this framework, we show that space-filling sample designs, such as blue noise and Poisson disk sampling, which optimize spectral properties, outperform random designs in terms of the generalization gap and characterize this gain in a closed-form. Our analysis also sheds light on design principles for constructing optimal task-agnostic sample designs that minimize the generalization gap. We corroborate our findings using regression experiments with neural networks on: a) synthetic functions, and b) a complex scientific simulator for inertial confinement fusion (ICF). Bhavya Kailkhura, Jayaraman J. Thiagarajan, Qunwei Li, Jize Zhang, Yi Zhou 0017, Peer-Timo Bremer |
NeurIPS | 6 |
| 2020 | MimicGAN: Robust Projection onto Image Manifolds with Corruption Mimicking
Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Peer-Timo Bremer |
Int. J. Comput. Vis. | 4 |
| 2020 | Toward Localized Topological Data Structures: Querying the Forest for the TreeabstractTopological approaches to data analysis can answer complex questions about the number, connectivity, and scale of intrinsic features in scalar data. However, the global nature of many topological structures makes their computation challenging at scale, and thus often limits the size of data that can be processed. One key quality to achieving scalability and performance on modern architectures is data locality, i.e., a process operates on data that resides in a nearby memory system, avoiding frequent jumps in data access patterns. From this perspective, topological computations are particularly challenging because the implied data structures represent features that can span the entire data set, often requiring a global traversal phase that limits their scalability. Traditionally, expensive preprocessing is considered an acceptable trade-off as it accelerates all subsequent queries. Most published use cases, however, explore only a fraction of all possible queries, most often those returning small, local features. In these cases, much of the global information is not utilized, yet computing it dominates the overall response time. We address this challenge for merge trees, one of the most commonly used topological structures. In particular, we propose an alternative representation, the merge forest, a collection of local trees corresponding to regions in a domain decomposition. Local trees are connected by a bridge set that allows us to recover any necessary global information at query time. The resulting system couples (i) a preprocessing that scales linearly in practice with (ii) fast runtime queries that provide the same functionality as traditional queries of a global merge tree. We test the scalability of our approach on a shared-memory parallel computer and demonstrate how data structure locality enables the analysis of large data with an order of magnitude performance improvement over the status quo. Furthermore, a merge forest reduces the memory overhead compared to a global merge tree and enables the processing of data sets that are an order of magnitude larger than possible with previous algorithms. Pavol Klacansky, Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Scalable Topological Data Analysis and Visualization for Evaluating Data-Driven Models in Scientific ApplicationsabstractWith the rapid adoption of machine learning techniques for large-scale applications in science and engineering comes the convergence of two grand challenges in visualization. First, the utilization of black box models (e.g., deep neural networks) calls for advanced techniques in exploring and interpreting model behaviors. Second, the rapid growth in computing has produced enormous datasets that require techniques that can handle millions or more samples. Although some solutions to these interpretability challenges have been proposed, they typically do not scale beyond thousands of samples, nor do they provide the high-level intuition scientists are looking for. Here, we present the first scalable solution to explore and analyze high-dimensional functions often encountered in the scientific data analysis pipeline. By combining a new streaming neighborhood graph construction, the corresponding topology computation, and a novel data aggregation scheme, namely topology aware datacubes, we enable interactive exploration of both the topological and the geometric aspect of high-dimensional data. Following two use cases from high-energy-density (HED) physics and computational biology, we demonstrate how these capabilities have led to crucial new insights in both applications. Shusen Liu 0001, Jim Gaffney, Jayson Luc Peterson, Peter B. Robinson, Harsh Bhatia, Valerio Pascucci, Brian K. Spears, Peer-Timo Bremer, Dan Maljovec, Rushil Anirudh, Jayaraman J. Thiagarajan, Sam Ade Jacobs, Brian Van Essen, David Hysom, Jae-Seung Yeom |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2019 | Parallelizing Training of Deep Generative Models on Massive Scientific DatasetsabstractTraining deep neural networks on large scientific data is a challenging task that requires enormous compute power, especially if no pre-trained models exist to initialize the process. We present a novel tournament method to train traditional as well as generative adversarial networks built on LBANN, a scalable deep learning framework optimized for HPC systems. LBANN combines multiple levels of parallelism and exploits some of the worlds largest supercomputers.We demonstrate our framework by creating a complex predictive model based on multi-variate data from high-energy-density physics containing hundreds of millions of images and hundreds of millions of scalar values derived from tens of millions of simulations of inertial confinement fusion. Our approach combines an HPC workflow and extends LBANN with optimized data ingestion and the new tournament-style training algorithm to produce a scalable neural network architecture using a CORAL-class supercomputer. Experimental results show that 64 trainers (1024 GPUs) achieve a speedup of 70.2× over a single trainer (16 GPUs) baseline, and an effective 109% parallel efficiency. Sam Ade Jacobs, Jim Gaffney, Tom Benson, Peter B. Robinson, Jayson Luc Peterson, Brian K. Spears, Brian Van Essen, David Hysom, Jae-Seung Yeom, Tim Moon, Rushil Anirudh, Jayaraman J. Thiagarajan, Shusen Liu 0001, Peer-Timo Bremer |
CLUSTER | 14 |
| 2019 | Unsupervised Dimension Selection Using a Blue Noise Graph SpectrumabstractUnsupervised dimension selection is an important problem that seeks to reduce dimensionality of data, while preserving the most useful characteristics. While dimensionality reduction is commonly utilized to construct low-dimensional embeddings, they produce feature spaces that are hard to interpret. Further, in applications such as sensor design, one needs to perform reduction directly in the input domain, instead of constructing transformed spaces. Consequently, dimension selection (DS) aims to solve the combinatorial problem of identifying the top-k dimensions, which is required for effective experiment design, reducing data while keeping it interpretable, and designing better sensing mechanisms. In this paper, we develop a novel approach for DS based on graph signal analysis to measure feature influence. By analyzing synthetic graph signals with a blue noise spectrum, we show that we can measure the importance of each dimension. Using experiments in supervised learning and image masking, we demonstrate the superiority of the proposed approach over existing techniques in capturing crucial characteristics of high dimensional spaces, using only a small subset of the original features. Jayaraman J. Thiagarajan, Rushil Anirudh, Rahul Sridhar, Peer-Timo Bremer |
ICASSP | 4 |
| 2019 | Understanding Deep Neural Networks through Input UncertaintiesabstractTechniques for understanding the functioning of complex machine learning models are becoming increasingly popular, not only to improve the validation process, but also to extract new insights about the data via exploratory analysis. Though a large class of such tools currently exists, most assume that predictions are point estimates and use a sensitivity analysis of these estimates to interpret the model. Using lightweight probabilistic networks we show how including prediction uncertainties in the sensitivity analysis leads to: (i) more robust and generalizable models; and (ii) a new approach for model interpretation through uncertainty decomposition. In particular, we introduce a new regularization that takes both the mean and variance of a prediction into account and demonstrate that the resulting networks provide improved generalization to unseen data. Furthermore, we propose a new technique to explain prediction uncertainties through uncertainties in the input domain, thus providing new ways to validate and interpret deep learning models. Jayaraman J. Thiagarajan, Irene Kim, Rushil Anirudh, Peer-Timo Bremer |
ICASSP | 4 |
| 2019 | A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancerabstractComputational models can define the functional dynamics of complex systems in exceptional detail. However, many modeling studies face seemingly incommensurate requirements: to gain meaningful insights into some phenomena requires models with high resolution (microscopic) detail that must nevertheless evolve over large (macroscopic) length- and time-scales. Multiscale modeling has become increasingly important to bridge this gap. Executing complex multiscale models on current petascale computers with high levels of parallelism and heterogeneous architectures is challenging. Many distinct types of resources need to be simultaneously managed, such as GPUs and CPUs, memory size and latencies, communication bottlenecks, and filesystem bandwidth. In addition, robustness to failure of compute nodes, network, and filesystems is critical. Francesco Di Natale, Harsh Bhatia, Timothy S. Carpenter, Chris Neale, Sara Kokkila Schumacher, Tomas Oppelstrup, Liam Stanton, Shiv Sundram, Thomas Scogland, Gautham Dharuman, Michael P. Surh, Yue Yang 0034, Claudia Misale, Lars Schneidenbach, Carlos H. A. Costa, Changhoan Kim, Bruce D'Amora, Sandrasegaram Gnanakaran, Dwight V. Nissley, Frederick H. Streitz, Felice C. Lightstone, Peer-Timo Bremer, James N. Glosli, Helgi I. Ingólfsson |
SC | 23 |
| 2019 | Shared-Memory Parallel Computation of Morse-Smale Complexes with Improved AccuracyabstractTopological techniques have proven to be a powerful tool in the analysis and visualization of large-scale scientific data. In particular, the Morse-Smale complex and its various components provide a rich framework for robust feature definition and computation. Consequently, there now exist a number of approaches to compute Morse-Smale complexes for large-scale data in parallel. However, existing techniques are based on discrete concepts which produce the correct topological structure but are known to introduce grid artifacts in the resulting geometry. Here, we present a new approach that combines parallel streamline computation with combinatorial methods to construct a high-quality discrete Morse-Smale complex. In addition to being invariant to the orientation of the underlying grid, this algorithm allows users to selectively build a subset of features using high-quality geometry. In particular, a user may specifically select which ascending/descending manifolds are reconstructed with improved accuracy, focusing computational effort where it matters for subsequent analysis. This approach computes Morse-Smale complexes for larger data than previously feasible with significant speedups. We demonstrate and validate our approach using several examples from a variety of different scientific domains, and evaluate the performance of our method. Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | A Study of the Trade-off Between Reducing Precision and Reducing Resolution for Data Analysis and VisualizationabstractThere currently exist two dominant strategies to reduce data sizes in analysis and visualization: reducing the precision of the data, e.g., through quantization, or reducing its resolution, e.g., by subsampling. Both have advantages and disadvantages and both face fundamental limits at which the reduced information ceases to be useful. The paper explores the additional gains that could be achieved by combining both strategies. In particular, we present a common framework that allows us to study the trade-off in reducing precision and/or resolution in a principled manner. We represent data reduction schemes as progressive streams of bits and study how various bit orderings such as by resolution, by precision, etc., impact the resulting approximation error across a variety of data sets as well as analysis tasks. Furthermore, we compute streams that are optimized for different tasks to serve as lower bounds on the achievable error. Scientific data management systems can use the results presented in this paper as guidance on how to store and stream data to make efficient use of the limited storage and bandwidth in practice. Duong Hoang, Pavol Klacansky, Harsh Bhatia, Peer-Timo Bremer, Peter Lindstrom 0001, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | NLIZE: A Perturbation-Driven Visual Interrogation Tool for Analyzing and Interpreting Natural Language Inference ModelsabstractWith the recent advances in deep learning, neural network models have obtained state-of-the-art performances for many linguistic tasks in natural language processing. However, this rapid progress also brings enormous challenges. The opaque nature of a neural network model leads to hard-to-debug-systems and difficult-to-interpret mechanisms. Here, we introduce a visualization system that, through a tight yet flexible integration between visualization elements and the underlying model, allows a user to interrogate the model by perturbing the input, internal state, and prediction while observing changes in other parts of the pipeline. We use the natural language inference problem as an example to illustrate how a perturbation-driven paradigm can help domain experts assess the potential limitation of a model, probe its inner states, and interpret and form hypotheses about fundamental model mechanisms such as attention. Shusen Liu 0001, Tao Li 0039, Vivek Srikumar, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Lose the Views: Limited Angle CT Reconstruction via Implicit Sinogram CompletionabstractComputed Tomography (CT) reconstruction is a fundamental component to a wide variety of applications ranging from security, to healthcare. The classical techniques require measuring projections, called sinograms, from a full 180° view of the object. However, obtaining a full-view is not always feasible, such as when scanning irregular objects that limit flexibility of scanner rotation. The resulting limited angle sinograms are known to produce highly artifact-laden reconstructions with existing techniques. In this paper, we propose to address this problem using CTNet - a system of 1D and 2D convolutional neural networks, that operates directly on a limited angle sinogram to predict the reconstruction. We use the x-ray transform on this prediction to obtain a "completed" sinogram, as if it came from a full 180°view. We feed this to standard analytical and iterative reconstruction techniques to obtain the final reconstruction. We show with extensive experimentation on a challenging real world dataset that this combined strategy outperforms many competitive baselines. We also propose a measure of confidence for the reconstruction that enables a practitioner to gauge the reliability of a prediction made by CTNet. We show that this measure is a strong indicator of quality as measured by the PSNR, while not requiring ground truth at test time. Finally, using a segmentation experiment, we show that our reconstruction also preserves the 3D structure of objects better than existing solutions. Rushil Anirudh, Hyojin Kim 0001, Jayaraman J. Thiagarajan, K. Aditya Mohan, Kyle Champley, Peer-Timo Bremer |
CVPR | 6 |
| 2018 | BabelFlow: An Embedded Domain Specific Language for Parallel Analysis and VisualizationabstractThe rapid growth in simulation data requires large-scale parallel implementations of scientific analysis and visualization algorithms, both to produce results within an acceptable timeframe and to enable in situ deployment. However, efficient and scalable implementations, especially of more complex analysis approaches, require not only advanced algorithms, but also an in-depth knowledge of the underlying runtime. Furthermore, different machine configurations and different applications may favor different runtimes, i.e., MPI vs Charm++ vs Legion, etc., and different hardware architectures. This diversity makes developing and maintaining a broadly applicable analysis software infrastructure challenging. We address some of these problems by explicitly separating the implementation of individual tasks of an algorithm from the dataflow connecting these tasks. In particular, we present an embedded domain specific language (EDSL) to describe algorithms using a new task graph abstraction. This task graph is then executed on top of one of several available runtimes (MPI, Charm++, Legion) using a thin layer of library calls. We demonstrate the flexibility and performance of this approach using three different large scale analysis and visualization use cases, i.e., topological analysis, rendering and compositing dataflow, and image registration of large microscopy scans. Despite the unavoidable overheads of a generic solution, our approach demonstrates performance portability at scale, and, in some cases, outperforms hand-optimized implementations. Steve Petruzza, Sean Treichler, Valerio Pascucci, Peer-Timo Bremer |
IPDPS | 4 |
| 2018 | Interactive Investigation of Traffic Congestion on Fat-Tree Networks Using TreeScopeabstractAbstract Parallel simulation codes often suffer from performance bottlenecks due to network congestion, leaving millions of dollars of investments underutilized. Given a network topology, it is critical to understand how different applications, job placements, routing schemes, etc., are affected by and contribute to network congestion, especially for large and complex networks. Understanding and optimizing communication on large‐scale networks is an active area of research. Domain experts often use exploratory tools to develop both intuitive and formal metrics for network health and performance. This paper presentsTreeScope, an interactive, web‐based visualization tool for exploring network traffic on large‐scale fat‐tree networks.TreeScopeencodes the network topology using a tailored matrix‐based representation and provides detailed visualization of all traffic in the network. We report on the design process ofTreeScope, which has been received positively by network researchers as well as system administrators. Through case studies of real and simulated data, we demonstrate howTreeScope's visual design and interactive support for complex queries on network traffic can provide experts with new insights into the occurrences and causes of congestion in the network. Harsh Bhatia, Abhinav Bhatele, Yarden Livnat, Jens Domke, Valerio Pascucci, Peer-Timo Bremer |
Comput. Graph. Forum | 7 |
| 2018 | Rendering and Extracting Extremal Features in 3D FieldsabstractAbstract Visualizing and extracting three‐dimensional features is important for many computational science applications, each with their own feature definitions and data types. While some are simple to state and implement (e.g. isosurfaces), others require more complicated mathematics (e.g. multiple derivatives, curvature, eigenvectors, etc.). Correctly implementing mathematical definitions is difficult, so experimenting with new features requires substantial investments. Furthermore, traditional interpolants rarely support the necessary derivatives, and approximations can reduce numerical stability. Our new approach directly translates mathematical notation into practical visualization and feature extraction, with minimal mental and implementation overhead. Using a mathematically expressive domain‐specific language, Diderot, we compute direct volume renderings and particle‐based feature samplings for a range of mathematical features. Non‐expert users can experiment with feature definitions without any exposure to meshes, interpolants, derivative computation, etc. We demonstrate high‐quality results on notoriously difficult features, such as ridges and vortex cores, using working code simple enough to be presented in its entirety. Gordon L. Kindlmann, Charisee Chiw, T. Huynh, Attila Gyulassy, John H. Reppy, Peer-Timo Bremer |
Comput. Graph. Forum | 6 |
| 2018 | Exploring High-Dimensional Structure via Axis-Aligned Decomposition of Linear ProjectionsabstractAbstract Two‐dimensional embeddings remain the dominant approach to visualize high dimensional data. The choice of embeddings ranges from highly non‐linear ones, which can capture complex relationships but are difficult to interpret quantitatively, to axis‐aligned projections, which are easy to interpret but are limited to bivariate relationships. Linear project can be considered as a compromise between complexity and interpretability, as they allow explicit axes labels, yet provide significantly more degrees of freedom compared to axis‐aligned projections. Nevertheless, interpreting the axes directions, which are often linear combinations of many non‐trivial components, remains difficult. To address this problem we introduce a structure aware decomposition of (multiple) linear projections into sparse sets of axis‐aligned projections, which jointly capture all information of the original linear ones. In particular, we use tools from Dempster‐Shafer theory to formally define how relevant a given axis‐aligned project is to explain the neighborhood relations displayed in some linear projection. Furthermore, we introduce a new approach to discover a diverse set of high quality linear projections and show that in practice the information of k linear projections is often jointly encoded in ∼ k axis‐aligned plots. We have integrated these ideas into an interactive visualization system that allows users to jointly browse both linear projections and their axis‐aligned representatives. Using a number of case studies we show how the resulting plots lead to more intuitive visualizations and new insights. Jayaraman J. Thiagarajan, Shusen Liu 0001, Karthikeyan Natesan Ramamurthy, Peer-Timo Bremer |
Comput. Graph. Forum | 4 |
| 2018 | A Spectral Approach for the Design of Experiments: Design, Analysis and AlgorithmsabstractThis paper proposes a new approach to construct high quality space-filling sample designs. First, we propose a novel technique to quantify the space-filling property and optimally trade-off uniformity and randomness in sample designs in arbitrary dimensions. Second, we connect the proposed metric (defined in the spatial domain) to the quality metric of the design performance (defined in the spectral domain). This connection serves as an analytic framework for evaluating the qualitative properties of space-filling designs in general. Using the theoretical insights provided by this spatial-spectral analysis, we derive the notion of optimal space-filling designs, which we refer to as space-filling spectral designs. Third, we propose an efficient estimator to evaluate the space-filling properties of sample designs in arbitrary dimensions and use it to develop an optimization framework for generating high quality space-filling designs. Finally, we carry out a detailed performance comparison on two different applications in varying dimensions: a) image reconstruction and b) surrogate modeling for several benchmark optimization functions and a physics simulation code for inertial confinement fusion (ICF). Our results clearly evidence the superiority of the proposed space-filling designs over existing approaches, particularly in high dimensions. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Charvi Rastogi, Pramod K. Varshney, Peer-Timo Bremer |
J. Mach. Learn. Res. | 5 |
| 2018 | MemAxes: Visualization and Analytics for Characterizing Complex Memory Performance BehaviorsabstractMemory performance is often a major bottleneck for high-performance computing (HPC) applications. Deepening memory hierarchies, complex memory management, and non-uniform access times have made memory performance behavior difficult to characterize, and users require novel, sophisticated tools to analyze and optimize this aspect of their codes. Existing tools target only specific factors of memory performance, such as hardware layout, allocations, or access instructions. However, today's tools do not suffice to characterize the complex relationships between these factors. Further, they require advanced expertise to be used effectively. We present MemAxes, a tool based on a novel approach for analytic-driven visualization of memory performance data. MemAxes uniquely allows users to analyze the different aspects related to memory performance by providing multiple visual contexts for a centralized dataset. We define mappings of sampled memory access data to new and existing visual metaphors, each of which enabling a user to perform different analysis tasks. We present methods to guide user interaction by scoring subsets of the data based on known performance problems. This scoring is used to provide visual cues and automatically extract clusters of interest. We designed MemAxes in collaboration with experts in HPC and demonstrate its effectiveness in case studies. Alfredo Giménez, Todd Gamblin, Ilir Jusufi, Abhinav Bhatele, Martin Schulz 0001, Peer-Timo Bremer, Bernd Hamann |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Visual Exploration of Semantic Relationships in Neural Word EmbeddingsabstractConstructing distributed representations for words through neural language models and using the resulting vector spaces for analysis has become a crucial component of natural language processing (NLP). However, despite their widespread application, little is known about the structure and properties of these spaces. To gain insights into the relationship between words, the NLP community has begun to adapt high-dimensional visualization techniques. In particular, researchers commonly use t-distributed stochastic neighbor embeddings (t-SNE) and principal component analysis (PCA) to create two-dimensional embeddings for assessing the overall structure and exploring linear relationships (e.g., word analogies), respectively. Unfortunately, these techniques often produce mediocre or even misleading results and cannot address domain-specific visualization challenges that are crucial for understanding semantic relationships in word embeddings. Here, we introduce new embedding techniques for visualizing semantic and syntactic analogies, and the corresponding tests to determine whether the resulting views capture salient structures. Additionally, we introduce two novel views for a comprehensive study of analogy relationships. Finally, we augment t-SNE embeddings to convey uncertainty information in order to allow a reliable interpretation. Combined, the different views address a number of domain-specific tasks difficult to solve with existing tools. Shusen Liu 0001, Peer-Timo Bremer, Jayaraman J. Thiagarajan, Vivek Srikumar, Bei Wang 0001, Yarden Livnat, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | A Virtual Reality Visualization Tool for Neuron TracingabstractTracing neurons in large-scale microscopy data is crucial to establishing a wiring diagram of the brain, which is needed to understand how neural circuits in the brain process information and generate behavior. Automatic techniques often fail for large and complex datasets, and connectomics researchers may spend weeks or months manually tracing neurons using 2D image stacks. We present a design study of a new virtual reality (VR) system, developed in collaboration with trained neuroanatomists, to trace neurons in microscope scans of the visual cortex of primates. We hypothesize that using consumer-grade VR technology to interact with neurons directly in 3D will help neuroscientists better resolve complex cases and enable them to trace neurons faster and with less physical and mental strain. We discuss both the design process and technical challenges in developing an interactive system to navigate and manipulate terabyte-sized image volumes in VR. Using a number of different datasets, we demonstrate that, compared to widely used commercial software, consumer-grade VR presents a promising alternative for scientists. Will Usher 0001, Pavol Klacansky, Frederick Federer, Peer-Timo Bremer, Aaron Knoll, Jeff Yarch, Alessandra Angelucci, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Interactive Visualization and Exploration of Patient Progression in a Hospital Setting
Wathsala Widanagamaachchi, Yarden Livnat, Peer-Timo Bremer, Scott L. DuVall, Valerio Pascucci |
AMIA | 3 |
| 2017 | Exploring the evolution of pressure-perturbations to understand atmospheric phenomenaabstractAtmospheric sciences is the study of physical and chemical phenomena occurring within the Earth's atmosphere. The study entails understanding the state of the Earth's atmosphere, how it is changing over time and why. Understanding how various weather events develop and evolve is often conducted through retrospective analysis of past atmospheric events. Atmospheric scientists can then utilize tools to better predict potential hazards and provide earlier warnings for events that may impact life and property. Several atmospheric state variables can be measured to identify high-impact events, one of which is surface atmospheric pressure. Many weather events are characterized by variations in surface pressure from the mean pressure value (i.e., pressure-perturbations). Accordingly, there is significant interest in extracting and tracking pressure-perturbations both spatially and temporally to better understand the evolution of weather events. Here, we present a visualization and analysis environment that allows interactive exploration of pressure-perturbation data sets. Our system, for the first time, enables atmospheric scientists to interactively explore the spatiotemporal behaviors of pressure-perturbations for a range of values and provides support to leverage other conventional data sets such as radar imagery and wind observations. It also allows atmospheric scientists to evaluate model and parameter sensitivity, which is difficult if not impossible with conventional visualization tools in atmospheric sciences. Finally, we demonstrate the utility of our approach for retrospective analysis using different case studies of recorded severe weather events. Wathsala Widanagamaachchi, Alexander Jacques, Bei Wang 0001, Erik T. Crosman, Peer-Timo Bremer, Valerio Pascucci, John D. Horel |
PacificVis | 5 |
| 2017 | ScrubJay: deriving knowledge from the disarray of HPC performance dataabstractModern HPC centers comprise clusters, storage, networks, power and cooling infrastructure, and more. Analyzing the efficiency of these complex facilities is a daunting task. Increasingly, facilities deploy sensors and monitoring tools, but with millions of instrumented components, analyzing collected data manually is intractable. Data from an HPC center comprises different formats, granularities, and semantics, and handwritten scripts no longer suffice to transform the data into a digestible form. Alfredo Giménez, Todd Gamblin, Abhinav Bhatele, Chad Wood, Kathleen Shoga, Aniruddha Marathe, Peer-Timo Bremer, Bernd Hamann, Martin Schulz 0001 |
SC | 7 |
| 2017 | Visualizing High-Dimensional Data: Advances in the Past DecadeabstractMassive simulations and arrays of sensing devices, in combination with increasing computing resources, have generated large, complex, high-dimensional datasets used to study phenomena across numerous fields of study. Visualization plays an important role in exploring such datasets. We provide a comprehensive survey of advances in high-dimensional data visualization that focuses on the past decade. We aim at providing guidance for data practitioners to navigate through a modular view of the recent advances, inspiring the creation of new visualizations along the enriched visualization pipeline, and identifying future opportunities for visualization research. Shusen Liu 0001, Dan Maljovec, Bei Wang 0001, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Two-dimensional shape retrieval using the distribution of extrema of Laplacian eigenfunctions
Dongmei Niu, Peer-Timo Bremer, Peter Lindstrom 0001, Bernd Hamann, Yuanfeng Zhou, Caiming Zhang 0001 |
Vis. Comput. | 2 |
| 2016 | Interactive exploration of atomic trajectories through relative-angle distribution and associated uncertaintiesabstractExploration of atomic trajectories is fundamental to understanding and characterizing complex chemical systems important in many applications. For instance, any new insight into the mechanisms of ionic migration in catalytic materials could lead to a substantial increase in battery performance. A new statistical measure, called the relative-angle distribution, has been proposed to understand complex motion - whether Brownian, ballistic, or diffusive. The relative-angle distribution can be represented as a collection of 1D histograms, but is currently created in a slow, offline process, making any parameter exploration a tedious and time-consuming task. Furthermore, the resulting plot can hide uncertainty in both the data and the visualization. As a result, once rastered or printed at a fixed resolution, these histograms can be misleading. We present a new analysis tool for the exploration of atomic trajectories that combines an interactive histogram visualization with uncertainty information for both data and plotting errors, and is also linked to an interactive 3D display of trajectories. Our tool enables a holistic exploration of trajectories previously not feasible, with the potential for significant scientific impact. In collaboration with domain experts, we have deployed our tool ta analyze molecular dynamics simulations of lithium-ion diffusion. Users have found that the tool significantly accelerates the exploration process and have used it to validate a number of previously unconfirmed hypotheses. Harsh Bhatia, Attila Gyulassy, Valerio Pascucci, Martina Bremer, Mitchell T. Ong, Vincenzo Lordi, Erik W. Draeger, John E. Pask, Peer-Timo Bremer |
PacificVis | 9 |
| 2016 | Evaluation of In-Situ Analysis Strategies at Scale for Power Efficiency and ScalabilityabstractThe increasing gap between available compute power and I/O capabilities is resulting in simulation pipelines running on leadership computing facilities being reformulated. In particular, in-situ processing is complementing conventional post-process analysis, however, it can be performed by using the same compute resources as the simulation or using secondary dedicated resources. In this paper, we focus on three different in-situ analysis strategies, which use the same compute resources as the ongoing simulation but different data movement strategies. We evaluate the costs incurred by these strategies in terms of run time, scalability and power/energy consumption. Furthermore, we extrapolate power behavior to peta-scale and investigate different design choices through projections. Experimental evaluation at full machine scale on Titan supports that using fewer cores per node for in-situ analysis is the optimum choice in terms of scalability. Hence, further research effort should be devoted towards developing in-situ analysis techniques following this strategy in future high-end systems. Ivan Rodero, Manish Parashar, Aaditya G. Landge, Sidharth Kumar, Valerio Pascucci, Peer-Timo Bremer |
CCGrid | 6 |
| 2016 | Theoretical guarantees for poisson disk sampling using pair correlation functionabstractIn this paper, we study the problem of generating uniform random point samples on a domain of d dimensional space based on a minimum distance criterion between point samples (Poisson-disk sampling or PDS). First, we formally define PDS via the pair correlation function (PCF) to quantitatively evaluate properties of the sampling process. Surprisingly, none of the existing PDS techniques satisfy both uniformity and minimum distance criterion, simultaneously. These approaches typically create an approximate PDS with high regularity, and inherently present high risk for sample aliasing. Our new formulation based on PCF introduces a new approach to evaluate PDS properties which leads to theoretical bounds on the size of a PDS in arbitrary dimensions as well as a faster algorithm to create better quality samplings than the current PDS approaches. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Pramod K. Varshney |
ICASSP | 3 |
| 2016 | Analyzing Network Health and Congestion in Dragonfly-Based SupercomputersabstractThe dragonfly topology is a popular choice for building high-radix, low-diameter, hierarchical networks with high-bandwidth links. On Cray installations of the dragonfly network, job placement policies and routing inefficiencies can lead to significant network congestion for a single job and multi-job workloads. In this paper, we explore the effects of job placement, parallel workloads and network configurations on network health to develop a better understanding of inter-job interference. We have developed a functional network simulator, Damselfly, to model the network behavior of Cray Cascade, and a visual analytics tool, DragonView, to analyze the simulation output. We simulate several parallel workloads based on five representative communication patterns on up to 131,072 cores. Our simulations and visualizations provide unique insight into the buildup of network congestion and present a trade-off between deployment dollar costs and performance of the network. Abhinav Bhatele, Yarden Livnat, Valerio Pascucci, Peer-Timo Bremer |
IPDPS | 5 |
| 2016 | Caliper: performance introspection for HPC software stacksabstractMany performance engineering tasks, from long-term performance monitoring to post-mortem analysis and online tuning, require efficient runtime methods for introspection and performance data collection. To understand interactions between components in increasingly modular HPC software, performance introspection hooks must be integrated into runtime systems, libraries, and application codes across the software stack. This requires an interoperable, cross-stack, general-purpose approach to performance data collection, which neither application-specific performance measurement nor traditional profile or trace analysis tools provide. With Caliper, we have developed a general abstraction layer to provide performance data collection as a service to applications, runtime systems, libraries, and tools. Individual software components connect to Caliper in independent data producer, data consumer, and measurement control roles, which allows them to share performance data across software stack boundaries. We demonstrate Caliper's performance analysis capbilities with two case studies of production scenarios. David Böhme, Todd Gamblin, D. A. Beckingsale, Peer-Timo Bremer, Alfredo Giménez, Matthew P. LeGendre, Olga Pearce, Martin Schulz 0001 |
SC | 4 |
| 2016 | The Grassmannian Atlas: A General Framework for Exploring Linear Projections of High-Dimensional DataabstractAbstract Linear projections are one of the most common approaches to visualize high‐dimensional data. Since the space of possible projections is large, existing systems usually select a small set of interesting projections by ranking a large set of candidate projections based on a chosen quality measure. However, while highly ranked projections can be informative, some lower ranked ones could offer important complementary information. Therefore, selection based on ranking may miss projections that are important to provide a global picture of the data. The proposed work fills this gap by presenting the Grassmannian Atlas, a framework that captures the global structures of quality measures in the space of all projections, which enables a systematic exploration of many complementary projections and provides new insights into the properties of existing quality measures. Shusen Liu 0001, Peer-Timo Bremer, J. J. Jayaraman, Bei Wang 0001, Brian Summa, Valerio Pascucci |
Comput. Graph. Forum | 2 |
| 2016 | Stair blue noise samplingabstractA common solution to reducing visible aliasing artifacts in image reconstruction is to employ sampling patterns with a blue noise power spectrum. These sampling patterns can prevent discernible artifacts by replacing them with incoherent noise. Here, we propose a new family of blue noise distributions, Stair blue noise , which is mathematically tractable and enables parameter optimization to obtain the optimal sampling distribution. Furthermore, for a given sample budget, the proposed blue noise distribution achieves a significantly larger alias-free low-frequency region compared to existing approaches, without introducing visible artifacts in the mid-frequencies. We also develop a new sample synthesis algorithm that benefits from the use of an unbiased spatial statistics estimator and efficient optimization strategies. Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Pramod K. Varshney |
ACM Trans. Graph. | 3 |
| 2016 | Ordering Traces Logically to Identify Lateness in Message Passing ProgramsabstractEvent traces are valuable for understanding the behavior of parallel programs. However, automatically analyzing a large parallel trace is difficult, especially without a specific objective. We aid this endeavor by extracting a trace's logical structure, an ordering of trace events derived from happened-before relationships, while taking into account developer intent. Using this structure, we can calculate an operation's delay relative to its peers on other processes. The logical structure also serves as a platform for comparing and clustering processes as well as highlighting communication patterns in a trace visualization. We present an algorithm for determining this idealized logical structure from traces of message passing programs, and we develop metrics to quantify delays and differences among processes. We implement our techniques in Ravel, a parallel trace visualization tool that displays both logical and physical timelines. Rather than showing the duration of each operation, we display where delays begin and end, and how they propagate. We apply our approach to the traces of several message passing applications, demonstrating the accuracy of our extracted structure and its utility in analyzing these codes. Katherine E. Isaacs, Todd Gamblin, Abhinav Bhatele, Martin Schulz 0001, Bernd Hamann, Peer-Timo Bremer |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2016 | Interstitial and Interlayer Ion Diffusion Geometry Extraction in Graphitic Nanosphere Battery MaterialsabstractLarge-scale molecular dynamics (MD) simulations are commonly used for simulating the synthesis and ion diffusion of battery materials. A good battery anode material is determined by its capacity to store ion or other diffusers. However, modeling of ion diffusion dynamics and transport properties at large length and long time scales would be impossible with current MD codes. To analyze the fundamental properties of these materials, therefore, we turn to geometric and topological analysis of their structure. In this paper, we apply a novel technique inspired by discrete Morse theory to the Delaunay triangulation of the simulated geometry of a thermally annealed carbon nanosphere. We utilize our computed structures to drive further geometric analysis to extract the interstitial diffusion structure as a single mesh. Our results provide a new approach to analyze the geometry of the simulated carbon nanosphere, and new insights into the role of carbon defect size and distribution in determining the charge capacity and charge dynamics of these carbon based battery materials. Attila Gyulassy, Aaron Knoll, Kah Chun Lau, Bei Wang 0001, Peer-Timo Bremer, Michael E. Papka, Larry A. Curtiss, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2015 | A Randomized Ensemble Approach to Industrial CT SegmentationabstractTuning the models and parameters of common segmentation approaches is challenging especially in the presence of noise and artifacts. Ensemble-based techniques attempt to compensate by randomly varying models and/or parameters to create a diverse set of hypotheses, which are subsequently ranked to arrive at the best solution. However, these methods have been restricted to cases where the underlying models are well-established, e.g. natural images. In practice, it is difficult to determine a suitable base-model and the amount of randomization required. Furthermore, for multi-object scenes no single hypothesis may perform well for all objects, reducing the overall quality of the results. This paper presents a new ensemble-based segmentation framework for industrial CT images demonstrating that comparatively simple models and randomization strategies can significantly improve the result over existing techniques. Furthermore, we introduce a per-object based ranking, followed by a consensus inference that can outperform even the best case scenario of existing hypothesis ranking approaches. We demonstrate the effectiveness of our approach using a set of noise and artifact rich CT images from baggage security and show that it significantly outperforms existing solutions in this area. Hyojin Kim 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICCV | 3 |
| 2015 | Identifying the Culprits Behind Network CongestionabstractNetwork congestion is one of the primary causes of performance degradation, performance variability and poor scaling in communication-heavy parallel applications. However, the causes and mechanisms of network congestion on modern interconnection networks are not well understood. We need new approaches to analyze, model and predict this critical behaviour in order to improve the performance of large-scale parallel applications. This paper applies supervised learning algorithms, such as forests of extremely randomized trees and gradient boosted regression trees, to perform regression analysis on communication data and application execution time. Using data derived from multiple executions, we create models to predict the execution time of communication-heavy parallel applications. This analysis also identifies the features and associated hardware components that have the most impact on network congestion and intern, on execution time. The ideas presented in this paper have wide applicability: predicting the execution time on a different number of nodes, or different input datasets, or even for an unknown code, identifying the best configuration parameters for an application, and finding the root causes of network congestion on different architectures. Abhinav Bhatele, Andrew R. Titus, Jayaraman J. Thiagarajan, Todd Gamblin, Peer-Timo Bremer, Martin Schulz 0001, Laxmikant V. Kalé |
IPDPS | 6 |
| 2015 | Recovering logical structure from Charm++ event tracesabstractAsynchrony and non-determinism in Charm++ programs present a significant challenge in analyzing their event traces. We present a new framework to organize event traces of parallel programs written in Charm++. Our reorganization allows one to more easily explore and analyze such traces by providing context through logical structure. We describe several heuristics to compensate for missing dependencies between events that currently cannot be easily recorded. We introduce a new task ordering that recovers logical structure from the non-deterministic execution order. Using the logical structure, we define several metrics to help guide developers to performance problems. We demonstrate our approach through two proxy applications written in Charm++. Finally, we discuss the applicability of this framework to other task-based runtimes and provide guidelines for tracing to support this form of analysis. Katherine E. Isaacs, Abhinav Bhatele, Jonathan Lifflander, David Böhme, Todd Gamblin, Martin Schulz 0001, Bernd Hamann, Peer-Timo Bremer |
SC | 8 |
| 2015 | Visual Exploration of High-Dimensional Data through Subspace Analysis and Dynamic ProjectionsabstractAbstract We introduce a novel interactive framework for visualizing and exploring high‐dimensional datasets based on subspace analysis and dynamic projections. We assume the high‐dimensional dataset can be represented by a mixture of low‐dimensional linear subspaces with mixed dimensions, and provide a method to reliably estimate the intrinsic dimension and linear basis of each subspace extracted from the subspace clustering. Subsequently, we use these bases to define unique 2D linear projections as viewpoints from which to visualize the data. To understand the relationships among the different projections and to discover hidden patterns, we connect these projections through dynamic projections that create smooth animated transitions between pairs of projections. We introduce the view transition graph, which provides flexible navigation among these projections to facilitate an intuitive exploration. Finally, we provide detailed comparisons with related systems, and use real‐world examples to demonstrate the novelty and usability of our proposed framework. Shusen Liu 0001, Bei Wang 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Valerio Pascucci |
Comput. Graph. Forum | 4 |
| 2015 | Local, smooth, and consistent Jacobi set simplification
Harsh Bhatia, Bei Wang 0001, Gregory Norgard, Valerio Pascucci, Peer-Timo Bremer |
Comput. Geom. | 5 |
| 2015 | Distributed Seams for Gigapixel PanoramasabstractGigapixel panoramas are an increasingly popular digital image application. They are often created as a mosaic of many smaller images. The mosaic acquisition can take many hours causing the individual images to differ in exposure and lighting conditions. A blending operation is often necessary to give the appearance of a seamless image. The blending quality depends on the magnitude of discontinuity along the image boundaries. Often, new boundaries, or seams, are first computed that minimize this transition. Current techniques based on multi-labeling Graph Cuts are too slow and memory intensive for gigapixel sized panoramas. In this paper, we present a parallel, out-of-core seam computing technique that is fast, has small memory footprint, and is capable of running efficiently on different types of parallel systems. Its maximum memory usage is configurable, in the form of a cache, which can improve performance by reducing redundant disk I/O and computations. It shows near-perfect scaling on symmetric multiprocessing systems and good scaling on clusters and distributed shared memory systems. Our technique improves the time required to compute seams for gigapixel imagery from many hours (or even days) to just a few minutes, while still producing boundaries with energy that is on-par with Graph Cuts. Sujin Philip, Brian Summa, Julien Tierny, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Multiple kernel interpolation for inverting non-linear dimensionality reduction and dimension estimationabstractThe problem of stably inverting a non-linear dimensionality reduction map has applications in data visualization and machine learning, besides being of theoretical interest. In this paper, we propose a meshfree interpolation method for obtaining such inverse maps using a non-negative linear combination of multiple interpolants. We show that the proposed scheme can improve upon the approximation power of its individual constituent kernels, and discuss the conditions under which its parameters can be uniquely estimated. We also provide an approach for estimating the intrinsic dimensionality (ID) of manifolds using the proposed inverse map. Experiments using multiple kernel interpolation for reconstruction of novel test data and ID estimation show an improved or similar performance compared to existing techniques. Jayaraman J. Thiagarajan, Peer-Timo Bremer, Karthikeyan Natesan Ramamurthy |
ICASSP | 2 |
| 2014 | Image segmentation using consensus from hierarchical segmentation ensemblesabstractUnsupervised, automatic image segmentation without contextual knowledge, or user intervention is a challenging problem. The key to robust segmentation is an appropriate selection of local features and metrics. However, a single aggregation of the local features using a greedy merging order often results in incorrect segmentation. This paper presents an unsupervised approach, which uses the consensus inferred from hierarchical segmentation ensembles, for partitioning images into foreground and background regions. By exploring an expanded set of possible aggregations of the local features, the proposed method generates meaningful segmentations that are not often revealed when only the optimal hierarchy is considered. A graph cuts-based approach is employed to combine the consensus along with a foreground-background model estimate, obtained using the ensemble, for effective segmentation. Experiments with a standard dataset show promising results when compared to several existing methods including the state-of-the-art weak supervised techniques that use co-segmentation. Hyojin Kim 0001, Jayaraman J. Thiagarajan, Peer-Timo Bremer |
ICIP | 3 |
| 2014 | Automatic image annotation using inverse maps from semantic embeddingsabstractHuman annotation in large scale image databases is time-consuming and error-prone. Since it is very hard to mine image databases using just visual features or textual descriptors, it is common to transform the image features into a semantically meaningful space. In this paper, we propose to perform image annotation in a semantic space inferred based on sparse representations. By constructing a semantic embedding for the visual features, that is constrained to be close to the tag embedding, we show that a robust inverse map can be used to predict the tags. Experiments using standard datasets show the effectiveness of the proposed approach in automatic image annotation when compared to existing methods. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Peer-Timo Bremer, Andreas Spanias |
ICIP | 4 |
| 2014 | Extracting logical structure and identifying stragglers in parallel execution tracesabstractWe introduce a new approach to automatically extract an idealized logical structure from a parallel execution trace. We use this structure to define intuitive metrics such as the lateness of a process involved in a parallel execution. By analyzing and illustrating traces in terms of logical steps, we leverage a developer's understanding of the happened-before relations in a parallel program. This technique can uncover dependency chains, elucidate communication patterns, and highlight sources and propagation of delays, all of which may be obscured in a traditional trace visualization. Katherine E. Isaacs, Todd Gamblin, Abhinav Bhatele, Peer-Timo Bremer, Martin Schulz 0001, Bernd Hamann |
PPoPP | 4 |
| 2014 | Dissecting On-Node Memory Access Performance: A Semantic ApproachabstractOptimizing memory access is critical for performance and power efficiency. CPU manufacturers have developed sampling-based performance measurement units (PMUs) that report precise costs of memory accesses at specific addresses. However, this data is too low-level to be meaningfully interpreted and contains an excessive amount of irrelevant or uninteresting information. We have developed a method to gather fine-grained memory access performance data for specific data objects and regions of code with low overhead and attribute semantic information to the sampled memory accesses. This information provides the context necessary to more effectively interpret the data. We have developed a tool that performs this sampling and attribution and used the tool to discover and diagnose performance problems in real-world applications. Our techniques provide useful insight into the memory behaviour of applications and allow programmers to understand the performance ramifications of key design decisions: domain decomposition, multi-threading, and data motion within distributed memory systems. Alfredo Giménez, Todd Gamblin, Barry Rountree, Abhinav Bhatele, Ilir Jusufi, Peer-Timo Bremer, Bernd Hamann |
SC | 6 |
| 2014 | Efficient I/O and Storage of Adaptive-Resolution DataabstractWe present an efficient, flexible, adaptive-resolution I/O framework that is suitable for both uniform and Adaptive Mesh Refinement (AMR) simulations. In an AMR setting, current solutions typically represent each resolution level as an independent grid which often results in inefficient storage and performance. Our technique coalesces domain data into a unified, multiresolution representation with fast, spatially aggregated I/O. Furthermore, our framework easily extends to importance-driven storage of uniform grids, for example, by storing regions of interest at full resolution and nonessential regions at lower resolution for visualization or analysis. Our framework, which is an extension of the PIDX framework, achieves state of the art disk usage and I/O performance regardless of resolution of the data, regions of interest, and the number of processes that generated the data. We demonstrate the scalability and efficiency of our framework using the Uintah and S3D large-scale combustion codes on the Mira and Edison supercomputers. Sidharth Kumar, John Edwards 0002, Peer-Timo Bremer, Aaron Knoll, Cameron Christensen, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Valerio Pascucci |
SC | 3 |
| 2014 | In-Situ Feature Extraction of Large Scale Combustion Simulations Using Segmented Merge TreesabstractThe ever increasing amount of data generated by scientific simulations coupled with system I/O constraints are fueling a need for in-situ analysis techniques. Of particular interest are approaches that produce reduced data representations while maintaining the ability to redefine, extract, and study features in a post-process to obtain scientific insights. This paper presents two variants of in-situ feature extraction techniques using segmented merge trees, which encode a wide range of threshold based features. The first approach is a fast, low communication cost technique that generates an exact solution but has limited scalability. The second is a scalable, local approximation that nevertheless is guaranteed to correctly extract all features up to a predefined size. We demonstrate both variants using some of the largest combustion simulations available on leadership class supercomputers. Our approach allows state-of-the-art, feature-based analysis to be performed in-situ at significantly higher frequency than currently possible and with negligible impact on the overall simulation runtime. Aaditya G. Landge, Valerio Pascucci, Attila Gyulassy, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer |
SC | 7 |
| 2014 | Extracting Features from Time-Dependent Vector Fields Using Internal Reference FramesabstractAbstract Extracting features from complex, time‐dependent flow fields remains a significant challenge despite substantial research efforts, especially because most flow features of interest are defined with respect to a given reference frame. Pathline‐based techniques, such as the FTLE field, are complex to implement and resource intensive, whereas scalar transforms, such as λ2, often produce artifacts and require somewhat arbitrary thresholds. Both approaches aim to analyze the flow in a more suitable frame, yet neither technique explicitly constructs one. This paper introduces a new data‐driven technique to compute internal reference frames for large‐scale complex flows. More general than uniformly moving frames, these frames can transform unsteady fields, which otherwise require substantial processing of resources, into a sequence of individual snapshots that can be analyzed using the large body of steady‐flow analysis techniques. Our approach is simple, theoretically well‐founded, and uses an embarrassingly parallel algorithm for structured as well as unstructured data. Using several case studies from fluid flow and turbulent combustion, we demonstrate that internal frames are distinguished, result in temporally coherent structures, and can extract well‐known as well as notoriously elusive features one snapshot at a time. Harsh Bhatia, Valerio Pascucci, Robert M. Kirby, Peer-Timo Bremer |
Comput. Graph. Forum | 4 |
| 2014 | Stability of Dissipation Elements: A Case Study in CombustionabstractAbstract Recently, dissipation elements have been gaining popularity as a mechanism for measurement of fundamental properties of turbulent flow, such as turbulence length scales and zonal partitioning. Dissipation elements segment a domain according to the source and destination of streamlines in the gradient flow field of a scalar function f : → ℝ. They have traditionally been computed by numerically integrating streamlines from the center of each voxel in the positive and negative gradient directions, and grouping those voxels whose streamlines terminate at the same extremal pair. We show that the same structures map well to combinatorial topology concepts developed recently in the visualization community. Namely, dissipation elements correspond to sets of cells of the Morse‐Smale complex. The topology‐based formulation enables a more exploratory analysis of the nature of dissipation elements, in particular, in understanding their stability with respect to small scale variations. We present two examples from combustion science that raise significant questions about the role of small scale perturbation and indeed the definition of dissipation elements themselves. Attila Gyulassy, Peer-Timo Bremer, Ray W. Grout, Hemanth Kolla, Jacqueline Chen, Valerio Pascucci |
Comput. Graph. Forum | 2 |
| 2014 | Distortion-Guided Structure-Driven Interactive Exploration of High-Dimensional DataabstractAbstract Dimension reduction techniques are essential for feature selection and feature extraction of complex high‐dimensional data. These techniques, which construct low‐dimensional representations of data, are typically geometrically motivated, computationally efficient and approximately preserve certain structural properties of the data. However, they are often used as black box solutions in data exploration and their results can be difficult to interpret. To assess the quality of these results, quality measures, such as co‐ranking [ LV09 ], have been proposed to quantify structural distortions that occur between high‐dimensional and low‐dimensional data representations. Such measures could be evaluated and visualized point‐wise to further highlight erroneous regions [ MLGH13 ]. In this work, we provide an interactive visualization framework for exploring high‐dimensional data via its two‐dimensional embeddings obtained from dimension reduction, using a rich set of user interactions. We ask the following question: what new insights do we obtain regarding the structure of the data, with interactive manipulations of its embeddings in the visual space? We augment the two‐dimensional embeddings with structural abstractions obtained from hierarchical clusterings, to help users navigate and manipulate subsets of the data. We use point‐wise distortion measures to highlight interesting regions in the domain, and further to guide our selection of the appropriate level of clusterings that are aligned with the regions of interest. Under the static setting, point‐wise distortions indicate the level of structural uncertainty within the embeddings. Under the dynamic setting, on‐the‐fly updates of point‐wise distortions due to data movement and data deletion reflect structural relations among different parts of the data, which may lead to new and valuable insights. Shusen Liu 0001, Bei Wang 0001, Peer-Timo Bremer, Valerio Pascucci |
Comput. Graph. Forum | 3 |
| 2014 | The Natural Helmholtz-Hodge Decomposition for Open-Boundary Flow AnalysisabstractThe Helmholtz-Hodge decomposition (HHD), which describes a flow as the sum of an incompressible, an irrotational, and a harmonic flow, is a fundamental tool for simulation and analysis. Unfortunately, for bounded domains, the HHD is not uniquely defined, traditionally, boundary conditions are imposed to obtain a unique solution. However, in general, the boundary conditions used during the simulation may not be known known, or the simulation may use open boundary conditions. In these cases, the flow imposed by traditional boundary conditions may not be compatible with the given data, which leads to sometimes drastic artifacts and distortions in all three components, hence producing unphysical results. This paper proposes the natural HHD, which is defined by separating the flow into internal and external components. Using a completely data-driven approach, the proposed technique obtains uniqueness without assuming boundary conditions a priori. As a result, it enables a reliable and artifact-free analysis for flows with open boundaries or unknown boundary conditions. Furthermore, our approach computes the HHD on a point-wise basis in contrast to the existing global techniques, and thus supports computing inexpensive local approximations for any subset of the domain. Finally, the technique is easy to implement for a variety of spatial discretizations and interpolated fields in both two and three dimensions. Harsh Bhatia, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Combing the Communication Hairball: Visualizing Parallel Execution Traces using Logical TimeabstractWith the continuous rise in complexity of modern supercomputers, optimizing the performance of large-scale parallel programs is becoming increasingly challenging. Simultaneously, the growth in scale magnifies the impact of even minor inefficiencies--potentially millions of compute hours and megawatts in power consumption can be wasted on avoidable mistakes or sub-optimal algorithms. This makes performance analysis and optimization critical elements in the software development process. One of the most common forms of performance analysis is to study execution traces, which record a history of per-process events and interprocess messages in a parallel application. Trace visualizations allow users to browse this event history and search for insights into the observed performance behavior. However, current visualizations are difficult to understand even for small process counts and do not scale gracefully beyond a few hundred processes. Organizing events in time leads to a virtually unintelligible conglomerate of interleaved events and moderately high process counts overtax even the largest display. As an alternative, we present a new trace visualization approach based on transforming the event history into logical time inferred directly from happened-before relationships. This emphasizes the code's structural behavior, which is much more familiar to the application developer. The original timing data, or other information, is then encoded through color, leading to a more intuitive visualization. Furthermore, we use the discrete nature of logical timelines to cluster processes according to their local behavior leading to a scalable visualization of even long traces on large process counts. We demonstrate our system using two case studies on large-scale parallel codes. Katherine E. Isaacs, Peer-Timo Bremer, Ilir Jusufi, Todd Gamblin, Abhinav Bhatele, Martin Schulz 0001, Bernd Hamann |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Exploring power behaviors and trade-offs of in-situ data analyticsabstractAs scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems. Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky |
SC | 7 |
| 2013 | Robust computation of Morse-Smale complexes of bilinear functions
Gregory Norgard, Peer-Timo Bremer |
Comput. Aided Geom. Des. | 2 |
| 2013 | Ridge-Valley graphs: Combinatorial ridge detection using Jacobi sets
Gregory Norgard, Peer-Timo Bremer |
Comput. Aided Geom. Des. | 2 |
| 2013 | Comments on the "Meshless Helmholtz-Hodge Decomposition"abstractThe Helmholtz-Hodge decomposition (HHD) is one of the fundamental theorems of fluids describing the decomposition of a flow field into its divergence-free, curl-free, and harmonic components. Solving for the HHD is intimately connected to the choice of boundary conditions which determine the uniqueness and orthogonality of the decomposition. This article points out that one of the boundary conditions used in a recent paper "Meshless Helmholtz-Hodge Decomposition" is, in general, invalid and provides an analytical example demonstrating the problem. We hope that this clarification on the theory will foster further research in this area and prevent undue problems in applying and extending the original approach. Harsh Bhatia, Gregory Norgard, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | The Helmholtz-Hodge Decomposition - A SurveyabstractThe Helmholtz-Hodge Decomposition (HHD) describes the decomposition of a flow field into its divergence-free and curl-free components. Many researchers in various communities like weather modeling, oceanology, geophysics, and computer graphics are interested in understanding the properties of flow representing physical phenomena such as incompressibility and vorticity. The HHD has proven to be an important tool in the analysis of fluids, making it one of the fundamental theorems in fluid dynamics. The recent advances in the area of flow analysis have led to the application of the HHD in a number of research communities such as flow visualization, topological analysis, imaging, and robotics. However, because the initial body of work, primarily in the physics communities, research on the topic has become fragmented with different communities working largely in isolation often repeating and sometimes contradicting each others results. Additionally, different nomenclature has evolved which further obscures the fundamental connections between fields making the transfer of knowledge difficult. This survey attempts to address these problems by collecting a comprehensive list of relevant references and examining them using a common terminology. A particular focus is the discussion of boundary conditions when computing the HHD. The goal is to promote further research in the field by creating a common repository of techniques to compute the HHD as well as a large collection of example applications in a broad range of areas. Harsh Bhatia, Gregory Norgard, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | ManyVis: Multiple Applications in an Integrated Visualization EnvironmentabstractAs the visualization field matures, an increasing number of general toolkits are developed to cover a broad range of applications. However, no general tool can incorporate the latest capabilities for all possible applications, nor can the user interfaces and workflows be easily adjusted to accommodate all user communities. As a result, users will often chose either substandard solutions presented in familiar, customized tools or assemble a patchwork of individual applications glued through ad-hoc scripts and extensive, manual intervention. Instead, we need the ability to easily and rapidly assemble the best-in-task tools into custom interfaces and workflows to optimally serve any given application community. Unfortunately, creating such meta-applications at the API or SDK level is difficult, time consuming, and often infeasible due to the sheer variety of data models, design philosophies, limits in functionality, and the use of closed commercial systems. In this paper, we present the ManyVis framework which enables custom solutions to be built both rapidly and simply by allowing coordination and communication across existing unrelated applications. ManyVis allows users to combine software tools with complementary characteristics into one virtual application driven by a single, custom-designed interface. Atul Rungta, Brian Summa, Dogan Demir, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | Combining in-situ and in-transit processing to enable extreme-scale scientific analysisabstractWith the onset of extreme-scale computing, I/O constraints make it increasingly difficult for scientists to save a sufficient amount of raw simulation data to persistent storage. One potential solution is to change the data analysis pipeline from a post-process centric to a concurrent approach based on either in-situ or in-transit processing. In this context computations are considered in-situ if they utilize the primary compute resources, while in-transit processing refers to offloading computations to a set of secondary resources using asynchronous data transfers. In this paper we explore the design and implementation of three common analysis techniques typically performed on large-scale scientific simulations: topological analysis, descriptive statistics, and visualization. We summarize algorithmic developments, describe a resource scheduling system to coordinate the execution of various analysis workflows, and discuss our implementation using the DataSpaces and ADIOS frameworks that support efficient data movement between in-situ and in-transit computations. We demonstrate the efficiency of our lightweight, flexible framework by deploying it on the Jaguar XK6 to analyze data generated by S3D, a massively parallel turbulent combustion code. Our framework allows scientists dealing with the data deluge at extreme scale to perform analyses at increased temporal resolutions, mitigate I/O costs, and significantly improve the time to insight. Janine Bennett, Hasan Abbasi, Peer-Timo Bremer, Ray W. Grout, Attila Gyulassy, Tong Jin 0002, Scott Klasky, Hemanth Kolla, Manish Parashar, Valerio Pascucci, Philippe P. Pébay, David C. Thompson 0001, Hongfeng Yu 0001, Fan Zhang 0004, Jacqueline Chen |
SC | 3 |
| 2012 | Novel views of performance data to analyze large-scale adaptive applicationsabstractPerformance analysis of parallel scientific codes is becoming increasingly difficult due to the rapidly growing complexity of applications and architectures. Existing tools fall short in providing intuitive views that facilitate the process of performance debugging and tuning. In this paper, we extend recent ideas of projecting and visualizing performance data for faster, more intuitive analysis of applications. We collect detailed per-level and per-phase measurements for a dynamically load-balanced, structured AMR library and project per-core data collected in the hardware domain on to the application's communication topology. We show how our projections and visualizations lead to a rapid diagnosis of and mitigation strategy for a previously elusive scaling bottleneck in the library that is hard to detect using conventional tools. Our new insights have resulted in a 22% performance improvement for a 65,536-core run of the AMR library on an IBM Blue Gene/P system. Abhinav Bhatele, Todd Gamblin, Katherine E. Isaacs, Brian T. N. Gunney, Martin Schulz 0001, Peer-Timo Bremer, Bernd Hamann |
SC | 6 |
| 2012 | Mapping applications with collectives over sub-communicators on torus networksabstractThe placement of tasks in a parallel application on specific nodes of a supercomputer can significantly impact performance. Traditionally, this task mapping has focused on reducing the distance between communicating tasks on the physical network. This minimizes the number of hops that point-to-point messages travel and thus reduces link sharing between messages and contention. However, for applications that use collectives over sub-communicators, this heuristic may not be optimal. Many collectives can benefit from an increase in bandwidth even at the cost of an increase in hop count, especially when sending large messages. For example, placing communicating tasks in a cube configuration rather than a plane or a line on a torus network increases the number of possible paths messages might take. This increases the available bandwidth which can lead to significant performance gains. We have developed Rubik, a tool that provides a simple and intuitive interface to create a wide variety of mappings for structured communication patterns. Rubik supports a number of elementary operations such as splits, tilts, or shifts, that can be combined into a large number of unique patterns. Each operation can be applied to disjoint groups of processes involved in collectives to increase the effective bandwidth. We demonstrate the use of Rubik for improving performance of two parallel codes, pF3D and Qbox, which use collectives over sub-communicators. Abhinav Bhatele, Todd Gamblin, Steve H. Langer, Peer-Timo Bremer, Erik W. Draeger, Bernd Hamann, Katherine E. Isaacs, Aaditya G. Landge, Joshua A. Levine, Valerio Pascucci, Martin Schulz 0001, Charles H. Still |
SC | 4 |
| 2012 | A Quantized Boundary Representation of 2D FlowsabstractAbstract Analysis and visualization of complex vector fields remain major challenges when studying large scale simulation of physical phenomena. The primary reason is the gap between the concepts of smooth vector field theory and their computational realization. In practice, researchers must choose between either numerical techniques, with limited or no guarantees on how they preserve fundamental invariants, or discrete techniques which limit the precision at which the vector field can be represented. We propose a new representation of vector fields that combines the advantages of both approaches. In particular, we represent a subset of possible streamlines by storing their paths as they traverse the edges of a triangulation. Using only a finite set of streamlines creates a fully discrete version of a vector field that nevertheless approximates the smooth flow up to a user controlled error bound. The discrete nature of our representation enables us to directly compute and classify analogues of critical points, closed orbits, and other common topological structures. Further, by varying the number of divisions (quantizations) used per edge, we vary the resolution used to represent the field, allowing for controlled precision. This representation is compact in memory and supports standard vector field operations. Joshua A. Levine, Shreeraj Jadhav, Harsh Bhatia, Valerio Pascucci, Peer-Timo Bremer |
Comput. Graph. Forum | 5 |
| 2012 | Flow Visualization with Quantified Spatial and Temporal Errors Using Edge MapsabstractRobust analysis of vector fields has been established as an important tool for deriving insights from the complex systems these fields model. Traditional analysis and visualization techniques rely primarily on computing streamlines through numerical integration. The inherent numerical errors of such approaches are usually ignored, leading to inconsistencies that cause unreliable visualizations and can ultimately prevent in-depth analysis. We propose a new representation for vector fields on surfaces that replaces numerical integration through triangles with maps from the triangle boundaries to themselves. This representation, called edge maps, permits a concise description of flow behaviors and is equivalent to computing all possible streamlines at a user defined error threshold. Independent of this error streamlines computed using edge maps are guaranteed to be consistent up to floating point precision, enabling the stable extraction of features such as the topological skeleton. Furthermore, our representation explicitly stores spatial and temporal errors which we use to produce more informative visualizations. This work describes the construction of edge maps, the error quantification, and a refinement procedure to adhere to a user defined error bound. Finally, we introduce new visualizations using the additional information provided by edge maps to indicate the uncertainty involved in computing streamlines and topological structures. Harsh Bhatia, Shreeraj Jadhav, Peer-Timo Bremer, Guoning Chen, Joshua A. Levine, Luis Gustavo Nonato, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | Computing Morse-Smale Complexes with Accurate GeometryabstractTopological techniques have proven highly successful in analyzing and visualizing scientific data. As a result, significant efforts have been made to compute structures like the Morse-Smale complex as robustly and efficiently as possible. However, the resulting algorithms, while topologically consistent, often produce incorrect connectivity as well as poor geometry. These problems may compromise or even invalidate any subsequent analysis. Moreover, such techniques may fail to improve even when the resolution of the domain mesh is increased, thus producing potentially incorrect results even for highly resolved functions. To address these problems we introduce two new algorithms: (i) a randomized algorithm to compute the discrete gradient of a scalar field that converges under refinement; and (ii) a deterministic variant which directly computes accurate geometry and thus correct connectivity of the MS complex. The first algorithm converges in the sense that on average it produces the correct result and its standard deviation approaches zero with increasing mesh resolution. The second algorithm uses two ordered traversals of the function to integrate the probabilities of the first to extract correct (near optimal) geometry and connectivity. We present an extensive empirical study using both synthetic and real-world data and demonstrates the advantages of our algorithms in comparison with several popular approaches. Attila Gyulassy, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Visualizing Network Traffic to Understand the Performance of Massively Parallel SimulationsabstractThe performance of massively parallel applications is often heavily impacted by the cost of communication among compute nodes. However, determining how to best use the network is a formidable task, made challenging by the ever increasing size and complexity of modern supercomputers. This paper applies visualization techniques to aid parallel application developers in understanding the network activity by enabling a detailed exploration of the flow of packets through the hardware interconnect. In order to visualize this large and complex data, we employ two linked views of the hardware network. The first is a 2D view, that represents the network structure as one of several simplified planar projections. This view is designed to allow a user to easily identify trends and patterns in the network traffic. The second is a 3D view that augments the 2D view by preserving the physical network topology and providing a context that is familiar to the application developers. Using the massively parallel multi-physics code pF3D as a case study, we demonstrate that our tool provides valuable insight that we use to explain and optimize pF3D's performance on an IBM Blue Gene/P system. Aaditya G. Landge, Joshua A. Levine, Abhinav Bhatele, Katherine E. Isaacs, Todd Gamblin, Martin Schulz 0001, Steve H. Langer, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2011 | Edge maps: Representing flow with bounded errorabstractRobust analysis of vector fields has been established as an important tool for deriving insights from the complex systems these fields model. Many analysis techniques rely on computing streamlines, a task often hampered by numerical instabilities. Approaches that ignore the resulting errors can lead to inconsistencies that may produce unreliable visualizations and ultimately prevent in-depth analysis. We propose a new representation for vector fields on surfaces that replaces numerical integration through triangles with linear maps defined on its boundary. This representation, called edge maps, is equivalent to computing all possible streamlines at a user defined error threshold. In spite of this error, all the streamlines computed using edge maps will be pairwise disjoint. Furthermore, our representation stores the error explicitly, and thus can be used to produce more informative visualizations. Given a piecewise-linear interpolated vector field, a recent result [15] shows that there are only 23 possible map classes for a triangle, permitting a concise description of flow behaviors. This work describes the details of computing edge maps, provides techniques to quantify and refine edge map error, and gives qualitative and visual comparisons to more traditional techniques. Harsh Bhatia, Shreeraj Jadhav, Peer-Timo Bremer, Guoning Chen, Joshua A. Levine, Luis Gustavo Nonato, Valerio Pascucci |
PacificVis | 3 |
| 2011 | Hybrid CPU-GPU Solver for Gradient Domain Processing of Massive ImagesabstractGradient domain processing is a computationally expensive image processing technique. Its use for processing massive images, giga or terapixels in size, can take several hours with serial techniques. To address this challenge, parallel algorithms are being developed to make this class of techniques applicable to the largest images available with running times that are more acceptable to the users. To this end we target the most ubiquitous form of computing power available today, which is small or medium scale clusters of commodity hardware. Such clusters are continuously increasing in scale, not only in the number of nodes, but also in the amount of parallelism available within each node in the form of multicore CPUs and GPUs. In this paper we present a hybrid parallel implementation of gradient domain processing for seamless stitching of gigapixel panoramas that utilizes MPI, threading and a CUDA based GPU component. We demonstrate the performance and scalability of our implementation by presenting results from two GPU clusters processing two large data sets. Sujin Philip, Brian Summa, Valerio Pascucci, Peer-Timo Bremer |
ICPADS | 4 |
| 2011 | Interpreting Performance Data across Intuitive DomainsabstractTo exploit the capabilities of current and future systems, developers must understand the interplay between on-node performance, domain decomposition, and an application's intrinsic communication patterns. While tools exist to gather and analyze data for each of these components individually, the resulting information is generally processed in isolation and presented in an abstract, categorical fashion unintuitive to most users. In this paper we present the HAC model, in which we identify the three domains of performance data most familiar to the user: (i)the application domain containing the application's working set, (ii) the hardware domain of the compute and network devices, and (iii) the communication domain of logical data transfers. We show that taking data from each of these domains and projecting, visualizing, and correlating it to the other domains can give valuable insights into the behavior of parallel application codes. The HAC abstraction opens the door for a new generation of tools that can help users more easily and intuitively associate performance data with root causes in the hardware system, the application's structure, and in its communication behavior, and by doing so leads to an improved understanding of the performance of their codes. Martin Schulz 0001, Joshua A. Levine, Peer-Timo Bremer, Todd Gamblin, Valerio Pascucci |
ICPP | 3 |
| 2011 | Topology-based Visualization of Transformation Pathways in Complex Chemical SystemsabstractAbstract Studying transformation in a chemical system by considering its energy as a function of coordinates of the system's components provides insight and changes our understanding of this process. Currently, a lack of effective visualization techniques for high‐dimensional energy functions limits chemists to plot energy with respect to one or two coordinates at a time. In some complex systems, developing a comprehensive understanding requires new visualization techniques that show relationships between all coordinates at the same time. We propose a new visualization technique that combines concepts from topological analysis, multi‐dimensional scaling, and graph layout to enable the analysis of energy functions for a wide range of molecular structures. We demonstrate our technique by studying the energy function of a dimer of formic and acetic acids and a LTA zeolite structure, in which we consider diffusion of methane. Kenes Beketayev, Gunther H. Weber, Maciej Haranczyk, Peer-Timo Bremer, Mario Hlawitschka, Bernd Hamann |
Comput. Graph. Forum | 4 |
| 2011 | Interactive editing of massive imagery made simple: Turning Atlanta into AtlantisabstractThis article presents a simple framework for progressive processing of high-resolution images with minimal resources. We demonstrate this framework's effectiveness by implementing an adaptive, multi-resolution solver for gradient-based image processing that, for the first time, is capable of handling gigapixel imagery in real time. With our system, artists can use commodity hardware to interactively edit massive imagery and apply complex operators, such as seamless cloning, panorama stitching, and tone mapping. We introduce a progressive Poisson solver that processes images in a purely coarse-to-fine manner, providing near instantaneous global approximations for interactive display (see Figure 1). We also allow for data-driven adaptive refinements to locally emulate the effects of a global solution. These techniques, combined with a fast, cache-friendly data access mechanism, allow the user to interactively explore and edit massive imagery, with the illusion of having a full solution at hand. In particular, we demonstrate the interactive modification of gigapixel panoramas that previously required extensive offline processing. Even with massive satellite images surpassing a hundred gigapixels in size, we enable repeated interactive editing in a dynamically changing environment. Images at these scales are significantly beyond the purview of previous methods yet are processed interactively using our techniques. Finally our system provides a robust and scalable out-of-core solver that consistently offers high-quality solutions while maintaining strict control over system resources. Brian Summa, Giorgio Scorzelli, Ming Jiang 0005, Peer-Timo Bremer, Valerio Pascucci |
ACM Trans. Graph. | 4 |
| 2011 | Feature-Based Statistical Analysis of Combustion Simulation DataabstractWe present a new framework for feature-based statistical analysis of large-scale scientific data and demonstrate its effectiveness by analyzing features from Direct Numerical Simulations (DNS) of turbulent combustion. Turbulent flows are ubiquitous and account for transport and mixing processes in combustion, astrophysics, fusion, and climate modeling among other disciplines. They are also characterized by coherent structure or organized motion, i.e. nonlocal entities whose geometrical features can directly impact molecular mixing and reactive processes. While traditional multi-point statistics provide correlative information, they lack nonlocal structural information, and hence, fail to provide mechanistic causality information between organized fluid motion and mixing and reactive processes. Hence, it is of great interest to capture and track flow features and their statistics together with their correlation with relevant scalar quantities, e.g. temperature or species concentrations. In our approach we encode the set of all possible flow features by pre-computing merge trees augmented with attributes, such as statistical moments of various scalar fields, e.g. temperature, as well as length-scales computed via spectral analysis. The computation is performed in an efficient streaming manner in a pre-processing step and results in a collection of meta-data that is orders of magnitude smaller than the original simulation data. This meta-data is sufficient to support a fully flexible and interactive analysis of the features, allowing for arbitrary thresholds, providing per-feature statistics, and creating various global diagnostics such as Cumulative Density Functions (CDFs), histograms, or time-series. We combine the analysis with a rendering of the features in a linked-view browser that enables scientists to interactively explore, visualize, and analyze the equivalent of one terabyte of simulation data. We highlight the utility of this new framework for combustion science; however, it is applicable to many other science domains. Janine Bennett, Vaidyanathan Krishnamoorthy, Shusen Liu 0001, Ray W. Grout, Evatt R. Hawkes, Jacqueline Chen, Jason F. Shepherd, Valerio Pascucci, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2011 | Interactive Exploration and Analysis of Large-Scale Simulations Using Topology-Based Data SegmentationabstractLarge-scale simulations are increasingly being used to study complex scientific and engineering phenomena. As a result, advanced visualization and data analysis are also becoming an integral part of the scientific process. Often, a key step in extracting insight from these large simulations involves the definition, extraction, and evaluation of features in the space and time coordinates of the solution. However, in many applications, these features involve a range of parameters and decisions that will affect the quality and direction of the analysis. Examples include particular level sets of a specific scalar field, or local inequalities between derived quantities. A critical step in the analysis is to understand how these arbitrary parameters/decisions impact the statistical properties of the features, since such a characterization will help to evaluate the conclusions of the analysis as a whole. We present a new topological framework that in a single-pass extracts and encodes entire families of possible features definitions as well as their statistical properties. For each time step we construct a hierarchical merge tree a highly compact, yet flexible feature representation. While this data structure is more than two orders of magnitude smaller than the raw simulation data it allows us to extract a set of features for any given parameter selection in a postprocessing step. Furthermore, we augment the trees with additional attributes making it possible to gather a large number of useful global, local, as well as conditional statistic that would otherwise be extremely difficult to compile. We also use this representation to create tracking graphs that describe the temporal evolution of the features over time. Our system provides a linked-view interface to explore the time-evolution of the graph interactively alongside the segmentation, thus making it possible to perform extensive data analysis in a very efficient manner. We demonstrate our framework by extracting and analyzing burning cells from a large-scale turbulent combustion simulation. In particular, we show how the statistical analysis enabled by our techniques provides new insight into the combustion process. Peer-Timo Bremer, Gunther H. Weber, Julien Tierny, Valerio Pascucci, Marcus S. Day, John B. Bell |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Topological Spines: A Structure-preserving Visual Representation of Scalar FieldsabstractWe present topological spines--a new visual representation that preserves the topological and geometric structure of a scalar field. This representation encodes the spatial relationships of the extrema of a scalar field together with the local volume and nesting structure of the surrounding contours. Unlike other topological representations, such as contour trees, our approach preserves the local geometric structure of the scalar field, including structural cycles that are useful for exposing symmetries in the data. To obtain this representation, we describe a novel mechanism based on the extraction of extremum graphs--sparse subsets of the Morse-Smale complex that retain the important structural information without the clutter and occlusion problems that arise from visualizing the entire complex directly. Extremum graphs form a natural multiresolution structure that allows the user to suppress noise and enhance topological features via the specification of a persistence range. Applications of our approach include the visualization of 3D scalar fields without occlusion artifacts, and the exploratory analysis of high-dimensional functions. Carlos D. Correa, Peter Lindstrom 0001, Peer-Timo Bremer |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Adaptive Extraction and Quantification of Geophysical VorticesabstractWe consider the problem of extracting discrete two-dimensional vortices from a turbulent flow. In our approach we use a reference model describing the expected physics and geometry of an idealized vortex. The model allows us to derive a novel correlation between the size of the vortex and its strength, measured as the square of its strain minus the square of its vorticity. For vortex detection in real models we use the strength parameter to locate potential vortex cores, then measure the similarity of our ideal analytical vortex and the real vortex core for different strength thresholds. This approach provides a metric for how well a vortex core is modeled by an ideal vortex. Moreover, this provides insight into the problem of choosing the thresholds that identify a vortex. By selecting a target coefficient of determination (i.e., statistical confidence), we determine on a per-vortex basis what threshold of the strength parameter would be required to extract that vortex at the chosen confidence. We validate our approach on real data from a global ocean simulation and derive from it a map of expected vortex strengths over the global ocean. Sean Williams, Mark R. Petersen, Peer-Timo Bremer, Matthew Hecht, Valerio Pascucci, James P. Ahrens, Mario Hlawitschka, Bernd Hamann |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | Lessons learned from moving earth system grid data sets over a 20 Gbps wide-area networkabstractIn preparation for the Intergovernmental Panel on Climate Change (IPCC) Fifth Assessment Report, the climate community will run the Coupled Model Intercomparison Project phase 5 (CMIP-5) experiments, which are designed to answer crucial questions about future regional climate change and the results of carbon feedback for different mitigation scenarios. The CMIP-5 experiments will generate petabytes of data that must be replicated seamlessly, reliably, and quickly to hundreds of research teams around the globe. As an end-to-end test of the technologies that will be used to perform this task, a multi-disciplinary team of researchers moved a small portion (10 TB) of the multimodel Coupled Model Intercomparison Project, Phase 3 data set used in the IPCC Fourth Assessment Report from three sources---the Argonne Leadership Computing Facility (ALCF), Lawrence Livermore National Laboratory (LLNL) and National Energy Research Scientific Computing Center (NERSC)---to the 2009 Supercomputing conference (SC09) show floor in Portland, Oregon, over circuits provided by DOE's ESnet. The team achieved a sustained data rate of 15 Gb/s on a 20 Gb/s network. More important, this effort provided critical feedback on how to deploy, tune, and monitor the middleware that will be used to replicate the upcoming petascale climate datasets. We report on obstacles overcome and the key lessons learned from this successful bandwidth challenge effort. Rajkumar Kettimuthu, Alex Sim, Dan Gunter, William E. Allcock, Peer-Timo Bremer, John Bresnahan, Andrew Cherry, Lisa Childers, Eli Dart, Ian T. Foster, Kevin Harms, Jason Hick, Jason Lee 0001, Michael Link, Jeff Long, Keith Miller 0005, Vijaya Natarajan, Valerio Pascucci, Kenneth Raffenetti, David Ressman, Dean N. Williams, Loren Wilson, Linda Winkler |
HPDC | 5 |
| 2010 | Analyzing and Tracking Burning Structures in Lean Premixed Hydrogen FlamesabstractThis paper presents topology-based methods to robustly extract, analyze, and track features defined as subsets of isosurfaces. First, we demonstrate how features identified by thresholding isosurfaces can be defined in terms of the Morse complex. Second, we present a specialized hierarchy that encodes the feature segmentation independent of the threshold while still providing a flexible multiresolution representation. Third, for a given parameter selection, we create detailed tracking graphs representing the complete evolution of all features in a combustion simulation over several hundred time steps. Finally, we discuss a user interface that correlates the tracking information with interactive rendering of the segmented isosurfaces enabling an in-depth analysis of the temporal behavior. We demonstrate our approach by analyzing three numerical simulations of lean hydrogen flames subject to different levels of turbulence. Due to their unstable nature, lean flames burn in cells separated by locally extinguished regions. The number, area, and evolution over time of these cells provide important insights into the impact of turbulence on the combustion process. Utilizing the hierarchy, we can perform an extensive parameter study without reprocessing the data for each set of parameters. The resulting statistics enable scientists to select appropriate parameters and provide insight into the sensitivity of the results with respect to the choice of parameters. Our method allows for the first time to quantitatively correlate the turbulence of the burning process with the distribution of burning regions, properly segmented and selected. In particular, our analysis shows that counterintuitively stronger turbulence leads to larger cell structures, which burn more intensely than expected. This behavior suggests that flames could be stabilized under much leaner conditions than previously anticipated. Peer-Timo Bremer, Gunther H. Weber, Valerio Pascucci, Marcus S. Day, John B. Bell |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | Visual Exploration of High Dimensional Scalar FunctionsabstractAn important goal of scientific data analysis is to understand the behavior of a system or process based on a sample of the system. In many instances it is possible to observe both input parameters and system outputs, and characterize the system as a high-dimensional function. Such data sets arise, for instance, in large numerical simulations, as energy landscapes in optimization problems, or in the analysis of image data relating to biological or medical parameters. This paper proposes an approach to analyze and visualizing such data sets. The proposed method combines topological and geometric techniques to provide interactive visualizations of discretely sampled high-dimensional scalar fields. The method relies on a segmentation of the parameter space using an approximate Morse-Smale complex on the cloud of point samples. For each crystal of the Morse-Smale complex, a regression of the system parameters with respect to the output yields a curve in the parameter space. The result is a simplified geometric representation of the Morse-Smale complex in the high dimensional input domain. Finally, the geometric representation is embedded in 2D, using dimension reduction, to provide a visualization platform. The geometric properties of the regression curves enable the visualization of additional information about each crystal such as local and global shape, width, length, and sampling densities. The method is illustrated on several synthetic examples of two dimensional functions. Two use cases, using data sets from the UCI machine learning repository, demonstrate the utility of the proposed approach on real data. Finally, in collaboration with domain experts the proposed method is applied to two scientific challenges. The analysis of parameters of climate simulations and their relationship to predicted global energy flux and the concentrations of chemical species in a combustion simulation and their integration with temperature. Samuel Gerber, Peer-Timo Bremer, Valerio Pascucci, Ross T. Whitaker |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2009 | A Topological Framework for the Interactive Exploration of Large Scale Turbulent CombustionabstractThe advent of highly accurate, large scale volumetric simulations has made data analysis and visualization techniques an integral part of the modern scientific process. To develop new insights from raw data, scientists need the ability to define features of interest in a flexible manner and to understand how changes in the feature definition impact the subsequent analysis of the data. Therefore, simply exploring the raw data is not sufficient. This paper presents a new topological framework for the analysis of large scale, time-varying, turbulent combustion simulations. It allows the scientists to interactively explore the complete parameter space of fuel consumption thresholds for an entire time-dependent combustion simulation. By computing augmented merge trees and their corresponding data segmentations, the system allows the user complete flexibility to segment, select, and track burning cells through time thanks to a linked view interface. We developed this technique in the context of low-swirl turbulent pre-mixed same simulation analysis, where the topological abstractions enable an efficient tracking through time of the burning cells and provide new qualitative and quantitative insights into the dynamics of the combustion process. Peer-Timo Bremer, Gunther H. Weber, Julien Tierny, Valerio Pascucci, Marcus S. Day, John B. Bell |
eScience | 1 |
| 2008 | A Practical Approach to Morse-Smale Complex Computation: Scalability and GeneralityabstractThe Morse-Smale (MS) complex has proven to be a useful tool in extracting and visualizing features from scalar-valued data. However, efficient computation of the MS complex for large scale data remains a challenging problem. We describe a new algorithm and easily extensible framework for computing MS complexes for large scale data of any dimension where scalar values are given at the vertices of a closure-finite and weak topology (CW) complex, therefore enabling computation on a wide variety of meshes such as regular grids, simplicial meshes, and adaptive multiresolution (AMR) meshes. A new divide-and-conquer strategy allows for memory-efficient computation of the MS complex and simplification on-the-fly to control the size of the output. In addition to being able to handle various data formats, the framework supports implementation-specific optimizations, for example, for regular data. We present the complete characterization of critical point cancellations in all dimensions. This technique enables the topology based analysis of large data on off-the-shelf computers. In particular we demonstrate the first full computation of the MS complex for a 1 billion/1024(3) node grid on a laptop computer with 2Gb memory. Attila Gyulassy, Peer-Timo Bremer, Bernd Hamann, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2007 | Robust on-line computation of Reeb graphs: simplicity and speedabstractReeb graphs are a fundamental data structure for understanding and representing the topology of shapes. They are used in computer graphics, solid modeling, and visualization for applications ranging from the computation of similarities and finding defects in complex models to the automatic selection of visualization parameters. We introduce an on-line algorithm that reads a stream of elements (vertices, triangles, tetrahedra, etc.) and continuously maintains the Reeb graph of all elements already reed. The algorithm is robust in handling non-manifold meshes and general in its applicability to input models of any dimension. Optionally, we construct a skeleton-like embedding of the Reeb graph, and/or remove topological noise to reduce the output size. For interactive multi-resolution navigation we also build a hierarchical data structure which allows real-time extraction of approximated Reeb graphs containing all topological features above a given error threshold. Our extensive experiments show both high performance and practical linear scalability for meshes ranging from thousands to hundreds of millions of triangles. We apply our algorithm to the largest, most general, triangulated surfaces available to us, including 3D, 4D and 5D simplicial meshes. To demonstrate one important application we use Reeb graphs to find and highlight topological defects in meshes, including some widely believed to be "clean." Valerio Pascucci, Giorgio Scorzelli, Peer-Timo Bremer, Ajith Mascarenhas |
ACM Trans. Graph. | 3 |
| 2007 | Topological Landscapes: A Terrain Metaphor for Scientific DataabstractScientific visualization and illustration tools are designed to help people understand the structure and complexity of scientific data with images that are as informative and intuitive as possible. In this context the use of metaphors plays an important role since they make complex information easily accessible by using commonly known concepts. In this paper we propose a new metaphor, called "Topological Landscapes," which facilitates understanding the topological structure of scalar functions. The basic idea is to construct a terrain with the same topology as a given dataset and to display the terrain as an easily understood representation of the actual input data. In this projection from an $n$-dimensional scalar function to a two-dimensional (2D) model we preserve function values of critical points, the persistence (function span) of topological features, and one possible additional metric property (in our examples volume). By displaying this topologically equivalent landscape together with the original data we harness the natural human proficiency in understanding terrain topography and make complex topological information easily accessible. Gunther H. Weber, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Segmenting molecular surfaces
Vijay Natarajan, Yusu Wang 0001, Peer-Timo Bremer, Valerio Pascucci, Bernd Hamann |
Comput. Aided Geom. Des. | 3 |
| 2006 | Spectral surface quadrangulationabstractResampling raw surface meshes is one of the most fundamental operations used by nearly all digital geometry processing systems. The vast majority of this work has focused on triangular remeshing, yet quadrilateral meshes are preferred for many surface PDE problems, especially fluid dynamics, and are best suited for defining Catmull-Clark subdivision surfaces. We describe a fundamentally new approach to the quadrangulation of manifold polygon meshes using Laplacian eigenfunctions, the natural harmonics of the surface. These surface functions distribute their extrema evenly across a mesh, which connect via gradient flow into a quadrangular base mesh. An iterative relaxation algorithm simultaneously refines this initial complex to produce a globally smooth parameterization of the surface. From this, we can construct a well-shaped quadrilateral mesh with very few extraordinary vertices. The quality of this mesh relies on the initial choice of eigenfunction, for which we describe algorithms and hueristics to efficiently and effectively select the harmonic most appropriate for the intended application. Shen Dong, Peer-Timo Bremer, Michael Garland, Valerio Pascucci, John C. Hart |
ACM Trans. Graph. | 2 |
| 2006 | A Topological Approach to Simplification of Three-Dimensional Scalar FunctionsabstractThis paper describes an efficient combinatorial method for simplification of topological features in a 3D scalar function. The Morse-Smale complex, which provides a succinct representation of a function's associated gradient flow field, is used to identify topological features and their significance. The simplification process, guided by the Morse-Smale complex, proceeds by repeatedly applying two atomic operations that each remove a pair of critical points from the complex. Efficient storage of the complex results in execution of these atomic operations at interactive rates. Visualization of the simplified complex shows that the simplification preserves significant topological features while removing small features and noise. Attila Gyulassy, Vijay Natarajan, Valerio Pascucci, Peer-Timo Bremer, Bernd Hamann |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2006 | Understanding the Structure of the Turbulent Mixing Layer in Hydrodynamic InstabilitiesabstractWhen a heavy fluid is placed above a light fluid, tiny vertical perturbations in the interface create a characteristic structure of rising bubbles and falling spikes known as Rayleigh-Taylor instability. Rayleigh-Taylor instabilities have received much attention over the past half-century because of their importance in understanding many natural and man-made phenomena, ranging from the rate of formation of heavy elements in supernovae to the design of capsules for Inertial Confinement Fusion. We present a new approach to analyze Rayleigh-Taylor instabilities in which we extract a hierarchical segmentation of the mixing envelope surface to identify bubbles and analyze analogous segmentations of fields on the original interface plane. We compute meaningful statistical information that reveals the evolution of topological features and corroborates the observations made by scientists. We also use geometric tracking to follow the evolution of single bubbles and highlight merge/split events leading to the formation of the large and complex structures characteristic of the later stages. In particular we (i) Provide a formal definition of a bubble; (ii) Segment the envelope surface to identify bubbles; (iii) Provide a multi-scale analysis technique to produce statistical measures of bubble growth; (iv) Correlate bubble measurements with analysis of fields on the interface plane; (v) Track the evolution of individual bubbles over time. Our approach is based on the rigorous mathematical foundations of Morse theory and can be applied to a more general class of applications. David E. Laney, Peer-Timo Bremer, Ajith Mascarenhas, Paul L. Miller, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2005 | Maximizing Adaptivity in Hierarchical Topological ModelsabstractWe present an approach to hierarchically encode the topology of functions over triangulated surfaces. Its Morse-Smale complex, a well known structure in computational topology, describes the topology of a function. Following concepts of Morse theory, a Morse-Smale complex (and therefore a function's topology) can be simplified by successively canceling pairs of critical points. We demonstrate how cancellations can be effectively encoded to produce a highly adaptive topology-based multi-resolution representation of a given function. Contrary to the approach, we avoid encoding the complete complex in a traditional mesh hierarchy. Instead, the information is split into a new structure we call a cancellation forest and a traditional dependency graph. The combination of this new structure with a traditional mesh hierarchy proofs to be significantly more flexible than the one previously reported. In particular, we can create hierarchies that are guaranteed to be of logarithmic height. Peer-Timo Bremer, Valerio Pascucci, Bernd Hamann |
SMI | 1 |
| 2005 | Topology-based Simplification for Feature Extraction from 3D Scalar FieldsabstractIn this paper, we present a topological approach for simplifying continuous functions defined on volumetric domains. We introduce two atomic operations that remove pairs of critical points of the function and design a combinatorial algorithm that simplifies the Morse-Smale complex by repeated application of these operations. The Morse-Smale complex is a topological data structure that provides a compact representation of gradient flow between critical points of a function. Critical points paired by the Morse-Smale complex identify topological features and their importance. The simplification procedure leaves important critical points untouched, and is therefore useful for extracting desirable features. We also present a visualization of the simplified topology. Attila Gyulassy, Vijay Natarajan, Valerio Pascucci, Peer-Timo Bremer, Bernd Hamann |
IEEE Visualization | 4 |
| 2004 | A Topological Hierarchy for Functions on Triangulated SurfacesabstractWe combine topological and geometric methods to construct a multiresolution representation for a function over a two-dimensional domain. In a preprocessing stage, we create the Morse-Smale complex of the function and progressively simplify its topology by cancelling pairs of critical points. Based on a simple notion of dependency among these cancellations, we construct a hierarchical data structure supporting traversal and reconstruction operations similarly to traditional geometry-based representations. We use this data structure to extract topologically valid approximations that satisfy error bounds provided at runtime. Peer-Timo Bremer, Herbert Edelsbrunner, Bernd Hamann, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2003 | A Multi-Resolution Data Structure for 2-Dimensional Morse FunctionsabstractWe combine topological and geometric methods to construct a multi-resolution data structure for functions over two-dimensional domains. Starting with the Morse-Smale complex, we construct a topological hierarchy by progressively canceling critical points in pairs. Concurrently, we create a geometric hierarchy by adapting the geometry to the changes in topology. The data structure supports mesh traversal operations similarly to traditional multi-resolution representations. Peer-Timo Bremer, Herbert Edelsbrunner, Bernd Hamann, Valerio Pascucci |
IEEE Visualization | 1 |