Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Patricia K. Fasel

dblp:64/1228 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 49% High-performance computing · 48% Distributed systems · 2%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › topological data analysis
contour tree
0.512021
Scalable Contour Tree Computation by Data Parallel Peak Pruning · IEEE Trans. Vis. Comput. Graph. 2021
Visualization and visual analytics
topological data analysis
0.512021
Scalable Contour Tree Computation by Data Parallel Peak Pruning · IEEE Trans. Vis. Comput. Graph. 2021
Parallel and multicore computing
data-parallel programming
0.512021
Scalable Contour Tree Computation by Data Parallel Peak Pruning · IEEE Trans. Vis. Comput. Graph. 2021
Parallel and multicore computing
parallel algorithms
0.512021
Scalable Contour Tree Computation by Data Parallel Peak Pruning · IEEE Trans. Vis. Comput. Graph. 2021
High-performance computing › scientific data analysis
in-situ analysis
0.422021
Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approach · SC 2015
Scalable Contour Tree Computation by Data Parallel Peak Pruning · IEEE Trans. Vis. Comput. Graph. 2021
High-performance computing › performance optimization at scale
extreme-scale scalability
0.112012
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing
performance optimization at scale
0.112012
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing
scientific computing systems
0.112012
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing
scientific computing
0.012001
PAWS: Collective Interactions and Data Transfers · HPDC 2001
Parallel and multicore computing › parallel libraries
communication library
0.011998
Efficient Coupling of Parallel Applications Using PAWS · HPDC 1998
High-performance computing
data transfer
0.011998
Efficient Coupling of Parallel Applications Using PAWS · HPDC 1998

Methods — techniques the papers use, named apart from their topics

shared-memory parallelism · 1.0peak pruning · 1.0GPU-CPU hybrid computation · 1.0particle-grid methods · 0.3distributed type casts · 0.0collective ports · 0.0
YearPublicationVenuePosition
2021 Scalable Contour Tree Computation by Data Parallel Peak Pruning
abstract
As data sets grow to exascale, automated data analysis and visualization are increasingly important, to intermediate human understanding and to reduce demands on disk storage via in situ analysis. Trends in architecture of high performance computing systems necessitate analysis algorithms to make effective use of combinations of massively multicore and distributed systems. One of the principal analytic tools is the contour tree, which analyses relationships between contours to identify features of more than local importance. Unfortunately, the predominant algorithms for computing the contour tree are explicitly serial, and founded on serial metaphors, which has limited the scalability of this form of analysis. While there is some work on distributed contour tree computation, and separately on hybrid GPU-CPU computation, there is no efficient algorithm with strong formal guarantees on performance allied with fast practical performance. We report the first shared SMP algorithm for fully parallel contour tree computation, with formal guarantees of O(lg V lg t) parallel steps and O(V lg V) work for data with V samples and t contour tree supernodes, and implementations with more than 30× parallel speed up on both CPU using TBB and GPU using Thrust and up 70× speed up compared to the serial sweep and merge algorithm.
Hamish A. Carr, Gunther H. Weber, Christopher M. Sewell, Oliver Rübel, Patricia K. Fasel, James P. Ahrens
IEEE Trans. Vis. Comput. Graph.5
2015 Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approach
abstract
Large-scale simulations can produce hundreds of terabytes to petabytes of data, complicating and limiting the efficiency of workflows. Traditionally, outputs are stored on the file system and analyzed in post-processing. With the rapidly increasing size and complexity of simulations, this approach faces an uncertain future. Trending techniques consist of performing the analysis in-situ, utilizing the same resources as the simulation, and/or off-loading subsets of the data to a compute-intensive analysis system. We introduce an analysis framework developed for HACC, a cosmological N-body code, that uses both in-situ and co-scheduling approaches for handling petabyte-scale outputs. We compare different analysis set-ups ranging from purely off-line, to purely in-situ to in-situ/co-scheduling. The analysis routines are implemented using the PISTON/VTK-m framework, allowing a single implementation of an algorithm that simultaneously targets a variety of GPU, multi-core, and many-core architectures.
Christopher M. Sewell, Katrin Heitmann, Hal Finkel, George Zagaris, Suzanne Parete-Koon, Patricia K. Fasel, Adrian Pope, Nicholas Frontiere, Li-Ta Lo, O. E. Bronson Messer, Salman Habib 0002, James P. Ahrens
SC6
2012 The universe at extreme scale: multi-petaflop sky simulation on the BG/Q
abstract
Remarkable observational advances have established a compelling cross-validated model of the Universe. Yet, two key pillars of this model -- dark matter and dark energy -- remain mysterious. Next-generation sky surveys will map billions of galaxies to explore the physics of the 'Dark Universe'. Science requirements for these surveys demand simulations at extreme scales; these will be delivered by the HACC (Hybrid/Hardware Accelerated Cosmology Code) framework. HACC's novel algorithmic structure allows tuning across diverse architectures, including accelerated and multi-core systems. On the IBM BG/Q, HACC attains unprecedented scalable performance - currently 6.23 PFlops at 62% of peak and 92% parallel efficiency on 786,432 cores (48 racks) - at extreme problem sizes with up to almost two trillion particles, larger than any cosmological simulation yet performed. HACC simulations at these scales will for the first time enable tracking individual galaxies over the entire volume of a cosmological survey.
Salman Habib 0002, Vitali A. Morozov, Hal Finkel, Adrian Pope, Katrin Heitmann, Kalyan Kumaran, Tom Peterka, Joseph A. Insley, David Daniel, Patricia K. Fasel, Nicholas Frontiere, Zarija Lukic
SC10
2001 PAWS: Collective Interactions and Data Transfers
abstract
The authors discuss problems and solutions pertaining to the interaction of components representing parallel applications. We introduce the notion of a collective port which is an extension of the Common Component Architecture (CCA) ports and allows collective components representing parallel applications to interact as one entity. We further describe a class of translation components, which translate between the distributed data format used by one parallel implementation to that used by another. A well known example of such components is the MxN component which translates between data distributed on M processors to data distributed on N processors. We describe its implementation in Parallel Application Work Space (PAWS), as well as the data structures PAWS uses to support it. We also present a mechanism allowing the framework to invoke this component on the programmer's behalf whenever such translation is necessary, freeing the programmer from treating collective component interactions as a special case. In doing that, we introduce framework-based, user-defined distributed type casts. Finally, we discuss our initial experiments in building optimized complex translation components out of atomic functionalities.
Kate Keahey, Patricia K. Fasel, Susan M. Mniszewski
HPDC2
1998 Efficient Coupling of Parallel Applications Using PAWS
abstract
PAWS (Parallel Application WorkSpace) is a software infrastructure for use in connecting separate parallel applications within a component-like model. A central PAWS Controller coordinates the linking of serial or parallel applications across a network to allow them to share parallel data structures such as multidimensional arrays. Applications use the PAWS API to indicate which data structures are to be shared and at what points the data is ready to be sent or received. PAWS implements a general parallel data descriptor and automatically carries out parallel layout remapping when necessary. Connections can be dynamically established and dropped, and can use multiple data transfer pathways between applications. PAWS uses the NEXUS communication library and is independent of the application's parallel communication mechanism.
Pete Beckman, Patricia K. Fasel, William F. Humphrey, Susan M. Mniszewski
HPDC2