VLDB 2026 Research / reviewers in the wild / expert
Michela Taufer
dblp:52/4034
· DBLP profile ↗
117ranked-venue papers
24as first author
48since 2021 · last 2026
0000-0002-0031-6377ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 79 · 20 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 17 since 2021Software engineering, systems software and programming languages · 23 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's RabbitsabstractModern HPC systems are placing increasing demands on job schedulers due to their scale and novel hardware. El Capitan’s Rabbit nodes exemplify this challenge: unlike traditional systems where storage is remote and shared, Rabbit nodes wire local NVMe SSDs directly to compute nodes via PCIe, forcing schedulers to actively track storage topology, capacity, and cross-job persistence, concerns they were never designed to handle. We introduce Flux Fiction, a fully plugin-based HPC system emulator built on top of Flux that replays historical job traces to evaluate scheduling policies in Flux. We validate Flux Fiction against the LLNL Tuolumne cluster using two workloads across four queueing policies, achieving a P99-bounded slowdown error below 1 in 7 of 8 experiments and a maximum utilization error of 1.2%. We then use Flux Fiction to explore Rabbit storage scheduling, demonstrating its ability to explore novel scheduling scenarios. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 10 |
| 2026 | Modernizing VPIC-Kokkos I/O: From Legacy Binary Output to Adaptive HDF5 WorkflowsabstractExascale particle-in-cell (PIC) simulations like Vector Particle-In-Cell (VPIC) face critical parallel input/output (I/O) bottlenecks due to legacy proprietary formats and storage bloat from redundant ghost cells. We present two architectural contributions: a Kokkos-aware staging pipeline that eliminates ghost cell padding – yielding a 67% reduction in grid-based export sizes – and a parallel Hierarchical Data Format 5 (HDF5) backend validated against h5bench in a weak-scaling study up to 896 MPI ranks. Our pipelined architecture achieves file-per-process (FPP) throughput comparable to the legacy binary format, outpacing monolithic bulk-writing benchmarks while isolating collective I/O (CIO) synchronization overheads, establishing a definitive performance baseline for emerging Exascale architectures. Connor Browne, Nigel Tan, Scott V. Luedtke, Michela Taufer, Brian J. Albright |
HPDC | 4 |
| 2026 | Merkle-Tree Weight Snapshot Deduplication for Provenance-Aware Auditing of Neural Network TrainingabstractWeight snapshots taken during neural network training provide a foundation for reproducibility and for understanding how models evolve during learning. They indicate whether networks progress toward higher accuracy or diverge toward poor generalization, yet their size and frequency impose severe storage and I/O burdens. As models scale, snapshots exhibit substantial cross-epoch redundancy, making them increasingly difficult to archive and analyze efficiently. We introduce a Merkle-tree deduplication pipeline that removes redundancy while exposing metadata about training dynamics. Chunking and deduplicating weights yields 70–80% storage savings across CIFAR-10/100 and protein diffraction datasets, outperforming list-based deduplication and per-snapshot compression baselines. Beyond space savings, Merkle-tree metadata categorizes chunks as fixed duplicates, shifted duplicates, or first occurrences. These signals predict validation accuracy with mean absolute error below 1% and provide an optional, metadata-driven signal to inform early stopping, enabling savings of 16–72% of the training epochs with negligible accuracy loss. Our work demonstrates that Merkle-tree deduplication provides a unified approach to reduce overhead, preserve reproducibility, and explain training dynamics within user-defined error tolerances, without disrupting the learning loop. Kin Wai Ng, Francesco Antici, Nigel Tan, Befikir Bogale, Caleb Han, Florence Tama, Osamu Miyashita, Bogdan Nicolae, Michela Taufer |
HPDC | 9 |
| 2026 | Editorial on future generation computer systems (FGCS) special collection on advances in quantum computing: methods, algorithms, and systems Vol II
Stefano Markidis, Lucio Grandinetti, Michela Taufer |
Future Gener. Comput. Syst. | 3 |
| 2026 | Recognition of best paper, outstanding editors, and reviewers for future generation computer systems in 2025
Michela Taufer |
Future Gener. Comput. Syst. | 1 |
| 2025 | Large Data Acquisition and Analytics at Synchrotron Radiation Facilities
Aashish Panta, Giorgio Scorzelli, Amy Ashurst Gooch, Werner Sun, Katherine S. Shanks, Suchismita Sarker, Devin Bougie, Keara Soloway, Rolf Verberg, Tracy Berman, Glenn Tarcea, John Allison, Michela Taufer, Valerio Pascucci |
IEEE Big Data | 13 |
| 2025 | ParaDyMS: Parallel Dynamic Motif Counting at ScaleabstractMotifs in graphs (or networks) are subgraphs induced by a small set of vertices such as triangles and cliques. The frequency of motifs is used to compare and align networks across various domains, such as biology, epidemiology, and social sciences. Recent advances have made it feasible to solve the computational challenge of counting motifs in networks with over a billion edges. However, these algorithms apply only for static networks where the structure remains unchanged. In reality, networks dynamically evolve, and understanding how motifs change with the dynamic nature of these networks remains an unsolved challenge despite its potential to provide essential insights into the system. We present ParaDyMS (Parallel Dynamic Motif Counting at Scale), the first parallel algorithm for updating motif counts in fully dynamic networks using batched updates. Our algorithm updates the frequencies of motifs only in the modified parts of the network instead of recomputing them from scratch. We provide proof of the algorithm's correctness and complexity and empirically compare its execution time with another state-of-the-art static algorithm on shared memory and GPUs using realworld networks. Our results show that our algorithm is highly scalable and can significantly reduce the time to compute motifs by more than 90% in the best case and, on average, by 69%. Nigel Tan, Jack D. Marquez, Michela Taufer, Sanjukta Bhowmick |
CCGrid | 4 |
| 2025 | Visualizing Temperature Hotspots and Their Impact Using the NEX-GDDP-CMIP6 DatasetabstractClimate events like tornadoes, hurricanes, and droughts are major twenty-first-century challenges. Mitigation efforts lag due to costly monitoring tools and complex data access. We present a dashboard with long-term temperature and weather trends, using nine global metrics from the NEX-GDDP-CMIP6 dataset, projecting to 2100. Our analysis flags countries facing major wet-bulb temperature hikes, linking these to expected economic impacts. Initial findings suggest that, in a worst-case scenario, some Northern Hemisphere nations may experience over a 10°C wet-bulb temperature rise between 2020 and 2090, potentially causing GDP losses exceeding 30%. Walter J. Ashworth, Jack D. Marquez, Michela Taufer |
eScience | 3 |
| 2025 | GEOtiled-SG: A Scalable Framework for High-Resolution Terrain Parameter ComputationabstractGEOtiled enables efficient computation of high-resolution terrain parameters from digital elevation models (DEMs) by decomposing large regions into smaller, parallelizable tiles. Originally developed as GEOtiled-G with support for three parameters (Slope, Aspect, Hillshade) via GDAL, we present GEOtiled-SG, an enhanced version that integrates the SAGA GIS library to compute over 15 parameters, expanding its utility in Earth science. To offset SAGA’s computational overhead, GEOtiled-SG introduces three optimizations: concurrent DEM cropping, buffer-aware mosaicking, and unified concurrency across workflow stages. Evaluations show that GEOtiled-SG maintains GEOtiled-G’s performance on the original parameters and offers consistent speedups across the expanded set. The framework is open source on GitHub, with data hosted on Dataverse, supporting reproducible, scalable terrain analysis. Gabriel Laboy, Ian Lumsden, Paula Olaya, Jack D. Marquez, Kin Wai Ng, Rodrigo Vargas, Michela Taufer |
eScience | 7 |
| 2025 | Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El CapitanabstractWe present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures. Ian Lumsden, Stefano Markidis, Andong Hu, Ivy Bo Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Thomas Scogland, Hariharan Devarajan, Bronis R. de Supinski, Gian Luca Delzanno, Michela Taufer |
eScience | 13 |
| 2025 | Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPCabstractEl Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 11 |
| 2025 | Advancing the GEOtiled Framework Through Scalable Terrain Parameter ComputationabstractThe GEOtiled framework facilitates the scalable and efficient computation of high-resolution terrain parameters using Digital Elevation Models (DEMs) across the Continental United States (CONUS). These parameters are essential in Earth Science applications. This paper presents significant advances to optimizing GEOtiled's performance. In GEOtiled 2.0, we introduce three key optimizations to speed up GEOtiled's overall runtime: concurrent cropping, efficient mosaic operation, and unified parallel processing. These optimizations improve data distribution and handling, reducing computation times while maintaining accuracy. We evaluate performance using DEM data at 30-meter resolution over a region covering the state of Tennessee. Our results demonstrate that the optimized framework provides substantial performance improvements for generating terrain parameters. GEOtiled 2.0 software, openly available on GitHub, advances reproducible, data-driven scientific discovery in Earth Science. Gabriel Laboy, Paula Olaya, Jack D. Marquez, Michael Sutherlin, Rodrigo Vargas, Michela Taufer |
HPDC | 6 |
| 2025 | On Optimizing Checkpoint Restoration for HPC Applications: Leveraging Merkle Trees and Asynchronous I/OabstractEfficient checkpoint restoration is critical in high-performance computing (HPC) and AI applications, where slow recovery times disrupt workflows, waste resources, and hinder reproducibility. This work introduces a Merkle tree checkpoint restoration method to accelerate failure recovery and improve explainability. Our method integrates asynchronous I/O via the Liburing library to optimize scattered reads in HPC applications. Tested on the Polaris system at Argonne National Laboratory, it exhibits lower restoration time and memory consumption than state-of-the-art checkpoint restoration methods, reaching near-full efficiency with duplicated data. Our work advances scalable and efficient checkpointing solutions for HPC, ensuring reliable and fast failure recovery for large-scale simulations. Zackary Malkmus, Nigel Tan, Ian Lumsden, Kevin Assogba, M. Mustafa Rafique, Bogdan Nicolae, Michela Taufer |
HPDC | 7 |
| 2025 | Cross-Architecture Performance Analysis Using the RAJA Performance SuiteabstractModern supercomputer architectures are diverse and becoming increasingly complex. Scientists are constantly porting code and re-optimizing it for the new architecture, but achieving good performance is challenging. Performance portability programming models such as RAJA, Kokkos, and OpenMP enable codes to maintain a single-source code rather than rewriting for each target architecture. However, portability models alone will not result in optimal performance as hardware has varying specifications (e.g., cache sizes and speeds) and parallel algorithms may use varying amounts of memory and compute resources. We present a systematic analysis of application behaviors across a diverse set of CPU and GPU hardware. We leverage the RAJA Performance Suite, which contains a curated set of kernels commonly found in HPC applications, to perform an in-depth GPU and memory analysis as well as a quantitative performance portability evaluation across different compute platforms. In analyzing the performance portability scores, we identify gaps and opportunities to achieve consistent performance across platforms. We provide a comprehensive analysis across seven architectures, including the most recent GPU systems with new physical memory layouts, where kernels demonstrate a runtime speedup of up to 44 ×. Although the speedup highlights the baseline improvements of newer hardware, the performance portability scores calculated, ranging from 0% to 92%, showcase where opportunities remain for scientists to increase utilization of the newer systems. Dewi Yokelson, Stephanie Brink, Jason Burmark, Michael McKinsey, Befikir Bogale, Ian Lumsden, Michela Taufer, Thomas Scogland, Olga Pearce |
ICPP | 7 |
| 2025 | Special Collection on Advances in Quantum Computing: Methods, Algorithms, and Systems
Stefano Markidis, Michela Taufer, Lucio Grandinetti |
Future Gener. Comput. Syst. | 2 |
| 2025 | Recognition of Best Paper, Outstanding Editors, and Reviewers for Future Generation Computer Systems in 2024
Michela Taufer |
Future Gener. Comput. Syst. | 1 |
| 2024 | PerSSD: Persistent, Shared, and Scalable Data with Node-Local Storage for Scientific Workflows in Cloud InfrastructureabstractComputational workflows need to retain data from both intermediate stages and final results to ensure the reproducibility and trustworthiness of scientific discoveries. While cloud infrastructure offers advantages like elasticity and automation, it compromises the persistence of intermediate data to ensure performance and reduce costs. Utilizing node-local storage can enhance performance but requires manual data transfers to persistent storage, making the technique impractical. To address these challenges, we propose a software architecture called Persistent, Shared, and Scalable Data (PerSSD) that integrates cloud operators and a Network File System (NFS) to make node-local data persistent and shareable across cloud nodes while ensuring performance. PerSSD outperforms traditional cloud object storage, achieving 35% reduction in the overall execution time of an earth science workflow, all while ensuring data persistence and shareability. Paula Olaya, Sophia Wen, Jay F. Lofstead, Michela Taufer |
IEEE Big Data | 4 |
| 2024 | Towards Affordable Reproducibility Using Scalable Capture and Comparison of Intermediate Multi-Run ResultsabstractEnsuring reproducibility in high-performance computing (HPC) applications is a significant challenge, particularly when nondeterministic execution can lead to untrustworthy results. Traditional methods that compare final results from multiple runs often fail because they provide sources of discrepancies only a posteriori and require substantial resources, making them impractical and unfeasible. This paper introduces an innovative method to address this issue by using scalable capture and comparing intermediate multi-run results. By capitalizing on intermediate checkpoints and hash-based techniques with user-defined error bounds, our method identifies divergences early in the execution paths. We employ Merkle trees for checkpoint data to reduce the I/O overhead associated with loading historical data. Our evaluations on the nondeterministic HACC cosmology simulation show that our method effectively captures differences above a predefined error bound and significantly reduces I/O overhead. Our solution provides a robust and scalable method for improving reproducibility, ensuring that scientific applications on HPC systems yield trustworthy and reliable results. Nigel Tan, Kevin Assogba, Walter J. Ashworth, Befikir Bogale, Franck Cappello, M. Mustafa Rafique, Michela Taufer, Bogdan Nicolae |
Middleware | 7 |
| 2024 | DYAD: Locality-aware Data Management for accelerating Deep Learning TrainingabstractDeep Learning (DL) is increasingly applied across various fields to solve complex scientific challenges in modern high-performance computing (HPC) systems that are beyond the reach of traditional algorithms. Training DL models for scientific applications involves processing multi-terabyte datasets in each epoch. The data access behavior during DL training exposes optimization opportunities to cache these datasets in near-compute storage accelerators in HPC systems, enhancing I/O throughput. However, current middleware solutions employ near-compute storage accelerators primarily as exclusive caches, which limits the effectiveness of cache access locality. To address this problem, we introduce DYAD, a system designed to maximize sample locality in the cache, thereby significantly increasing I/O throughput in HPC systems.DYAD optimizes I/O for DL training based on three key features. First, DYAD boosts inter-node access speeds by using a novel streaming RPC with RDMA protocol, achieving a 1.25x performance gain over state-of-the-art solutions. Second, DYAD further enhances inter-node access by coordinating data movement, which mitigates network congestion and increases throughput for inter-node accesses by up to 8.78x. Last, DYAD uses smart metadata caching that outperforms traditional global metadata access methods by several orders of magnitude in terms of lookup throughput. We demonstrate how DYAD accelerates large-scale DL training on a high-end HPC cluster with 512 GPUs by up to 10.82x faster epochs compared to UnifyFS by performing locality-aware caching on near-compute storage accelerators. Hariharan Devarajan, Ian Lumsden, Chen Wang 0004, Konstantia Georgouli, Thomas Scogland, Jae-Seung Yeom, Michela Taufer |
SBAC-PAD | 7 |
| 2024 | Recognition of best paper, outstanding editors, and outstanding reviewers for Future Generation Computer Systems in 2023
Michela Taufer |
Future Gener. Comput. Syst. | 1 |
| 2024 | Design Concerns for Integrated Scripting and Interactive Visualization in Notebook EnvironmentsabstractInteractive visualization can support fluid exploration but is often limited to predetermined tasks. Scripting can support a vast range of queries but may be more cumbersome for free-form exploration. Embedding interactive visualization in scripting environments, such as computational notebooks, provides an opportunity to leverage the strengths of both direct manipulation and scripting. We investigate interactive visualization design methodology, choices, and strategies under this paradigm through a design study of calling context trees used in performance analysis, a field which exemplifies typical exploratory data analysis workflows with Big Data and hard to define problems. We first produce a formal task analysis assigning tasks to graphical or scripting contexts based on their specificity, frequency, and suitability. We then design a notebook-embedded interactive visualization and validate it with intended users. In a follow-up study, we present participants with multiple graphical and scripting interaction modes to elicit feedback about notebook-embedded visualization design, finding consensus in support of the interaction model. We report and reflect on observations regarding the process and design implications for combining visualization and scripting in notebooks. Connor Scully-Allison, Ian Lumsden, Katy Williams, Jesse Bartels, Michela Taufer, Stephanie Brink, Abhinav Bhatele, Olga Pearce, Katherine E. Isaacs |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use CaseabstractScientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services. Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer |
CLOUD | 11 |
| 2023 | Online Boosted Gaussian Learners for In-Situ Detection and Characterization of Protein Folding States in Molecular Dynamics SimulationsabstractMolecular Dynamics (MD) simulations are a crucial tool for understanding how proteins fold. In its easiest form, MD simulations can be scaled through data parallelism, this means that multiple folding trajectories can be spawned and executed in parallel, facilitating a more efficient exploration of the protein folding space. However, due to data dependencies, the analysis of MD simulations remains largely as a centralized process. In this work, we propose a data parallel, lightweight technique to learn the characteristics of protein folding states in MD simulations. Contrary to other methods, ours can differentiate relevant states in a single protein folding trajectory without requiring centralized global knowledge of the protein dynamics. As its processing and memory overheads are negligible (in the order of milliseconds per window of frames, and kilo bytes respectively) this technique can be coupled with the simulation for in-situ analysis. Harshita Sahni, Hector Carrillo-Cabada, Ekaterina D. Kots, Silvina Caíno-Lores, Jack D. Marquez, Ewa Deelman, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
e-Science | 9 |
| 2023 | Thicket: Seeing the Performance Experiment Forest for the Individual Run TreesabstractThicket is an open-source Python toolkit for Exploratory Data Analysis (EDA) of multi-run performance experiments. It enables an understanding of optimal performance configuration for large-scale application codes. Most performance tools focus on a single execution (e.g., single platform, single measurement tool, single scale). Thicket bridges the gap to convenient analysis in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets by providing an interface for interacting with the performance data. Thicket has a modular structure composed of three components. The first component is a data structure for multi-dimensional performance data, which is composed automatically on the portable basis of call trees, and accommodates any subset of dimensions present in the dataset. The second is the metadata, enabling distinction and sub-selection of dimensions in performance data. The third is a dimensionality reduction mechanism, enabling analysis such as computing aggregated statistics on a given data dimension. Extensible mechanisms are available for applying analyses (e.g., top-down on Intel CPUs), data science techniques (e.g., K-means clustering from scikit-learn), modeling performance (e.g., Extra-P), and interactive visualization. We demonstrate the power and flexibility of Thicket through two case studies, first with the open-source RAJA Performance Suite on CPU and GPU clusters and another with a large physics simulation run on both a traditional HPC cluster and an AWS Parallel Cluster instance. Stephanie Brink, Michael McKinsey, David Böhme, Connor Scully-Allison, Ian Lumsden, W. Daryl Hawkins, Treece Burgess, Vanessa Lama, Jakob Lüttgau, Katherine E. Isaacs, Michela Taufer, Olga Pearce |
HPDC | 11 |
| 2023 | Studying Latency and Throughput Constraints for Geo-Distributed Data in the National Science Data FabricabstractThe National Science Data Fabric (NSDF) is our solution to the problem of addressing the data-sharing needs of the growing data science community. NSDF is designed to make sharing data across geographically distributed sites easier for users who lack technical expertise and infrastructure. By developing an easy-to-install software stack, we promote the FAIR data-sharing principles in NSDF while leveraging existing high-speed data transfer infrastructures such as Globus and XRootD. This work shows how we leverage latency and throughput information between geo-distributed NSDF sites with NSDF entry points to optimize the automatic coordination of data placement and transfer across the data fabric, which can further improve the efficiency of data sharing. Jakob Lüttgau, Heberth F. Martinez, Glenn Tarcea, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer |
HPDC | 6 |
| 2023 | GEOtiled: A Scalable Workflow for Generating Large Datasets of High-Resolution Terrain ParametersabstractTerrain parameters such as slope, aspect, and hillshading are essential in various applications, including agriculture, forestry, and hydrology. However, generating high-resolution terrain parameters is computationally intensive, making it challenging to provide these value-added products to communities in need. We present a scalable workflow called GEOtiled that leverages data partitioning to accelerate the computation of terrain parameters from digital elevation models, while preserving accuracy. We assess our workflow in terms of its accuracy and wall time by comparing it to SAGA, which is highly accurate but slow to generate results, and to GDAL, which supports memory optimizations but not data parallelism. We obtain a coefficient of determination (R^2) between GEOtiled and SAGA of 0.794, ensuring accuracy in our terrain parameters. We achieve an X6 speedup compared to GDAL when generating the terrain parameters at a high-resolution (10 m) for the Contiguous United States (CONUS). Camila Roa, Paula Olaya, Ricardo M. Llamas, Rodrigo Vargas, Michela Taufer |
HPDC | 5 |
| 2023 | Composable Workflow for Accelerating Neural Architecture Search Using In Situ Analytics for Protein ClassificationabstractNeural architecture search (NAS), which automates the design of neural network (NN) architectures for scientific datasets, requires significant computational resources and time — often on the order of days or weeks of GPU hours and training time. We design the Analytics for Neural Network (A4NN) workflow, a composable workflow that significantly reduces the time and resources required to design accurate and efficient NN architectures. We introduce a parametric fitness prediction strategy and distribute training across multiple accelerators to decrease the aggregated NN training time. A4NN rigorously record neural architecture histories, model states, and metadata to reproduce the search for near-optimal NNs. We demonstrate A4NN’s ability to reduce training time and resource consumption on a dataset generated by an X-ray Free Electron Laser (XFEL) experiment simulation. When deploying A4NN, we decrease training time by up to 37% and epochs required by up to 38%. Georgia Channing, Ria Patel, Paula Olaya, Ariel Keller Rorabaugh, Osamu Miyashita, Silvina Caíno-Lores, Catherine D. Schuman, Florence Tama, Michela Taufer |
ICPP | 9 |
| 2023 | Scalable Incremental Checkpointing using GPU-Accelerated De-DuplicationabstractWriting large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs’ high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques. Nigel Tan, Jakob Lüttgau, Jack D. Marquez, Keita Teranishi, Nicolas M. Morales, Sanjukta Bhowmick, Franck Cappello, Michela Taufer, Bogdan Nicolae |
ICPP | 8 |
| 2023 | Performance assessment of ensembles of in situ workflows under resource constraintsabstractSummary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node. Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Silvina Caíno-Lores, Michela Taufer, Ewa Deelman |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Recognition of Outstanding Future Generation Computer Systems Reviewers for 2022
Michela Taufer |
Future Gener. Comput. Syst. | 1 |
| 2023 | Building Trust in Earth Science Findings through Data Traceability and Results ExplainabilityabstractTo trust findings in computational science, scientists need workflows that trace the data provenance and support results explainability. As workflows become more complex, tracing data provenance and explaining results become harder to achieve. In this paper, we propose a computational environment that automatically creates a workflow execution's record trail and invisibly attaches it to the workflow's output, enabling data traceability and results explainability. Our solution transforms existing container technology, includes tools for automatically annotating provenance metadata, and allows effective movement of data and metadata across the workflow execution. We demonstrate the capabilities of our environment with the study of SOMOSPIE, an earth science workflow. Through a suite of machine learning modeling techniques, this workflow predicts soil moisture values from the 27 km resolution satellite data down to higher resolutions necessary for policy making and precision agriculture. By running the workflow in our environment, we can identify the causes of different accuracy measurements for predicted soil moisture values in different resolutions of the input data and link different results to different machine learning methods used during the soil moisture downscaling, all without requiring scientists to know aspects of workflow design and implementation. Paula Olaya, Dominic Kennedy, Ricardo M. Llamas, Leobardo Valera, Rodrigo Vargas, Jay F. Lofstead, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | Augmenting Singularity to Generate Fine-grained Workflows, Record Trails, and Data ProvenanceabstractThe use of containerization technology in high performance computing (HPC) workflows has substantially increased recently because it makes workflows much easier to develop and deploy. Although many HPC workflows include multiple data and multiple applications, they have traditionally all been bundled together into one monolithic container. This hinders the ability to trace the thread of execution, thus preventing scientists from establishing data provenance, or having workflow reproducibility. To provide a solution to this problem we extend the functionality of a popular HPC container runtime, Singularity. We implement both the ability to compose fine-grained containerized workflows and execute these workflows within the Singularity runtime with automatic metadata collection. Specifically, the new functionality collects a record trail of execution and creates data provenance. The use of our augmented Singularity is demonstrated with an earth science workflow, SOMOSPIE. The workflow is composed via our augmented Singularity which creates fine-grained containers and collects the metadata to trace, explain, and reproduce the prediction of soil moisture at a fine resolution. Dominic Kennedy, Paula Olaya, Jay F. Lofstead, Rodrigo Vargas, Michela Taufer |
e-Science | 5 |
| 2022 | Enabling Call Path Querying in Hatchet to Identify Performance Bottlenecks in Scientific ApplicationsabstractAs computational science applications benefit from larger-scale, more heterogeneous high performance computing (HPC) systems, the process of studying their performance becomes increasingly complex. The performance data analysis library Hatchet provides some insights into this complexity, but is currently limited in its analysis capabilities. Missing capabilities include the handling of relational caller-callee data captured by HPC profilers. To address this shortcoming, we augment Hatchet with a Call Path Query Language that leverages relational data in the performance analysis of scientific applications. Specifically, our Query Language enables data reduction using call path pattern matching. We demonstrate the effectiveness of our Query Language in identifying performance bottlenecks and enhancing Hatchet's analysis capabilities through three case studies. In the first case study, we compare the performance of sequential and multi-threaded versions of the graph alignment application Fido. In doing so, we identify the existence of large memory inefficiencies in both versions. In the second case study, we examine the performance of MPI calls in the linear algebra mini-application AMG2013 when using MVAPICH and Spectrum-MPI. In doing so, we identify hidden performance losses in specific MPI functions. In the third case study, we illustrate the use of our Query Language in Hatchet's interactive visualization. In doing so, we show that our Query Language enables a simple and intuitive way to massively reduce profiling data. Ian Lumsden, Jakob Lüttgau, Vanessa Lama, Connor Scully-Allison, Stephanie Brink, Katherine E. Isaacs, Olga Pearce, Michela Taufer |
e-Science | 8 |
| 2022 | Reproducing and Extending Analytical Performance Models of Generalized Hierarchical SchedulingabstractWorkflows in High-Performance Computing (HPC) are rapidly changing towards more complex and large-scale workflows. In particular, high-throughput and ensemble workflows are becoming increasingly common. These workflows impose significant burden on current HPC scheduling systems which typically use slow, centralized schedulers. Generalized hierarchical scheduling (GHS) is a potential solution to face modern workflows but is not widely adopted in HPC yet. One difficulty hindering widespread adoption is the lack of performance models to configure and fit application requirements. The few existing models are often built on stick assumptions that can substantially reduce the analysis realism. In this paper, we reproduce the analysis and improve the realism of a state-of-the-art model presented in “An Analytical Performance Model of Generalized Hierarchical Scheduling” [1] by Herbein and co-authors. Specifically, we first reproduce four key analysis studies in the original paper and then expand the model by removing key assumptions, one at a time. In doing so, we extend the realism of the original model. We empirically validate our extended model using three different scenarios and discuss the observed accuracy. Jakob Lüttgau, Silvina Caíno-Lores, Kae Suarez, Dong H. Ahn, Stephen Herbein, Michela Taufer |
e-Science | 6 |
| 2022 | Identifying Structural Properties of Proteins from X-ray Free Electron Laser Diffraction PatternsabstractCapturing structural information of a biological molecule is crucial to determine its function and understand its mechanics. X-ray Free Electron Lasers (XFEL) are an experimental method used to create diffraction patterns (images) that can reveal structural information. In this work we design, implement, and evaluate XPSI (X-ray Free Electron Laser-based Protein Structure Identifier), a framework capable of predicting three structural properties in molecules (i.e., orientation, conformation, and protein type) from their diffraction patterns. XPSI predicts these properties with high accuracy in challenging scenarios, such as recognizing orientations despite symmetries in diffraction patterns, distinguishing conformations even when they have similar structures, and identifying protein types under different noise conditions. Our framework shows low computational cost and high prediction accuracy compared to other machine learning methods such as random forest and neural networks. Paula Olaya, Silvina Caíno-Lores, Vanessa Lama, Ria Patel, Ariel Keller Rorabaugh, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 8 |
| 2022 | A Methodology to Generate Efficient Neural Networks for Classification of Scientific DatasetsabstractNeural networks (NNs) are increasingly utilized in high-throughput scientific workflows. In this context, NN efficiency is essential for successful workflow management. We use a multi-objective Neural Architecture Search (NAS), NSGA-Net, to search for highly accurate NNs while optimizing for efficient use of computational resources by minimizing FLoating-point Operations Per Second (FLOPS). We define a domain-agnostic methodology to generate NNs with the support of NSGA-Net, select promising NNs that balance accuracy and FLOPS usage, and refine a subset of NNs in order to curate networks suitable for efficient data analysis. We apply this methodology to a protein diffraction use case. Preliminary results show NNs that efficiently classify conformation of proteins with a final accuracy of 97.7% or higher and using only 187 FLOPS. Ria Patel, Ariel Keller Rorabaugh, Paula Olaya, Silvina Caíno-Lores, Georgia Channing, Catherine D. Schuman, Osamu Miyashita, Florence Tama, Michela Taufer |
e-Science | 9 |
| 2022 | The Materials Commons Data RepositoryabstractRepositories are increasingly used for publishing and sharing scientific data. The Materials Commons is a data repository that follows the FAIR (Findable, Accessible, Inter-operable, Reusable) principles. We demonstrate the challenges with FAIR and how Materials Commons solves them. We also discuss the Nationals Science Data Fabric (NSDF) [1], a project that is democratizing data access, and show how Materials Commons with the NSDF software stack accelerates data access and scientific research. Glenn Tarcea, Brian Puchala, Tracy Berman, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer, John Allison |
e-Science | 6 |
| 2022 | Ubique: A New Model for Untangling Inter-task Data Dependence in Complex HPC WorkflowsabstractExploiting task parallelism is getting increasingly difficult for diverse and complex scientific workflows running on High Performance Computing (HPC) systems. In this paper, we argue that the difficulty rises from a void in the spectrum of existing data-transfer models for resolving inter-task data dependence within a workflow and propose a novel model to fill that gap: Ubique. The Ubique model combines the best from in-transit and in situ models in order for loosely coupled producer and consumer tasks to run concurrently and to resolve their data dependencies efficiently with little or no modifications to their codes, striking a balance between transparent optimization, productivity, and performance. Our preliminary evaluation suggests that Ubique can significantly outperform the parallel file system (PFS)-based model while offering automatic data transfer and synchronization which are the features lacking in many traditional models. It also identifies the performance characteristics of its key depending subsystems, which must be understood for further broadening its benefits. Jae-Seung Yeom, Dong H. Ahn, Ian Lumsden, Jakob Lüttgau, Silvina Caíno-Lores, Michela Taufer |
e-Science | 6 |
| 2022 | NSDF-Cloud: Enabling Ad-Hoc Compute Clusters Across Academic and Commercial CloudsabstractComputational resources are increasingly provisioned to users through cloud-like interfaces. Both academic and commercial cloud offerings exist, but no single standardized interface for common actions such as configuration, launching, and termination of virtual resources exists. This imposes huge technical burden on domain scientist that attempt to take advantage of these resources; even expert users spend considerable time to port their applications from one cloud platform to another. Jakob Lüttgau, Paula Olaya, Naweiluo Zhou, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer |
HPDC | 6 |
| 2022 | NSDF-FUSE: A Testbed for Studying Object Storage via FUSE File SystemsabstractThis work presents NSDF-FUSE, a testbed for evaluating settings and performance of FUSE-based file systems on top of S3-compatible object storage; the testbed is part of a suite of services from the National Science Data Fabric (NSDF) project (an NSF-funded project that is delivering cyberinfrastructures for data scientists). We demonstrate how NSDF-FUSE can be deployed to evaluate eight different mapping packages that mount S3-compatible object storage to a file system, as well as six data patterns representing different I/O operations on two cloud platforms. NSDF-FUSE is open-source and can be easily extended to run with other software mapping packages and different cloud platforms. Paula Olaya, Jakob Lüttgau, Naweiluo Zhou, Jay F. Lofstead, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer |
HPDC | 7 |
| 2022 | VPIC 2.0: Next Generation Particle-in-Cell SimulationsabstractVPIC is a general purpose particle-in-cell simulation code for modeling plasma phenomena such as magnetic reconnection, fusion, solar weather, and laser-plasma interaction in three dimensions using large numbers of particles. VPIC's capacity in both fidelity and scale makes it particularly well-suited for plasma research on pre-exascale and exascale platforms. In this article, we demonstrate the unique challenges involved in preparing the VPIC code for operation at exascale, outlining important optimizations to make VPIC efficient on accelerators. Specifically, we show the work undertaken in adapting VPIC to exploit the portability-enabling framework Kokkos and highlight the enhancements to VPIC's modeling capabilities to achieve performance at exascale. We assess the achieved performance-portability trade-off through a suite of studies on nine different varieties of modern pre-exascale hardware. Our performance-portability study includes weak-scaling runs on three of the top ten TOP500 supercomputers, as well as a comparison of low-level system performance of hardware from four different vendors. Robert F. Bird, Nigel Tan, Scott V. Luedtke, Stephen Lien Harrell, Michela Taufer, Brian J. Albright |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Building High-Throughput Neural Architecture Search Workflows via a Decoupled Fitness Prediction EngineabstractNeural networks (NN) are used in high-performance computing and high-throughput analysis to extract knowledge from datasets. Neural architecture search (NAS) automates NN design by generating, training, and analyzing thousands of NNs. However, NAS requires massive computational power for NN training. To address challenges of efficiency and scalability, we proposePENGUIN, a decoupled fitness prediction engine that informs the search without interfering in it.PENGUINuses parametric modeling to predict fitness of NNs. Existing NAS methods and parametric modeling functions can be plugged intoPENGUINto build flexible NAS workflows. Through this decoupling and flexible parametric modeling,PENGUINreduces training costs: it predicts the fitness of NNs, enabling NAS to terminate training NNs early. Early termination increases the number of NNs that fixed compute resources can evaluate, thus giving NAS additional opportunity to find better NNs. We assess the effectiveness of our engine on 6,000 NNs across three diverse benchmark datasets and three state of the art NAS implementations using the Summit supercomputer. Augmenting these NAS implementations withPENGUINcan increase throughput by a factor of 1.6 to 7.1. Furthermore, walltime tests indicate thatPENGUINcan reduce training time by a factor of 2.5 to 5.3. Ariel Keller Rorabaugh, Silvina Caíno-Lores, J. Travis Johnston, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Developers' PerspectiveabstractScientists using the high-throughput computing (HTC) paradigm for scientific discovery rely on complex software systems and heterogeneous architectures that must deliver robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Developers must overcome a variety of obstacles to pursue workflow interoperability, identify tools and libraries for robust science, port codes across different architectures, and establish trust in non-deterministic results. This poster presents recommendations to build a roadmap to overcome these challenges and enable robust science for HTC applications and workflows. The findings were collected from an international community of software developers during a Virtual World Cafe in May 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall, Miron Livny |
CLUSTER | 1 |
| 2021 | A Case Study in Scientific Reproducibility from the Event Horizon Telescope (EHT)abstractThis poster presents the first results of an interdisciplinary project aiming to develop and share sustainable knowledge necessary to analyze, understand, and use published scientific results to advance reproducibility in multi-messenger astrophysics. Specifically, the project targets breakthrough work associated with the First M87 Event Horizon Telescope (EHT) and delivers recommendations on how the published results of the first black hole can be effectively reproduced. The project has the potential to advance new discovery in multi-messenger astrophysics by providing guidance for generalizing methods and findings from use cases. Ross Ketron, Jacob Leonard, Brandan Roachell, Ria Patel, R. White, Silvina Caíno-Lores, Nigel Tan, Patrick R. Miles, Karan Vahi, Ewa Deelman, Duncan A. Brown, Michela Taufer |
e-Science | 12 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Scientists' PerspectiveabstractThis poster presents our first steps to define a roadmap to robust science for high-throughput applications used in scientific discovery. These applications combine multiple components into increasingly complex multi-modal workflows that are often executed in concert on heterogeneous systems. The increasing complexity hinders the ability of scientists to generate robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Scientists must withstand and overcome adverse conditions such as heterogeneous and unreliable architectures at all scales (including extreme scale), rigorous testing under uncertainties, unexplainable algorithms in machine learning, and black-box methods. This poster presents findings and recommendations to build a roadmap to overcome these challenges and enable robust science. The data was collected from an international community of scientists during a virtual world café in February 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall |
e-Science | 1 |
| 2021 | AI4IO: A Suite of Ai-Based Tools for IO-Aware HPC Resource ManagementabstractHigh performance computing (HPC) is undergoing many changes at the system level. While scientific applications can reach petaflops or more in computing performance, potentially resulting in larger data generation rates and more frequent checkpointing, the data movement to the parallel file system remains costly due to constraints imposed by HPC centers on the IO bandwidth. In other words, the bandwidth to file systems is outpaced by the rate of data generation; the associated IO contention increases job runtime and delays execution. This situation is aggravated by the fact that when users submit their jobs to a HPC system, they rely on resource managers and job schedulers to monitor and manage the computing resources (i.e., nodes). Both resource managers and job schedulers remain blind to the impact of IO contention on the overall simulation performance. In this talk we discuss how Artificial Intelligence (AI) can augment HPC systems to prevent and mitigate IO contention while dealing with IO bandwidth constraints. Our solution, called Analytics for IO (AI4IO), consists of a suite of AI-based tools that enable IO-awareness on HPC systems. Specifically, we present two AI4IO tools: PRIONN and CanarIO. PRIONN automates predictions about usersubmitted job resource usage, including per-job IO bandwidth; CanarIO detects, in real-time, the presence of IO contention on HPC systems and predicts which jobs are affected by that contention (e.g., because of their frequent checkpointing). By working in concert, PRIONN and CanarIO predict the a priori knowledge necessary to prevent and mitigate IO contention with IO-aware scheduling. We integrate AI4IO in the Flux scheduler and show how A4IO produce improvements in simulation performance: we observe up to 6.2% improvement in makespan of HPC job workloads, which amounts to more than 18,000 node-hours saved per week on a production-size cluster. Our work is the first step to implementing IO-aware scheduling on production HPC systems. Michela Taufer |
HiPC | 1 |
| 2021 | A Graphic Encoding Method for Quantitative Classification of Protein Structure and Representation of Conformational ChangesabstractIn order to successfully predict a proteins function throughout its trajectory, in addition to uncovering changes in its conformational state, it is necessary to employ techniques that maintain its 3D information while performing at scale. We extend a protein representation that encodes secondary and tertiary structure into fix-sized, color images, and a neural network architecture (called GEM-net) that leverages our encoded representation. We show the applicability of our method in two ways: (1) performing protein function prediction, hitting accuracy between 78 and 83 percent, and (2) visualizing and detecting conformational changes in protein trajectories during molecular dynamics simulations. Hector Carrillo-Cabada, Jeremy Benson, Asghar M. Razavi, Brianna Mulligan, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Identifying Degree and Sources of Non-Determinism in MPI Applications Via Graph KernelsabstractAs the scientific community prepares to deploy an increasingly complex and diverse set of applications on exascale platforms, the need to assess reproducibility of simulations and identify the root causes of reproducibility failures increases correspondingly. One of the greatest challenges facing reproducibility issues at exascale is the inherent non-determinism at the level of inter-process communication. The use of non-deterministic communication constructs is necessary to boost performance, but communication non-determinism can also hamper software correctness and result reproducibility. To address this challenge, we propose a software framework for identifying the percentage and sources of communication non-determinism. We model parallel executions as directed graphs and leverage graph kernels to characterize run-to-run variations in inter-process communication. We demonstrate the effectiveness of graph kernel similarity as a proxy for non-determinism, by showing that these kernels can quantify the type and degree of non-determinism present in communication patterns. To demonstrate our framework's ability to link and quantify runtime non-determinism to root sources, demonstrate with present for an adaptive mesh refinement application, where our framework automatically quantifies the impact of function calls on non-determinism, and a Monte Carlo application, where our framework automatically quantifies the impact of parameter configurations on non-determinism. Dylan Chapp, Nigel Tan, Sanjukta Bhowmick, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | CanarIO: Sounding the Alarm on IO-Related Performance DegradationabstractUsers interact with High Performance Computing (HPC) machines through batch systems, which take user job submissions and allocate them to computing resources. While some resource managers have a generalized resource model, in nearly all modern systems, nodes are the only resource managed. Other resources, such as parallel file systems, are also necessary for jobs to make progress, but schedulers are blind to these resources. Facility staff can manually detect critical problems and manually hold jobs that need particular file systems, but this requires manual monitoring. Without human intervention, modern schedulers will happily run jobs whose required resources are not available. As a result, resources are wasted when IO-intensive jobs are scheduled on file systems with degraded performance.We introduce CanarIO, a tool for predicting the IO-sensitivity of HPC jobs and detecting IO-related performance degradation on HPC systems. CanarIO uses a set of "canary" IO probes run at regular intervals on the system. Using performance measurements from these jobs, CanarIO builds classifiers that can determine which jobs are IO-sensitive and when file system performance is degraded. We demonstrate the accuracy of our tool with a simulation of system execution using real HPC data. Specifically, we detect 37.5% of IO degradation events and correctly identify >90% of IO-sensitive jobs. We show that with CanarIO predictions we recover >1,500 node-hours in 10 days, with a potential maximum of nearly 10,000 node-hours. CanarIO is the first step necessary for augmenting schedulers to be resource-aware. Michael R. Wyatt II, Stephen Herbein, Kathleen Shoga, Todd Gamblin, Michela Taufer |
IPDPS | 5 |
| 2020 | Flux: Overcoming scheduling challenges for exascale workflows
Dong H. Ahn, Ned Bass, Albert Chu, Jim Garlick, Mark Grondona, Stephen Herbein, Helgi I. Ingólfsson, Joe Koning, Tapasya Patki, Thomas Scogland, Becky Springmeyer, Michela Taufer |
Future Gener. Comput. Syst. | 12 |
| 2020 | Memory-Efficient and Skew-Tolerant MapReduce Over MPI for Supercomputing SystemsabstractData analytics has become an integral part of large-scale scientific computing. Among various data analytics frameworks, MapReduce has gained the most traction. Although some efforts have been made to enable efficient MapReduce for supercomputing systems, they are often limited to fairly homogeneous workloads where equal partitioning of input data across tasks results in essentially equal output or temporary data generated on each task. For workloads that are more skewed, however, current implementations can result in imbalance in memory usage and, consequently, can cause a slowdown in execution time and a loss in data scalability. To tackle this problem, we enhance a previously published memory-conscious MapReduce over MPI framework called Mimir. Our enhancements to Mimir include combiner and dynamic repartition optimizations to minimize and balance memory usage and to achieve close to optimal balance of the memory usage across processes and to reduce the execution time by up to 12 times. Experimental results show that Mimir can scale to at least 3072 processes on the Tianhe-2 supercomputer on skewed datasets. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2019 | SOMOSPIE: A Modular SOil MOisture SPatial Inference Engine Based on Data-Driven DecisionsabstractThe current availability of soil moisture data over large areas comes from satellite remote sensing technologies (i.e., radar-based systems), but these data have coarse resolution and often exhibit large spatial information gaps. Where data are too coarse or sparse for a given need (e.g., precision farming), one can leverage machine-learning techniques coupled with other sources of environmental information (e.g., topography) to generate gap-free information at a finer spatial resolution (i.e., increased granularity). To this end, we develop a spatial inference engine consisting of modular stages for processing spatial environmental data, generating predictions with machine-learning techniques, and analyzing these predictions. We demonstrate the functionality of this approach and the effects of data processing choices via multiple prediction maps over a United States ecological region with a highly diverse soil moisture profile (i.e., the Middle Atlantic Coastal Plains). The relevance of our work derives from a pressing need to improve the spatial representation of soil moisture for applications in environmental sciences (e.g., ecological niche modeling, carbon monitoring systems, and other Earth system models) and precision farming (e.g., optimizing irrigation practices and other land management decisions). Danny Rorabaugh, Mario Guevara, Ricardo M. Llamas, Joy Kitson, Rodrigo Vargas, Michela Taufer |
eScience | 6 |
| 2019 | Characterizing In Situ and In Transit Analytics of Molecular Dynamics Simulations for Next-Generation SupercomputersabstractMolecular Dynamics (MD) simulations executed on state-of-the-art supercomputers are producing data at rates faster than it can be written out to disk. In situ and in transit analysis of data generated by MD simulations reduce the original volume of information by several orders of magnitude, thereby alleviating the negative impact of I/O bottlenecks. This work focuses on characterizing the impact of in situ and in transit analytics on the overall MD workflow performance, and the capability for capturing rapid, rare events in the simulated molecular system. The MD simulation and analysis processes share data via remote direct memory access (RDMA) using DataSpaces. Our metrics of interest are time spent waiting in I/O by the MD simulation, lost frames of the MD simulation, and idle time of the analysis. We measure these metrics for a diverse set of molecular systems and characterize their trends for in situ and in transit configurations. We then model which frames are dropped and which ones are analyzed for a real use case. The insights gained from this study are generally applicable for in situ and in transit workflows that require optimization of parameters to minimize loss in workflow performance and analytic accuracy. Michela Taufer, Ewa Deelman, Michael R. Wyatt II, Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Harel Weinstein, Michel A. Cuendet, Trilce Estrada |
eScience | 1 |
| 2019 | Session details: Reliability and VariabilityabstractNo abstract available. Michela Taufer |
HPDC | 1 |
| 2018 | On the Power of Combiner Optimizations in MapReduce Over MPI WorkflowsabstractAnalyzing large volumes of data is becoming more and more important in various scientific computing domains. MapReduce over MPI frameworks are an appealing solution to enable scalable big data analytics on supercomputing systems. These systems can further leverage features of MapReduce applications by merging (key/value) pairs before the reduce function in combiner optimizations. In this paper, we propose a pipeline combiner workflow and integrate it into Mimir, a cutting-edge implementation of Map Reduce over MPI. Our results with real datasets on the Tianhe-2 supercomputer prove that our pipeline combiner workflow can reduce memory usage up to 51% and improve the overall performance up to 61%. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
ICPADS | 7 |
| 2018 | KeyBin2: Distributed Clustering for Scalable and In-Situ AnalysisabstractWe present KeyBin2, a key-based clustering method that is able to learn from distributed data in parallel. KeyBin2 uses random projections and discrete optimizations to efficiently clustering very high dimensional data. Because it is based on keys computed independently per dimension and per data point, KeyBin2 scales linearly. We perform accuracy and scalability tests to evaluate our algorithm's performance using synthetic and real datasets. The experiments show that KeyBin2 outperforms other parallel clustering methods for problems with increased complexity. Finally, we present an application of KeyBin2 for in-situ clustering of protein folding trajectories. Jeremy Benson, Matt Peterson, Michela Taufer, Trilce Estrada |
ICPP | 4 |
| 2018 | PRIONN: Predicting Runtime and IO using Neural NetworksabstractFor job allocation decision, current batch schedulers have access to and use only information on the number of nodes and runtime because it is readily available at submission time from user job scripts. User-provided runtimes are typically inaccurate because users overestimate or lack understanding of job resource requirements. Beyond the number of nodes and runtime, other system resources, including IO and network, are not available but play a key role in system performance. There is the need for automatic, general, and scalable tools that provide accurate resource usage information to schedulers so that, by becoming resource-aware, they can better manage system resources. Michael R. Wyatt II, Stephen Herbein, Todd Gamblin, Adam Moody, Dong H. Ahn, Michela Taufer |
ICPP | 6 |
| 2018 | Modeling the Next-Generation High Performance SchedulersabstractHigh performance computing (HPC) resources and workloads are undergoing tumultuous changes. HPC resources are growing more diverse with the adoption of accelerators; HPC workloads have increased in size by orders of magnitude. Despite these changes, when assigning workload jobs to resources, HPC schedulers still rely on users to accurately anticipate their applications' resource usage and remain stuck with the decades-old centralized scheduling model. In this talk we will discuss these ongoing changes and propose alternative models for HPC scheduling based on resource-awareness and fully hierarchical models. A key role in our models' evaluation is played by an emulator of a real open-source, next-generation resource management system. We will discuss the challenges of realistically mimicking the system's scheduling behavior. Our evaluation shows how our models improve scheduling scalability on a diverse set of synthetic and real-world workloads. This is joint work with Stephen Herbein and Michael Wyatt at the University of Delaware, and Tapasya Patki, Dong H. Ahn, Don Lipari, Thomas R.W. Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Tamara Dahlgren, David Domyancic, and Becky Springmeyer at the Lawrence Livermore National Laboratory. Michela Taufer |
SIGSIM-PADS | 1 |
| 2017 | Data analytics for modeling soil moisture patterns across united states ecoclimatic domainsabstractOur poster presents a data analytics strategy to enable scientists to model patterns of soil moisture data at different resolutions across the United States. We build upon previous work of Guevara and co-authors with three contributions. First, we introduce divisions of soil moisture into the climatic regions proposed by the National Ecology Observatory Network. Second, we reduce the topological parameters used in modeling soil moisture using Principal Component Analysis. Third, we present an efficient workflow for modeling and visualizing soil moisture data. Thomas Kitson, Paula Olaya, Elizabeth Racca, Michael R. Wyatt II, Mario Guevara, Rodrigo Vargas, Michela Taufer |
IEEE BigData | 7 |
| 2017 | Bloomfish: A Highly Scalable Distributed K-mer Counting FrameworkabstractK-mer counting is a fundamental operation in DNA research and genome analytics; its application includes estimating genome assembly, understanding similarities in genomic samples, and merging a newly processed genome with a reference genome. As the genome dataset becomes larger and larger, designing a highly optimized distributed-memory implementation becomes more and more important. Current distributed-memory solutions have two limitations: they have a high memory footprint, and they do not provide advanced optimizations for loading enormous genome datasets into memory. Based on these observations, we present Bloomfish, a distributed, memory-efficient, scalable solution to the limits of current work. To keep a low memory footprint, Bloomfish leverages the compact hash array design of the single-node Jellyfish system and the optimized workflow of the high-performance MapReduce framework Mimir. We have also codesigned Mimir's I/O to efficiently load enormous datasets. We ran Bloomfish on the Tianhe-2 supercomputer with large sequence datasets (up to 24 TB). Our results show that Bloomfish achieves unprecedented scalability in genome analytics. Yanfei Guo, Yanjie Wei, Bingqiang Wang, Yutong Lu, Pietro Cicotti, Pavan Balaji, Michela Taufer |
ICPADS | 8 |
| 2017 | Mimir: Memory-Efficient and Scalable MapReduce for Large Supercomputing SystemsabstractIn this paper we present Mimir, a new implementation of MapReduce over MPI. Mimir inherits the core principles of existing MapReduce frameworks, such as MR-MPI, while redesigning the execution model to incorporate a number of sophisticated optimization techniques that achieve similar or better performance with significant reduction in the amount of memory used. Consequently, Mimir allows significantly larger problems to be executed in memory, achieving large performance gains. We evaluate Mimir with three benchmarks on two highend platforms to demonstrate its superiority compared with that of other frameworks. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
IPDPS | 7 |
| 2017 | Enabling scalable and accurate clustering of distributed ligand geometries on supercomputers
Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Pavan Balaji, Michela Taufer |
Parallel Comput. | 5 |
| 2016 | Machine Learning Predictions of Runtime and IO Traffic on High-End ClustersabstractWe use supervised machine learning algorithms (i.e., Decision Trees, Random Forest, and K-nearest Neighbors) to predict performance characteristics such as runtime and IO traffic of batch jobs on high-end clusters, using only user job scripts as input. We show that decision trees outperform other algorithms and accurately predict the runtime of 73% of jobs within a error tolerance of 10 minutes, which is a 51% improvement over the user requested runtime. Ryan McKenna, Stephen Herbein, Adam Moody, Todd Gamblin, Michela Taufer |
CLUSTER | 5 |
| 2016 | Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC ClustersabstractThe economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-generation resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads. Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela Taufer |
HPDC | 9 |
| 2016 | Study of Neocortex Simulations with GENESIS on High Performance Computing ResourcesabstractOne significant challenge in neuroscience is understanding the cooperative behavior of large numbers of neurons. Models of neuronal networks allow scientists to explore the impact of differential neuronal connectivity using analysis techniques and information not available experimentally. However, modeling realistic neurobiological processes and encoding them in computer simulations is challenging, as increasing computing and data requirements are all of concern. In this work we study the performance of neocortex simulations using GEneral NEural SImulation System (GENESIS), a well-known multi-function brain simulation package, supported by high performance computing resources. The contribution of our work is threefold. First, we study the impact of platforms (i.e., single fat nodes versus high-end clusters) and their features on the performance and data generation for a small scale model of neocortex. Second, we assess the impact of the model complexity (i.e., number of cells and cell connectivity)on the performance and data generation for increasingly large versions of the neocortex model on high-end clusters. Third, we provide selected scientific results obtained by increasing the model complexity. We show that the more realistic and rigorous modeling of neocortex functions is computationally feasible but requires high performance computing resources to mitigate the growing computing and data requests of the associated simulations. Sean McDaniel, David L. Boothe, Joshua C. Crone, Song Jun Park, Dale R. Shires, Alfred B. Yu, Michela Taufer |
ICPADS | 7 |
| 2016 | Scheduling Matters: Area-Oriented Heuristic for Resource ManagementabstractParallel and distributed systems that provide compute resources on demand are convenient, cost-effective, and becoming increasingly common. Boosting workload performance in such environments through scheduling has been of great interest, as users and providers aim to increase parallelism and reduce execution times. For modern data centers, leaving a smaller carbon footprint while maintaining high performance and low cost is becoming the next big challenge. With this in mind, we analyze the relative impacts on resource utilization of three well-motivated platform-oblivious scheduling heuristics. We simulate over 50,000 DAG workflow executions and measure performance, cost, and resource utilization under the three scheduling heuristics. Our results provide insights to better enable high-performance execution of workflows and advanced capacity planning while increasing resource utilization and reducing costs. Jeremy Benson, Trilce Estrada, Arnold L. Rosenberg, Michela Taufer |
SBAC-PAD | 4 |
| 2016 | HYPPO: A Hybrid, Piecewise Polynomial Modeling Technique for Non-Smooth SurfacesabstractThe number and diversity of tunable parameters in applications makes predicting settings that achieve optimal performance challenging. Complicating matters is the fact that resources are increasingly shared among computational tasks (for example, in cloud environments). Choosing any setting that yields near-optimal performance runs the risk of overusing shared resources. Building accurate models that capture the complicated interplay of parameters is crucial in order to maximize performance with minimal resource impact. Traditional techniques tend to fall short when modeling performance. One reason is that performance surfaces are often irregular but most traditional techniques are designed to produce smooth models. In this paper we introduce a hybrid modeling technique that combines the strengths of surrogate-based modeling (SBM) and k nearest-neighbor regression (kNN) into a single method called HYPPO. The hybrid method is a piecewise polynomial model composed of many small, local models. We demonstrate that HYPPO significantly improves overall prediction accuracy compared with SBM and kNN. J. Travis Johnston, Connor Zanin, Michela Taufer |
SBAC-PAD | 3 |
| 2016 | Performance characterization of irregular I/O at the extreme scale
Stephen Herbein, Sean McDaniel, Norbert Podhorszki, Jeremy Logan, Scott Klasky, Michela Taufer |
Parallel Comput. | 6 |
| 2016 | Special Issue on Cluster Computing
Michela Taufer, Pavan Balaji, Satoshi Matsuoka |
Parallel Comput. | 1 |
| 2015 | Accurate Scoring of Drug Conformations at the Extreme ScaleabstractWe present a scalable method to extensively search for and accurately select pharmaceutical drug candidates in large spaces of drug conformations computationally generated and stored across the nodes of a large distributed system. For each legend conformation in the dataset, our method first extracts relevant geometrical properties and transforms the properties into a single metadata point in the three-dimensional space. Then, it performs an ochre-based clustering on the metadata to search for predominant clusters. Our method avoids the need to move legend conformations among nodes because it extracts relevant data properties locally and concurrently. By doing so, we can perform accurate and scalable distributed clustering analysis on large distributed datasets. We scale the analysis of our pharmaceutical datasets a factor of 400X higher in performance and 500X larger in size than ever before. We also show that our clustering achieves higher accuracy compared with that of traditional clustering methods and conformational scoring based on minimum energy. Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Pavan Balaji, Michela Taufer |
CCGRID | 5 |
| 2015 | On the Need for Reproducible Numerical Accuracy through Intelligent Runtime Selection of Reduction Algorithms at the Extreme ScaleabstractThe inherent nondeterminism present in reduction operations on an exascale system, coupled with the nonassociativity of floating-point arithmetic, makes achieving reproducible results difficult or impossible. Work investigating the irreproducibility phenomenon has generally proceeded along one of two veins: (1) development of algorithms that produce reproducible numerical results irrespective of nondeterminism in the reduction tree and (2) study of the system-level factors that induce nondeterminism. Our work builds on the latter and unveils the power of mathematical methods to mitigate error propagation at the exascale. We focus on floating-point error accumulation over global summations where enforcing any reduction order is expensive or impossible. We model parallel summations with reduction trees and identify those parameters that can be used to estimate the reduction's sensitivity to variability in the reduction tree. We assess the impact of these parameters on the ability of different reduction methods to successfully mitigate errors. Our results illustrate the pressing need for intelligent runtime selection of reduction operators that ensure a given degree of reproducible accuracy. Dylan Chapp, J. Travis Johnston, Michela Taufer |
CLUSTER | 3 |
| 2015 | Network Quality of Service in Docker ContainersabstractThis poster presents an extension to the currently limited Docker's networks. Specifically, to guarantee quality of service (QoS) on the network, our extension allows users to assign priorities to Docker's containers and configures the network to service these containers based on their assigned priority. Providing QoS not only improves the user experience but also reduces the operation cost by allowing for the efficient use of resources. Our implementation ensures that time-sensitive and critical applications, hosted in high-priority containers, get a greater share of network bandwidth, without starving other containers. Ayush Dusia, Michela Taufer |
CLUSTER | 3 |
| 2015 | A Two-Tiered Approach to I/O Quality of Service in Docker ContainersabstractLinux containers allow applications to run in complete isolation from one another without the extra overhead of running entirely separate operating systems. This approach eliminates memory overheads associated with virtualization and virtual machines and helps businesses run their day-today applications. Unfortunately, multiple applications sharing the same resources can result in substantial resource contention among the applications in the containers and substantial performance loss. One way to mitigate this loss in performance is by ensuring quality of service (QoS) guaranteeing that the application of interest meets the performance requirements. Existing work targets ways of managing CPU, network, and memory contention, however, no solutions exist for managing contention associated with I/O. To address the I/O contention challenge in containers, we propose a two-tiered approach (i.e., at both the cluster and node levels) that extends Docker and Docker Swarm, making both capable of monitoring and controlling the I/O of Dockers containers. We demonstrate how our two-tiered approach has the potential for higher resource utilization without the effects of contention. Sean McDaniel, Stephen Herbein, Michela Taufer |
CLUSTER | 3 |
| 2015 | Dynamic CPU Resource Allocation in Containerized Cloud EnvironmentsabstractIn recent years, lighter-weight virtualization solutions have begun to emerge as an alternative to virtual machines. Because these solutions are still in their infancy, however, several research questions remain open in terms of how to effectively manage computing resources. One important problem is the management of resources in the event of overutilization. For some applications, overutilization can severely affect performance. We provide a solution to this problem by extending the concept of timeslicing to the level of virtualization container. Through this approach we can control and mitigate some of the more detrimental performance effects oversubscription. Our results show significant improvement over standard scheduling with Docker. José Monsalve Diaz, Aaron Myles Landwehr, Michela Taufer |
CLUSTER | 3 |
| 2015 | From HPC Performance to Climate Modeling: Transforming Methods for HPC Predictions into Models of Extreme Climate ConditionsabstractIn the past forty years, the high-performance computing (HPC) community has been developing powerful and rigorous tools for predicting the performance of supercomputers from log traces. In this paper, we transform one of these approaches previously used for predicting idle resources in high-end clusters into a method for capturing extreme climate events in geographical locations of interest. Our method uses an analysis based on empirical cumulative distribution functions (ECDFs) to benchmark and model occurrences of climate events including extreme temperature and precipitation. The method comprises two phases: a learning phase and a prediction phase. The learning phase applies the ECDF-based empirical analysis to historical climate data in order to identify suitable modeling and forecasting windows, both given in years. The prediction phase applies the modeling window to the most recent climate data in order to estimate the likelihood that given portions of the region of interest can experience extreme climate events in the forecasting window. The research is the first of its kind to extend HPC performance modeling techniques to study extreme climate events. Ryan McKinney, Vivek K. Pallipuram, Rodrigo Vargas, Michela Taufer |
e-Science | 4 |
| 2015 | A Genetic Programming Approach to Design Resource Allocation Policies for Heterogeneous Workflows in the CloudabstractWhen dealing with very large applications in the cloud, higher costs do not always result in better turnaround times, particularly for complex workflows with multiple task dependencies. Thus, resource allocation policies are needed that can determine when using expensive but faster resources is best and when it is not. Manually developing such heuristics is time consuming and limited by the subjective beliefs of the developer. To overcome such impediments, we present an automatic method that designs and evaluates a large set of policies using a genetic programming approach. Our method finds a robust set of policies that adapt to changes in workload while using resources efficiently. Our results show that our genetic programming designed policies perform better than greedy and other human designed policies do. Trilce Estrada, Michael R. Wyatt II, Michela Taufer |
ICPADS | 3 |
| 2015 | A Testing Engine for High-Performance and Cost-Effective Workflow Execution in the CloudabstractWhile pursuing high performance and cost effectiveness for directed acyclic graph (DAG)-structured scientific workflow executions in the cloud, it is critical to identify appropriate resource instances and their quantity. This paper presents a testing engine that employs a resource-selection heuristic, which statically analyzes the DAG structure to guide the selection of resource instances, how many and which ones. The testing engine combines the heuristic with two platform-independent DAG-scheduling policies, the Area-oriented DAG-scheduling heuristic (AO) and the Locally-Optimal heuristic (L-OPT), to perform extensive validation assessments. The testing engine ensures the realism of these assessments by modeling the performance variability of the cloud platform using real traces. The testing engine also enables cost-effectiveness analysis that guides users to select a small set of instance candidates that provide performance-cost trade off. Our empirical results show that the pairing of the resource-selection heuristic with AO scheduling policy is a powerful method for cost-effective DAG-structured workflow execution in the cloud. Vivek K. Pallipuram, Trilce Estrada, Michela Taufer |
ICPP | 3 |
| 2014 | Applying frequency analysis techniques to dag-based workflows to benchmark and predict resource behavior on non-dedicated clustersabstractToday, scientific workflows on high-end non-dedicated clusters increasingly resemble directed acyclic graphs (DAGs). The execution trace analysis of the associated DAG-based workflows can provide valuable insights into the system behavior in general, and the occurrences of events like idle times in particular, thereby opening avenues for optimized resource utilization. In this paper, we propose a bipartite tool that uses frequency analysis techniques to benchmark and predict event occurrences in DAG-based workflows; highlighting the system behavior for a given cluster configuration. Using an empirically determined prediction window, the tool parses real-time traces to generate the cumulative distribution function (CDF) of the event occurrences. The CDF is then queried to predict the likelihood of a given number of event instances on the cluster resources in a future time frame. Our results yield average prediction hit-rates as high as 94%. The proposed research enables a runtime system to identify unfavorable event occurrences, thereby allowing for preventive scheduling strategies that maximize system utilization. Vivek K. Pallipuram, Jeffrey DiMarco, Michela Taufer |
CLUSTER | 3 |
| 2014 | Gender and volunteer computing: A survey studyabstractVolunteer computing is a form of citizen science that has a significant gender imbalance. Far fewer women than men participate; women are typically less than ten percent of a project's participants. To better understand the experience of women in volunteer computing and seek clues as to methods for using volunteer computing experience as a recruiting tool, we analyze participant survey data from a project that tries to engage new communities with an interactive infrastructure. Our results showed very few gender differences among the responding men and women in volunteer computing. Our findings add to evidence that men and women engaged in computing activities are overwhelmingly similar. The challenge for gender balance seems to be informing and engaging larger networks of diverse women to promote volunteer computing and contribute to achieving its goals. Feng Raoking, Joanne McGrath Cohoon, Kathryn Cooke, Michela Taufer, Trilce Estrada |
FIE | 4 |
| 2014 | Using surrogate-based modeling to predict optimal I/O parameters of applications at the extreme scaleabstractOn petascale systems, the selection of optimal values for I/O parameters without taking into account the I/O size and pattern can cause the I/O time to dominate the simulation time, compromising the application's scalability. In this paper, we adopt and adapt an engineering method called surrogate-based modeling to efficiently search for the optimal I/O parameter values and accurately predict the associated I/O times at the extreme scale. Our approach allows us to address both the search and prediction in a short time, even when the application's I/O is large and exhibits irregular patterns. Michael Matheny, Stephen Herbein, Norbert Podhorszki, Scott Klasky, Michela Taufer |
ICPADS | 5 |
| 2014 | Enabling In-Situ Data Analysis for Large Protein-Folding Trajectory DatasetsabstractThis paper presents a one-pass, distributed method that enables in-situ data analysis for large protein folding trajectory datasets by executing sufficiently fast, avoiding moving trajectory data, and limiting the memory usage. First, the method extracts the geometric shape features of each protein conformation in parallel. Then, it classifies sets of consecutive conformations into meta-stable and transition stages using a probabilistic hierarchical clustering method. Lastly, it rebuilds the global knowledge necessary for the intraand inter-trajectory analysis through a reduction operation. The comparison of our method with a traditional approach for a villin headpiece sub domain shows that our method generates significant improvements in execution time, memory usage, and data movement. Specifically, to analyze the same trajectory consisting of 20,000 protein conformations, our method runs in 41.5 seconds while the traditional approach takes approximately 3 hours, uses 6.9MB memory per core while the traditional method uses 16GB on one single node where the analysis is performed, and communicates only 4.4KB while the traditional method moves the entire dataset of 539MB. The overall results in this paper support our claim that our method is suitable for in-situ data analysis of folding trajectories. Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Michela Taufer |
IPDPS | 4 |
| 2014 | Bandwidth Modeling in Large Distributed Systems for Big Data ApplicationsabstractThe emergence of Big Data applications provides new challenges in data management such as processing and movement of masses of data. Volunteer computing has proven itself as a distributed paradigm that can fully support Big Data generation. This paradigm uses a large number of heterogeneous and unreliable Internet-connected hosts to provide Peta-scale computing power for scientific projects. With the increase in data size and number of devices that can potentially join a volunteer computing project, the host bandwidth can become a main hindrance to the analysis of the data generated by these projects, especially if the analysis is a concurrent approach based on either in-situ or in-transit processing. In this paper, we propose a bandwidth model for volunteer computing projects based on the real trace data taken from the Docking@Home project with more than 280,000 hosts over a 5-year period. We validate the proposed statistical model using model-based and simulation-based techniques. Our modeling provides us with valuable insights on the concurrent integration of data generation with in-situ and in-transit analysis in the volunteer computing paradigm. Bahman Javadi, Boyu Zhang 0002, Michela Taufer |
PDCAT | 3 |
| 2013 | Benchmarking Gender Differences in Volunteer Computing ProjectsabstractVolunteer Computing (VC) uses the computational resources of volunteers with Internet-connected personal computers to address fundamental problems in science. Docking Home (D@H) is a VC project targeting drug discovery through high throughput docking simulations i.e., by docking small molecules (ligands) into target proteins associated to diseases. Currently there are more than 27,000 volunteers (and 70,000 computers) worldwide supporting D@H. Similar to national trends in STEM fields, in general, the huge majority of volunteers engaged in VC projects, and in D@H in particular, are Caucasian males. This paper aims to characterize the current VC community supporting D@H and uses the information to define strategies that can help attract and retain female and ethnic minority volunteers. Trilce Estrada, Kathleen L. Pusecker, Manuel R. Torres, Joanne McGrath Cohoon, Michela Taufer |
e-Science | 5 |
| 2013 | Efficient SDS Simulations on Multi-GPU Nodes of XSEDE High-End ClustersabstractEfficiently studying Sodium Dodecyl Sulfate (SDS) molecules' formations in the presence of different molar concentrations on high-end GPU clusters whose nodes share accelerators exposes us to several challenges, including the need to dynamically adapt the job lengths. Neither virtualization nor lightweight OS solutions can easily support generality, portability, and maintainability in concert. Our solution complements rather than rewrites existing workflow and resource managers with a companion module that complements functions of the workflow manager and a wrapper module that extends functions of the resource managers. Results on the Keene land cluster show how, by using our modules, accelerated SDS simulations more efficiently use the cluster's GPUs while leading to relevant scientific observations. Samuel Schlachter, Stephen Herbein, Michela Taufer, Shuching Ou, Sandeep Patel, Jeremy Logan |
e-Science | 3 |
| 2013 | On the powerful use of simulations in the Quake-Catcher Network to efficiently position low-cost earthquake sensors
Kyle E. Benson, Samuel Schlachter, Trilce Estrada, Michela Taufer, Jesse Lawrence, Elizabeth Cochran |
Future Gener. Comput. Syst. | 4 |
| 2012 | ExSciTecH: Expanding volunteer computing to Explore Science, Technology, and HealthabstractThis paper presents ExSciTecH, an NSF-funded project deploying volunteer computing (VC) systems to Explore Science, Tecenology, and Health. ExSciTecH aims at radically transforming VC systems and the volunteer's experience. To pursue this goal, ExSciTecH integrates and uses gameplay environments into BOINC, a well-known VC middleware, to involve the volunteers not only for simply donating idle cycles but also for actively participating in scientific discovery, i.e., generating new simulations side by side with the scientists. More specifically, ExSciTecH plugs into the BOINC framework extending it with two main gaming components, i.e., a learning component that includes a suite of games for training users on relevant biochemical concepts, and an engaging component that includes a suite of games to engage volunteers in drug design and scientific discovery. We assessed the impact of a first implementation of the learning game on a group of students at the University of Delaware. Our tests clearly show how ExSciTecH can generate higher levels of enthusiasm than more traditional learning tools in our students. Michael Matheny, Samuel Schlachter, L. M. Crouse, E. T. Kimmel, Trilce Estrada, Marcel Schumann, Roger S. Armen, Gary M. Zoppetti, Michela Taufer |
eScience | 9 |
| 2012 | On the effectiveness of application-aware self-management for scientific discovery in volunteer computing systemsabstractAn important challenge faced by high-throughput, multiscale applications is that human intervention has a central role in driving their success. However, manual intervention is inefficient, error-prone and promotes resource wasting. This paper presents an application-aware modular framework that provides self-management for computational multiscale applications in volunteer computing (VC). Our framework consists of a learning engine and three modules that can be easily adapted to different distributed systems. The learning engine of this framework is based on our novel tree-like structure called KOTree. KOTree is a fully automatic method that organizes statistical information in a multidimensional structure that can be efficiently searched and updated at runtime. Our empirical evaluation shows that our framework can effectively provide application-aware self-management in VC systems. Additionally, we observed that our KOTree algorithm is able to predict accurately the expected length of new jobs, resulting in an average of 85% increased throughput with respect to other algorithms. Trilce Estrada, Michela Taufer |
SC | 2 |
| 2011 | On the Powerful Use of Simulations in the Quake-Catcher Network to Efficiently Position Low-cost Earthquake SensorsabstractThe Quake-Catcher Network (QCN) uses low-cost sensors connected to volunteer computers across the world to monitor seismic events. The location and density of these sensors' placement can impact the accuracy of the event detection. Because testing different special arrangements of new sensors could disrupt the currently active project, this would best be accomplished in a simulated environment. This paper presents an accurate and efficient framework for simulating the low cost QCN sensors and identifying their most effective locations and densities. Results presented show how our simulations are reliable tools to study diverse scenarios under different geographical and infrastructural constraints. Kyle E. Benson, Trilce Estrada, Michela Taufer, Jesse Lawrence, Elizabeth Cochran |
eScience | 3 |
| 2011 | Providing Quality of Science in Volunteer ComputingabstractWe propose an application-centric approach for tuning, at runtime, parameters in volunteer computing (VC). Our approach defines Quality of Science (QSc) as the ability of a system to meet application-specific goals within time constraints. We first estimate the impact of different application parameters on both QSc metrics and resource usage from empirical information gathered at runtime and then reconfigure the parameters accordingly. Our approach results in more accurate solutions in shorter times and lower resource consumption than other, more traditional approaches. Trilce Estrada, Michela Taufer |
HPCC | 2 |
| 2011 | Special issue of computer communications on information and future communication security
Jong Hyuk Park 0001, Sheikh Iqbal Ahamed, Willy Susilo, Michela Taufer |
Comput. Commun. | 4 |
| 2010 | Improving numerical reproducibility and stability in large-scale numerical simulations on GPUsabstractThe advent of general purpose graphics processing units (GPGPU's) brings about a whole new platform for running numerically intensive applications at high speeds. Their multi-core architectures enable large degrees of parallelism via a massively multi-threaded environment. Molecular dynamics (MD) simulations are particularly well-suited for GPU's because their computations are easily parallelizable. Significant performance improvements are observed when single precision floating-point arithmetic is used. However, this performance comes at the cost of accuracy: it is widely acknowledged that constant-energy (NVE) MD simulations accumulate errors as the simulation proceeds due to the inherent errors associated with integrators used for propagating the coordinates. A consequence of this numerical integration is the drift of potential energy as the simulation proceeds. Double precision arithmetic partially corrects this drifting, but is significantly slower than single precision, comparable to CPU performance. To address this problem, we extend the approaches of previous literature to improve numerical reproducibility and stability in MD simulations, while assuring efficiency and performance comparable to that when using the GPU hardware implementation of single precision arithmetic. We present development of a library of mathematical functions that use fast and efficient algorithms to fix the error produced by the equivalent operations performed by GPU. We successfully validate the library with a suite of synthetic codes emulating the MD behavior on GPUs. Michela Taufer, Omar Padron, Philip Saponaro, Sandeep Patel |
IPDPS | 1 |
| 2010 | Parallelization of tau-leap coarse-grained Monte Carlo simulations on GPUsabstractThe Coarse-Grained Monte Carlo (CGMC) method is a multi-scale stochastic mathematical and simulation framework for spatially distributed systems. CGMC simulations are important tools for studying phenomena such as catalysis, crystal growth, surface diffusion, phase transitions on single crystals, and cell membrane receptor dynamics. In parallel CGMC, the tau-leap method is used for parallel simulations that are executed on traditional CPU clusters in a master-slave setting. Unfortunately the communications between master and slaves negatively impact speedup and scalability. In this paper, we explore the potentials of GPUs for the tau-leap method and we present an extensive performance evaluation that leads to the most suitable degree of parallelism for this method under different simulation profiles. We show how the efficient parallelization of the tau-leap method for GPUs includes (1) the redefinition of its data structures, (2) the redesign of its algorithm, and (3) the selection of the most appropriate degree of parallelism (i.e., fine-grained or course-gained) on a single GPU or multiple GPUs. Exceptional performance improvements can thus be achieved for this method. Lifan Xu, Michela Taufer, Stuart Collins, Dionisios G. Vlachos |
IPDPS | 2 |
| 2009 | Modeling Job Lifespan Delays in Volunteer Computing ProjectsabstractVolunteer computing (VC) projects harness the power of computers owned by volunteers across the Internet to perform hundreds of thousands of independent jobs. In VC projects, the path leading from the generation of jobs to the validation of the job results is characterized by delays hidden in the job lifespan, i.e., distribution delay, in-progress delay, and validation delay. These delays are difficult to estimate because of the dynamic behavior and heterogeneity of VC resources. A wrong estimation of these delays can cause the loss of project throughput and job latency in VC projects. In this paper, we evaluate the accuracy of several probabilistic methods to model the upper time bounds of these delays. We show how our selected models predict up-and-down trends in traces from existing VC projects. The use of our models provides valuable insights on selecting project deadlines and taking scheduling decisions. By accurately predicting job lifespan delays, our models lead to more efficient resource use, higher project throughput, and lower job latency in VC projects. Trilce Estrada, Michela Taufer, Kevin Reed |
CCGRID | 2 |
| 2009 | MNEMONIC: A Network Environment for Automatic Optimization and Tuning of Data Movement over Advanced NetworksabstractComputational science is considered the third pillar of science but while we continue to make significant strides in high-performance computing, the ability to transfer the vast amounts of data generated and consumed by scientists has fallen behind. This divergence has significantly constrained scientific progress. This paper describes MNEMONIC, an end-to-end cyber-infrastructure system designed to improve scientists' productivity by automatically and transparently optimizing and tuning data movement over the network. MNEMONIC integrates advanced technologies, i.e., network interfaces, protocols, and measurement toolkits, to seamlessly manage high-performance data movement. We have integrated the Java implementation of GridFTP in MNEMONIC to provide a cross-platform data transfer capability. Preliminary results show that MNEMONIC+GridFTP outperforms the TCP+GridFTP implementation by over a factor of 10x, while remaining easy to install and operate. Patrick McClory, Ezra Kissel, D. Martin Swany, Michela Taufer |
ICCCN | 4 |
| 2009 | EmBOINC: An emulator for performance analysis of BOINC projectsabstractBOINC is a platform for volunteer computing. The server component of BOINC embodies a number of scheduling policies and parameters that have a large impact on the projects throughput and other performance metrics. We have developed a system, EmBOINC, for studying these policies and parameters. EmBOINC uses a hybrid approach: it simulates a population of volunteered clients (including heterogeneity, churn, availability, reliability) and it emulates the server component; that is, it uses the actual server software and its associated database. This paper describes the design of EmBOINC and validates its results based on trace data from an existing BOINC project. Trilce Estrada, Michela Taufer, Kevin Reed, David P. Anderson |
IPDPS | 2 |
| 2009 | Performance Prediction and Analysis of BOINC Projects: An Empirical Study with EmBOINCabstractMiddleware systems for volunteer computing convert a set of computers that is large and diverse (in terms of hardware, software, availability, reliability, and trustworthiness) into a unified computing resource. This involves a number of scheduling policies and parameters, which have a large impact on the throughput and other performance metrics. How can we study and refine these policies? Experimentation in the context of a working project is problematic, and it is difficult to accurately model complex middleware in a conventional simulator. Instead, we use an approach in which the policies being studied are “emulated”, using parts of the actual middleware. In this paper we describe EmBOINC, an emulator based on the BOINC middleware system. EmBOINC simulates a population of volunteered clients (including heterogeneity, churn, availability, and reliability) and emulates the BOINC server components. After describing the design of EmBOINC and its validation, we present three case studies in which the impact of different scheduling policies are quantified in terms of throughput, latency, and starvation metrics. Trilce Estrada, Michela Taufer, David P. Anderson |
J. Grid Comput. | 2 |
| 2008 | On the Effectiveness of Rebuilding RNA Secondary Structures from Sequence ChunksabstractDespite the computing power of emerging technologies, predicting long RNA secondary structures with thermodynamics-based methods is still infeasible, especially if the structures include complex motifs such as pseudoknots. This paper presents preliminary results on rebuilding RNA secondary structures by an extensive and systematic sampling of nucleotide chunks. The rebuilding approach merges the significant motifs found in the secondary structures of the single chunks. The extensive sampling and prediction of nucleotide chunks are supported by grid technology as part of the RNAVLab functionality. Significant motifs are identified in the chunk secondary structures and merged in a single structure based on their recurrences and other statistical insights. A critical analysis of the strengths, weaknesses, and future developments of our method is presented. Michela Taufer, Thamar Solorio, Abel Licon, David Mireles, Ming-Ying Leung |
IPDPS | 1 |
| 2008 | RNAVLab: A virtual laboratory for studying RNA secondary structures based on grid computing technology
Michela Taufer, Ming-Ying Leung, Thamar Solorio, Abel Licon, David Mireles, Roberto Araiza, Kyle L. Johnson |
Parallel Comput. | 1 |
| 2007 | Moving Volunteer Computing towards Knowledge-Constructed, Dynamically-Adaptive Modeling and SchedulingabstractVolunteer computing projects supported by BOINC have been exploring new research directions. For example, mature projects like Folding@home are moving towards the use of a broader range of architectures and computers. Other projects such as Docking@Home are exploring multi-scale, resource-driven and application-driven adaptations of the volunteer system. This paper presents results that enforce the need for knowledge-constructed capabilities in volunteer computing projects, i.e., the capability to drive simulations based on application-results and resource-status. The Docking@Home project, which uses volunteer resources to study putative drugs by computationally simulating the behavior of small molecules (ligands) when docking to a protein, serves as a case study to positively assess two key hypotheses. The first hypothesis claims that the adaptive selection of computational models for docking simulations based on the features of the protein and ligand can positively affect the final accuracy of the prediction. The second hypothesis claims that the adaptive selection of volunteer resources can ultimately improve project throughput. Michela Taufer, Andre Kerstens, Trilce Estrada, David A. Flores, Richard Zamudio, Patricia J. Teller, Roger S. Armen, Charles L. Brooks III |
IPDPS | 1 |
| 2007 | RNAVLab: A unified environment for computational RNA structure analysis based on grid computing technologyabstractRibonucleic acid (RNA) molecules play important roles in many biological processes including gene expression and regulation. An RNA molecule is a linear polymer which folds back on itself to form a three dimensional (3D) functional structure. In this paper we briefly address mathematical problems associated with the grid computing approach to RNA structure prediction. In particular, we introduce models to partition a large RNA molecule into smaller segments to be assigned to different computers on the grid. Based on these models, we formulate a sampling strategy to select RNA segments for computational prediction to maximize prediction consistency. This strategy is under construction as part of RNAVLab, our unified environment for computational RNA structure analysis, i.e., prediction, alignment, comparison, and classification. A first prototype of RNAVLab is presented and used to investigate the possible association of secondary structure types with RNA functions by analyzing secondary structures for a family of nodavirus genomes. Michela Taufer, Ming-Ying Leung, Kyle L. Johnson, Abel Licon |
IPDPS | 1 |
| 2007 | Topaz: Extending Firefox to Accommodate the GridFTP ProtocolabstractAs grid infrastructures mature, an increasing challenge is to provide end-user scientists with intuitive interfaces to computational services, data management capabilities, and visualization tools. One novel approach, being successfully applied in the domain of computational chemistry, is to leverage the capabilities of the Mozilla framework to provide rich end-user tools that seamlessly integrate with remote resources such as Web/grid services and data repositories. The Mozilla framework provides much of the infrastructure to build rich end-user applications, but lacks the capability to integrate with grid protocols and APIs. In this paper we present the design and evaluation of Topaz, a Mozilla-based component that provides GridFTP functionality to the popular Firefox browser. Topaz provides end-user scientists with a familiar and user-friendly interface with which to access arbitrary GridFTP servers by providing upload and download functionalities as well as by obtaining and managing users' grid certificates. Richard Zamudio, Daniel Catarino, Michela Taufer, Brent Stearn, Karan Bhatia |
IPDPS | 3 |
| 2006 | The Effectiveness of Threshold-Based Scheduling Policies in BOINC ProjectsabstractSeveral scientific projects use BOINC (Berkeley Open Infrastructure for Network Computing) to perform largescale simulations using volunteers' computers (workers) across the Internet. In general, the scheduling of tasks in BOINC uses a First-Come-First-Serve policy and no attention is paid to workers' past performance, such as whether or not they have tended to perform tasks promptly and correctly. In this paper we use SimBA, a discrete-event Simulator of BOINC Applications, to study new threshold-based scheduling strategies for BOINC projects that use availability and reliability metrics to classify workers and distribute tasks according to this classification. We show that if availability and reliability thresholds are selected properly, then the workers' throughput of valid results increases significantly in BOINC projects. Trilce Estrada, David A. Flores, Michela Taufer, Patricia J. Teller, Andre Kerstens, David P. Anderson |
e-Science | 3 |
| 2006 | A systematic multi-step methodology for performance analysis of communication traces of distributed applications based on hierarchical clusteringabstractOften parallel scientific applications are instrumented and traces are collected and analyzed to identify processes with performance problems or operations that cause delays in program execution. The execution of instrumented codes may generate large amounts of performance data, and the collection, storage, and analysis of such traces are time and space demanding. To address this problem, this paper presents an efficient, systematic, multi-step methodology, based on hierarchical clustering, for analysis of communication traces of parallel scientific applications. The methodology is used to discover potential communication performance problems of three applications: TRACE, REMO, and SWEEP3D. Maria Gabriela Aguilera, Patricia J. Teller, Michela Taufer, Felix Wolf 0001 |
IPDPS | 3 |
| 2006 | Poster reception - SimBA: a discrete event simulator for performance prediction of volunteer computing projectsabstractSimBA (Simulator of BOINC Applications) is a discrete event simulator that accurately models the main functions of BOINC, a master-worker runtime framework. SimBA generates, distributes, and monitors tasks executed in a highly volatile, heterogeneous, and distributed environment such as Volunteer Computing (VC). In addition, it collects and validates results of executed tasks.To understand the strengths and weaknesses of BOINC under distinct scenarios, project designers must study and quantify its performance without affecting the VC community. Although this is not possible on production systems, it is possible using SimBA, which is capable of testing a wide range of hypotheses in a short period of time.Our experience to date indicates that SimBA is a reliable tool for performance prediction of VC projects. Preliminary results show that SimBA's predictions of [email protected] performance are within approximately 5% of the performance reported by this BOINC project. David A. Flores, Trilce Estrada, Michela Taufer, Patricia J. Teller, Andre Kerstens |
SC | 3 |
| 2006 | Predictor@Home: A "Protein Structure Prediction Supercomputer' Based on Global ComputingabstractPredicting the structure of a protein from its amino acid sequence is a complex process, the understanding of which could be used to gain new insight into the nature of protein functions or provide targets for structure-based design of drugs to treat new and existing diseases. While protein structures can be accurately modeled using computational methods based on all-atom physics-based force fields including implicit solvation, these methods require extensive sampling of native-like protein conformations for successful prediction and, consequently, they are often limited by inadequate computing power. To address this problem, we developed Predictor@Home, a "structure prediction supercomputer” powered by the Berkeley Open Infrastructure for Network Computing (BOINC) framework and based on the global computing paradigm (i.e., volunteered computing resources interconnected to the Internet and owned by the public). In this paper, we describe the protocol we employed for protein structure prediction and its integration into a global computing architecture based on public resources. We show how Predictor@Home significantly improved our ability to predict protein structures by increasing our sampling capacity by one to two orders of magnitude. Michela Taufer, Chahm An, Andreas Kerstens, Charles L. Brooks III |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | Metrics for Effective Resource Management in Global Computing EnvironmentsabstractGlobal computing uses Internet-connected PCs volunteered by their owners. These PCs are diverse, volatile, and error-prone. Sophisticated scheduling methods commonly applied in grid computing may not be sufficiently scalable and flexible for global computing environments. This paper shows that it is possible to classify global computing hosts based on simple metrics such as availability and reliability, and that it is efficient to assign tasks to such hosts accordingly. The proposed classification of workers is applied to P@H, a global computing project for protein structure prediction Michela Taufer, Patricia J. Teller, David P. Anderson, Charles L. Brooks III |
e-Science | 1 |
| 2005 | Study of a highly accurate and fast protein-ligand docking method based on molecular dynamicsabstractAbstract Few methods use molecular dynamics simulations in concert with atomically detailed force fields to perform protein–ligand docking calculations because they are considered too time demanding, despite their accuracy. In this paper we present a docking algorithm based on molecular dynamics which has a highly flexible computational granularity. We compare the accuracy and the time required with well‐known, commonly used docking methods such as AutoDock, DOCK, FlexX, ICM, and GOLD. We show that our algorithm is accurate, fast and, because of its flexibility, applicable even to loosely coupled distributed systems such as desktop Grids for docking. Copyright © 2005 John Wiley & Sons, Ltd. Michela Taufer, Michael F. Crowley, Daniel J. Price, Andrew A. Chien, Charles L. Brooks III |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | The Computational Chemistry Prototyping EnvironmentabstractEvolving technologies, as exemplified by computational grids and Web services, have made it possible to solve new scientific problems that would not have been feasible previously. In order to make such advances available to the community in general and to be able to solve new problems, not necessarily from the same discipline, it is imperative to build tools that provide a common user interface in order that application programmers and users do not have to be concerned with particulars of Web services and their underlying code, computational platforms, or with data file formats. We will describe our efforts in creating a computational chemistry environment that encompasses a general scientific workflow environment, a domain specific example for quantum chemistry, our ongoing design of a workflow user interface, and our efforts at database integration. Kim K. Baldridge, Jerry P. Greenberg, Wibke Sudholt, Stephen Mock, Ilkay Altintas, Céline Amoreira, Yohann Potier, Adam Birnbaum, Karan Bhatia, Michela Taufer |
Proc. IEEE | 10 |
| 2005 | DGMonitor: A Performance Monitoring Tool for Sandbox-Based Desktop Grid Platforms
Pietro Cicotti, Michela Taufer, Andrew A. Chien |
J. Supercomput. | 2 |
| 2004 | DGMonitor: A Performance Monitoring Tool for Sandbox-Based Desktop Grid PlatformsabstractSummary form only given. Accurate, continuous resource monitoring and profiling are critical for enabling performance tuning and scheduling optimization. In desktop grid systems that employ sandboxing, these issues are challenging because (1) subjobs inside sandboxes are executed in a virtual computing environment and (2) the state of the virtual computing environment within the sandboxes is reset to empty after each subjob completes. DGMonitor is a monitoring tool, which builds a global, accurate, and continuous view of real resource utilization for desktop grids with sandboxing. Our monitoring tool measures performance unobtrusively and reliably, uses a simple performance data model, and is easy to use. Our measurements demonstrate that DGMonitor can scale to large desktop grids (up to 12000 workers) with low monitoring overhead in terms of resource consumption (less than 0.1%) on desktop PCs. Though we developed DGMonitor with the Entropia DCGrid platform, our tool is easily integrated into other desktop grid systems. In all of these systems, DGMonitor data can support existing and novel information services, particularly for performance tuning and scheduling. Pietro Cicotti, Michela Taufer, Andrew A. Chien |
IPDPS | 2 |
| 2004 | Characterizing and Evaluating Desktop Grids: An Empirical StudyabstractSummary form only given. Desktop resources are attractive for running compute-intensive distributed applications. Several systems that aggregate these resources in desktop grids have been developed. While these systems have been successfully used for many high throughput applications there has been little insight into the detailed temporal structure of CPU availability of desktop grid resources. Yet, this structure is critical to characterize the utility of desktop grid platforms for both task parallel and even data parallel applications. We address the following questions: (i) What are the temporal characteristics of desktop CPU availability in an enterprise setting? (ii) How do these characteristics affect the utility of desktop grids? (iii) Based on these characteristics, can we construct a model of server "equivalents" for the desktop grids, which can be used to predict application performance? We present measurements of an enterprise desktop grid with over 220 hosts running the Entropia commercial desktop grid software. We utilize these measurements to characterize CPU availability and develop a performance model for desktop grid applications for various task granularities, showing that there is an optimal task size. We then use a cluster equivalence metric to quantify the utility of the desktop grid relative to that of a dedicated cluster. Derrick Kondo, Michela Taufer, Charles L. Brooks III, Henri Casanova, Andrew A. Chien |
IPDPS | 2 |
| 2004 | Study of a Highly Accurate and Fast Protein-Ligand Docking Algorithm Based on Molecular DynamicsabstractSummary form only given. Few methods use molecular dynamics simulations based on atomically detailed force fields to study the protein-ligand docking process because they are considered too time demanding despite their accuracy. We present a docking algorithm based on molecular dynamics simulations which has a highly flexible computational granularity. We compare the accuracy and the time required with well-known, commonly used docking methods like AutoDock, DOCK, FlexX, ICM, and GOLD. We show that our algorithm is accurate, fast and, because of its flexibility, applicable even to loosely coupled distributed systems like desktop grids for docking. Michela Taufer, Michael F. Crowley, Daniel J. Price, Andrew A. Chien, Charles L. Brooks III |
IPDPS | 1 |
| 2003 | Combining Task- and Data Parallelism to Speed up Protein Folding on a Desktop Grid PlatformabstractThe steady increase of computing power at lower and lower cost enables molecular dynamics simulations to investigate the process of protein folding with an explicit treatment of water molecules. Such simulations are typically done with well known computational chemistry codes like CHARMM. Desktop grids such as the United Devices MetaProcessor are highly attractive platforms, since scavenging for unused machines on Intra- and Internet delivers compute power that is almost free. However, the predominant programming paradigm for current desktop grids is pure task parallelism and might not fit the needs for protein folding simulations with explicit water molecules. A short overall turn-around time of a simulation remains highly important for research productivity, but the need for an accurate model and long simulation time-scales leads to tasks that are too large for optimal scheduling on a desktop grid. To address this problem, we introduce a combination of task- and data parallelism as a well suitable computing paradigm for protein folding investigations on grid platforms. As a proof of concept, we design and implement a simple system for protein folding simulations based on the notion of combined task and data parallelism with clustered workers. Clustered workers are machines grouped into small clusters according to network and CPU performance criteria and act as super-nodes within a desktop grid, permitting the utilization of data parallelism in addition to the task parallelism. We integrate our new paradigm into the existing software environment of the United Devices MetaProcessor. For a test protein, we reach a better quality of the folding calculations than we reached using just task parallelism on distributed systems. Bennet Uk, Michela Taufer, Thomas Stricker, Giovanni Settanni, Andrea Cavalli, Amedeo Caflisch |
CCGRID | 2 |
| 2003 | A Performance Monitor Based on Virtual Global Time for Clusters of PCsabstractDebugging the performance of parallel and distributed systems remains a difficult task despite the widespread use of middleware packages for automatic distribution, communication and tasking in clusters. In this paper we present a performance monitoring tool for clusters of PCs that is based on the simple concept of accounting for resource usage and on the simple idea of mapping all performance related state of hardware performance counters and operating system variables backwards to the application level. In this way a monitoring tool can explain the most relevant performance metrics at a higher level that is easily understood by the application developer. The most important metric for distributed high performance applications remains the total execution time vs. the number of compute nodes involved, since it translates into the scalability of an application. As a detailed contribution of this paper, we closely look into what is needed to reverse map the low level performance counters at each node back through the middleware layer responsible for the parallelization and distribution. The specific problems encountered and dealt with are the creation of a flexible notion of global time for time-stamping and the reassembling of performance data and an appropriate communication mechanism to minimize monitoring intrusion due to the additional networking traffic caused by the monitor. We show how our tool can be used to measure, explain and predict the performance and scalability of a distributed OLAP application running on clusters of PCs. Michela Taufer, Thomas Stricker |
CLUSTER | 1 |
| 2002 | On the Migration of the Scientific Code Dyana from SMPs to Clusters of PCs and on to the GridabstractDyana is a molecular biology code used in the study of infectious prion proteins. Like many other scientific codes, Dyana was migrated successfully from vector supercomputers to some more cost-effective cluster of commodity PCs. A further migration to a widely distributed grid computing platform looks very tempting because many of these platforms promise the use of nearly free compute-cycles on the Internet. Not all codes are equally suited for all platforms. Even embarrassingly parallel codes might require a significant re-engineering effort for a migration from one platform to another. A better understanding of the performance characteristics of a code is required before a migration is attempted. To address this problem, we present a systematic method to study the viability of a code migration from one platform to another, before it is actually undertaken. We construct an analytic performance model of the application. We use the previous migration from SMPs to commodity clusters of PCs to validate and calibrate the model. Finally, we extrapolate the performance of Dyana widely distributed computing on the grid and we suggest optimizations in the process of migration. Our general model predicts that Dyana can efficiently use up to 42000 processors with its current workload and is therefore well suited for grid computing on the Internet. Michela Taufer, Thomas Stricker, Gerard Roos, Peter Güntert |
CCGRID | 1 |
| 2002 | Scalability and resource usage of an OLAP benchmark on clusters of PCsabstractDesigning clusters of PCs for distributed databases processing OLAP (On Line Analytical Processing) workloads in parallel with good scalability remains a particular challenge as we are lacking a deep understanding of the architectural issues around resource usage by standard DBMSs on distributed platforms. To address this problem, we present a novel performance monitoring framework for filtering and abstracting samples of performance data from low level counters into a high level performance picture. Our framework is used side by side with the DBMS and delivers many interesting insights about the most critical resource in the different queries and systems configuration. As required for a larger distributed hardware/software system, our solution comprises software instrumentation at the OS level, tools for gathering performance relevant data and an analytical model for performance evaluation and performance prediction to future platforms. We demonstrate the viability of our approach with the in-depth analysis of distributed TPC-D, a standard OLAP benchmark running on clusters of commodity PCs. Based on the data provided by our framework, we isolate and resolve a few crucial performance issues of OLAP workloads on clusters. For different queries, we give a workload characterization in terms of resource usage, quantify the optimal scalability and investigate the impact of the networking speed on the overall application performance. We show that the disk performance and CPU speed remains the most critical resource bottleneck for most queries. Queries with a lot of inter-node communication are limited by the communication software inefficiency within the DBMS and not by the raw networking speeds. A systematic performance evaluation constitutes a solid basis for architectural decisions and system optimization in clusters of PCs that are dedicated to large parallel database systems. Michela Taufer, Thomas Stricker, Roger Weber |
SPAA | 1 |
| 1998 | Accurate Performance Evaluation, Modelling and Prediction of a Message Passing Simulation Code based on MiddlewareabstractIn distributed and vectorized computing there is a large number of highly different supercomputing platforms an application could run on. Therefore most traditional parallel codes are ill equipped to collect data about their resource usage or their behavior at run time and the corresponding data are rarely published and few scientists attack the planning of an application and its platform systematically. As an improvement over the current state of the art, we propose an integrated approach to performance evaluation, modeling and prediction for different platforms. Our approach uses a combination of analytical modeling and systematically designed experimentation with full application runs, reduced application kernels and some benchmarks. We studied our methodology of performance assessment with Opal, an example code in molecular biology, developed at our institution to run on our four Cray J90 ``Classic'' Vector SMPs. Besides a detailed assessment of performance achieved on the J90s, the primary goal of our study was to find the most suitable and most cost effective hardware platform for the application, in particular to check the suitability of this application for slow CoPs, SMP CoPs and fast CoPs, three flavors of Clusters of PCs built with off-the-shelf Intel Pentium processors. A performance assessment based on our model is much easier than porting and parallelizing the application for a new target machine and so we could easily obtain and include performance estimates for a T3E-900, a high end MPP system. The predicted execution times and speedup figures indicate that a well designed cluster of PCs achieves similar if not better performance than the J90 vector processors currently used and that the computational efficiency compares favorably to the T3E-900 for that particular application code. Michela Taufer, Thomas Stricker |
SC | 1 |