Michael E. Papka

dblp:21/5623 · DBLP profile ↗
← Back
89ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0002-6418-5767ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 54 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 since 2021Software engineering, systems software and programming languages · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 SmartCap: Coordinated CPU-GPU Power Capping for Performance-Assurance Energy Efficiency
abstract
Performance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling methods often rely on exhaustive offline profiling, which becomes impractical due to the large setting space and the high cost of profiling large-scale applications. In this paper, we present OPEN, a framework consists of offline and online phases. The offline phase involves building a performance predictor and constructing an initial dense matrix. In the online phase, OPEN performs lightweight online profiling, and leverages the performance predictor with collaborative filtering to make performance prediction. We evaluate OPEN on multiple heterogeneous systems, including those equipped with A100 and A30 GPUs. Results show that OPEN achieves prediction accuracy up to 98.29\%. This demonstrates that OPEN effectively reduces profiling cost while maintaining high accuracy, making it practical for power-aware performance modeling in modern HPC environments. Overall, OPEN provides a lightweight solution for performance prediction under power constraints, enabling better runtime decisions in power-aware computing environments.
Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan
ICS4
2026 Beyond Throughput: Performance and Energy Insights of LLM Inference Across AI Accelerators
Giacomo Brunetta, Varuni Sastry 0001, Xingfu Wu, Valerie Taylor 0001, Michael E. Papka, Zhiling Lan
IPDPS5
2025 Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf Benchmarks
abstract
The significantly increasing use of artificial intelligence (AI) has led to the availability of specialized AI accelerators, aiming to enhance the performance and energy efficiency of AI workloads. In this paper, we conduct an initial study to evaluate the energy requirements of four AI accelerators: Nvidia A100 GPUs, Intel Habana Gaudi Processing Units (HPUs), Graphcore Bow-Pod64 Intelligence Processing Units (IPUs), and GroqRack Language Processing Units (LPUs) using two popular MLPerf benchmarks: BERT-Large and ResNet50. We report the energy requirements for the two benchmarks to achieve a common MLPerfspecified target accuracy. The benchmarks and AI accelerators were chosen based on the following criteria: publicly available tools or libraries from the vendors to monitor power consumption, publicly available optimized models from the vendors, and access to the AI accelerators. Our experimental results indicate that for ResNet50, Intel Gaudi2 HPUs delivered the highest throughput, the lowest energy consumption for both training and inference, and the highest inference energy efficiency, while Graphcore demonstrated the highest training energy efficiency. For BERTLarge pre-training, Intel Gaudi2 outperformed both Nvidia A100 and Graphcore in terms of time, energy consumption, and training energy efficiency. However, for BERT-Large inference, Nvidia A100 achieved the shortest time and lowest energy consumption; Graphcore exhibited the highest throughput in both pre-training and inference, along with the highest inference energy efficiency. We discuss our observations and findings while exploring the associated tradeoffs.
Farah Ferdaus, Xingfu Wu, Valerie Taylor 0001, Zhiling Lan, Sanjif Shanmugavelu, Venkatram Vishwanath, Michael E. Papka
CCGrid7
2025 Interactive Exploration of HACC Cosmology Data using WebXR
abstract
This project introduces an interactive, web-based 3D data viewer specifically designed for the analysis of large-scale scientific cosmological datasets. The OpenCosmo Compute Portal [1], developed by Argonne National Laboratory’s Cosmological Physics and Advanced Computing group, provides easy access to cosmological simulations. Adding advanced visualization capabilities, such as those offered by this viewer, would significantly enhance the portal’s utility, allowing a broader audience to gain deeper insights into queried datasets without needing to build their own visualization tools.We address the critical need to democratize access to advanced scientific visualization, particularly for researchers who aren’t visualization experts. Leveraging modern web technologies, our viewer provides a fluid and responsive environment for exploring complex point cloud data. Key features include customizable visual properties, interactive selection, and robust state management, all accessible directly within a web browser, or virtual reality headset, if available.
Idunnuoluwa A. Adeniji, Joseph A. Insley, Mengjiao Han, Janet Knowles, Michael E. Papka, Victor A. Mateevitsi, Silvio Rizzi 0001
eScience5
2025 Toward Dynamic Gaussian Rendering for Digital Twins
abstract
Recent innovations in 3D reconstruction and the rise in popularity of Digital Twins present a unique opportunity for the integration of the two technologies for high quality data visualization. In this work, I present an architecture for integrating Gaussian Splat reconstructions with dynamic data and user interaction, implemented in Unreal Engine. I also explore the application of this method to digital twin creation and how it addresses specific visualization challenges.
Eero Dunham, Mengjiao Han, Victor A. Mateevitsi, Joseph A. Insley, Michael E. Papka, Silvio Rizzi 0001, Janet Knowles
eScience5
2025 Toward Distributed 3D Gaussian Splatting for High-Resolution Isosurface Visualization
abstract
We present a multi-GPU extension of the 3D Gaussian Splatting (3D-GS) pipeline for scientific visualization. Building on previous work that demonstrated high-fidelity isosurface reconstruction using Gaussian primitives, we incorporate a multi-GPU training backend adapted from Grendel-GS to enable scalable processing of large datasets. By distributing optimization across GPUs, our method improves training throughput and supports high-resolution reconstructions that exceed single-GPU capacity. In our experiments, the system achieves a 5.6× speedup on the Kingsnake dataset (4M Gaussians) using four GPUs compared to a single-GPU baseline, and successfully trains the Miranda dataset (18M Gaussians) that is an infeasible task on a single A100 GPU. This work lays the groundwork for integrating 3D-GS into HPC-based scientific workflows, enabling real-time post hoc and in situ visualization of complex simulations.
Mengjiao Han, Andres Sewell, Joseph A. Insley, Janet Knowles, Victor A. Mateevitsi, Michael E. Papka, Steve Petruzza, Silvio Rizzi 0001
eScience6
2025 Intuitive Computational Steering Using Ascent and Trame
abstract
Large-scale scientific simulations that run on high-performance computing resources are typically executed in batch mode, where jobs are queued and run once sufficient resources become available. Visualization and analysis of simulation data can be performed post hoc or in situ, but in both cases, scientists typically do not gain insights until after the simulation has completed. Recently, bidirectional steering capabilities integrated into in situ libraries have enabled scientists to investigate and control simulations during runtime. In this paper, we present our work on developing a bridge between Ascent (a flyweight in situ processing and visualization library) and Trame (a web-based framework for developing interactive visualizations) to create a means for domain users to intuitively steer large-scale simulations. We demonstrate the effectiveness of web-based, user-friendly, customizable steering through three real-world use cases involving large-scale production simulations.
Thomas Marrinan, Andres Sewell, Victor A. Mateevitsi, Steve Petruzza, Jifu Tan, Dimitrios K. Fytanidis, Michael E. Papka
eScience7
2025 GENIUS: AI Powered Assistant for Scientific Research
abstract
General Experimentation and Natural Interface Utility System (GENIUS), is an AI personal assistant specifically tailored for scientists engaged in experimentation and research. GENIUS integrates Large Language Models (LLM), with immersive mixed reality (XR) capabilities. Through natural speech recognition, users can interact effortlessly with GENIUS to ask questions, visualize and manipulate complex 3D models, and execute computational jobs on Argonne Leadership Computing Facility (ALCF) supercomputers.
Ricky Massa, Aaqel Shaik, Brian Ta, Mengjiao Han, Joseph A. Insley, Janet Knowles, Victor A. Mateevitsi, Michael E. Papka, Silvio Rizzi 0001, Shilpika
eScience8
2025 Synthesis of Ground Truth for Neuronal Segmentation
abstract
Segmenting neuropil data for connectomics is limited by the high cost and effort of manually labeling Electron Microscopy (EM) volumes. To address this, we propose a two-stage framework to generate large amounts of realistic image data with ground-truth segmentations. First, we use Procedural Content Generation (PCG) with a parametric evolution strategy to synthesize diverse 3D neuron morphologies, which are combined into dense, noise-free segmentation masks. Second, these masks guide a conditional denoising diffusion model (DDPM) to generate realistic EM image patches, trained adversarially with a PatchGAN discriminator. This approach enables scalable production of high-quality training data, aiming to improve model accuracy and reduce manual annotation efforts.
Jyotsna Rajaraman, Thomas D. Uram, Kevin M. Boergens, Michael E. Papka
eScience4
2025 Modular Agentic System for Scientific Visualization in Mixed Reality
abstract
Mixed reality (MR) enables immersive, intuitive engagement with scientific data. When paired with AI-driven assistants, it has the potential to transform traditional workflows. In this paper, we introduce a modular agentic architecture for scientific visualization in MR, designed to balance general-purpose flexibility with domain-specific extensibility. Our modular architecture supports composable tools, contextual reasoning, and dynamic task execution. We outline a three-layer design, domain module integration, and orchestration of multistep workflows. We demonstrate the system’s capabilities through use cases in biology and general-purpose scientific visualization, including protein interaction networks and remote ParaView-based rendering. The result is a flexible and extensible foundation for spatial scientific computing.
Aaqel Shaik, Ricky Massa, Brian Ta, Mengjiao Han, Joseph A. Insley, Janet Knowles, Victor A. Mateevitsi, Michael E. Papka, Silvio Rizzi 0001, Shilpika
eScience8
2025 Machine Learning Driven Auto-Tuning for Non-Uniform All-to-All Collectives
abstract
Non-uniform all-to-all communication patterns present optimization challenges in parallel computing due to their irregular data distribution and dynamic behavior. While MPI_Alltoallv provides the standard interface for such exchanges, achieving optimal performance requires careful selection among multiple implementation variants and tuning of algorithm-specific parameters. This paper presents a data-driven autotuning framework that combines machine learning-based runtime prediction with a lookup-table mechanism for fast configuration selection. The ML model estimates the communication time of each algorithm configuration under a given system setup, allowing the framework to identify the optimal implementation and parameter set based on predicted performance. We validate our approach through comprehensive benchmarking of MPI_Alltoallv and two specialized algorithms across varying process counts, message sizes, and tunable parameters. Applied to a real MPI-based transitive closure application on the Fugaku supercomputer, our framework achieves up to$6.03 \times$reduction in communication time over the vendor implementation, providing a detailed understanding of non-uniform collective communication behavior and a practical framework for automatic performance optimization in HPC applications.
Kunting Qi, Jens Domke, Seydou Ba, Venkatram Vishwanath, Michael E. Papka, Sidharth Kumar
HiPC6
2025 More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
Michael E. Papka, Zhiling Lan
JSSPP2
2025 CQSim+: Symbiotic Simulation for Multi-Resource Scheduling in High-Performance Computing
Yash Kurkure, Shambhawi Sharma, Xin Wang 0115, Michael E. Papka, Zhiling Lan
SIGSIM-PADS4
2025 On the Effectiveness of Unified Memory in Multi-GPU Collective Communication
abstract
Modern supercomputers are becoming increasingly dense with accelerators. Industry leaders offer multi-GPU architectures with high interconnection bandwidth between the devices to match the requirements of modern workloads. While those technologies advance, it is up to the programmer to successfully exploit them. Recognizing this burden, multiple abstractions have been built. We focus on the NVIDIA Collective Communication Library (NCCL) and Unified Memory (UM). The former provides MPI-like directives integrated within the GPU runtime, allowing lower latencies and increasing the bandwidth over previous approaches. The latter simplifies the programming paradigm, offering a unified virtual address space. Moreover, it enables memory oversubscription, drastically reducing the efforts towards handling larger problems without completely restructuring the codebase. This work provides the first joint analysis of NCCL and UM from single-node multi-GPU architectures to a production supercomputer. We explore all the available collective communication directives concerning their power requirements and overall throughput. Moreover, we study the effects of various hyperparameters, e.g., message sizes, oversubscription level, and memory advice, on the overall obtainable performance. Our findings showcase how using UM brings negligible increased energy consumption; moreover, in distributed settings, other restricting factors, such as network bottlenecks, surpass the overhead introduced by UM’s page-eviction mechanisms.
Riccardo Strina, Ian Di Dio Lavore, Marco D. Santambrogio, Michael E. Papka, Zhiling Lan
PDP4
2025 Minimizing Power Waste in Heterogenous Computing via Adaptive Uncore Scaling
abstract
High-performance computing (HPC) systems are essential for scientific discovery and engineering innovation. However, their growing power demands pose significant challenges, particularly as systems scale to the exascale level. Prior uncore frequency tuning studies have primarily focused on conventional HPC workloads running on CPU-only systems. As HPC advances toward heterogeneous computing, integrating diverse GPU workloads on heterogeneous CPU-GPU systems, it becomes imperative to revisit and enhance uncore scaling. Our investigation reveals that uncore frequency scales down only when CPU power approaches its thermal design power (TDP), which is rare in GPU-dominant applications. As a result, modern computing systems experience unnecessary power waste. In this study, we present MAGUS, a user-transparent uncore frequency scaling runtime for heterogeneous computing. MAGUS dynamically adjusts uncore frequencies according to distinct application execution phases, effectively minimizing power waste caused by consistently using maximum uncore frequencies. Our design incorporates several key techniques, including real-time monitoring and prediction of memory accesses, intelligent handling of frequent phase transitions, and leveraging vendor-provided power management features. We evaluate MAGUS with various GPU benchmarks and applications on multiple heterogeneous systems with different CPU and GPU architectures. Experimental results demonstrate that MAGUS achieves up to 27% energy savings compared to the default settings, while maintaining a performance loss of less than 5% and an overhead of under 1%.
Seyfal Sultanov, Michael E. Papka, Zhiling Lan
SC3
2025 Trust Your Gut: Comparing Human and Machine Inference from Noisy Visualizations
abstract
People commonly utilize visualizations not only to examine a given dataset, but also to draw generalizable conclusions about the underlying models or phenomena. Prior research has compared human visual inference to that of an optimal Bayesian agent, with deviations from rational analysis viewed as problematic. However, human reliance on non-normative heuristics may prove advantageous in certain circumstances. We investigate scenarios where human intuition might surpass idealized statistical rationality. In two experiments, we examine individuals' accuracy in characterizing the parameters of known data-generating models from bivariate visualizations. Our findings indicate that, although participants generally exhibited lower accuracy compared to statistical models, they frequently outperformed Bayesian agents, particularly when faced with extreme samples. Participants appeared to rely on their internal models to filter out noisy visualizations, thus improving their resilience against spurious data. However, participants displayed overconfidence and struggled with uncertainty estimation. They also exhibited higher variance than statistical machines. Our findings suggest that analyst gut reactions to visualizations may provide an advantage, even when departing from rationality. These results carry implications for designing visual analytics tools, offering new perspectives on how to integrate statistical models and analyst intuition for improved inference and decision-making. The data and materials for this paper are available at https://osf.io/qmfv6.
Ratanond Koonchanok, Michael E. Papka, Khairi Reda
IEEE Trans. Vis. Comput. Graph.2
2025 Distributed Neural Representation for Reactive In Situ Visualization
abstract
Implicit neural representations (INRs) have emerged as a powerful tool for compressing large-scale volume data. This opens up new possibilities for in situ visualization. However, the efficient application of INRs to distributed data remains an underexplored area. In this work, we develop a distributed volumetric neural representation and optimize it for in situ visualization. Our technique eliminates data exchanges between processes, achieving state-of-the-art compression speed, quality and ratios. Our technique also enables the implementation of an efficient strategy for caching large-scale simulation data in high temporal frequencies, further facilitating the use of reactive in situ visualization in a wider range of scientific problems. We integrate this system with the Ascent infrastructure and evaluate its performance and usability using real-world simulations.
Qi Wu 0015, Joseph A. Insley, Victor A. Mateevitsi, Silvio Rizzi 0001, Michael E. Papka, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.5
2024 A Multi-Level, Multi-Scale Visual Analytics Approach to Assessment of Multifidelity HPC Systems
abstract
The ability to monitor and interpret hardware system events and behaviors is crucial to improving the robustness and reliability of these systems, especially in a supercomputing facility. The growing complexity and scale of these systems demand an increase in monitoring data collected at multiple fidelity levels and varying temporal resolutions. In this work, we aim to build a holistic analytical system that helps make sense of such massive data, mainly the hardware logs, job logs, and environment logs collected from disparate subsystems and components of a supercomputer system. This end-to-end log analysis system, coupled with visual analytics support, allows users to glean and promptly extract supercomputer usage and error patterns at varying temporal and spatial resolutions. We use multi-resolution dynamic mode decomposition (mrDMD), a technique that depicts high-dimensional data as correlated spatial-temporal variations patterns or modes, to extract variation patterns isolated at specified frequencies. Our improvements to the mrDMD algorithm help promptly reveal useful information in the massive environment log dataset, which is then associated with the processed hardware and job log datasets using our visual analytics system. Furthermore, our system can identify the usage and error patterns filtered at user, project, and subcomponent levels. We exemplify the effectiveness of our approach with two use scenarios with the Cray XC40 supercomputer.
Shilpika, Bethany Lusch, Murali Emani, Filippo Simini, Venkatram Vishwanath, Michael E. Papka, Kwan-Liu Ma
CCGrid6
2024 Color Maker: a Mixed-Initiative Approach to Creating Accessible Color Maps
abstract
Quantitative data is frequently represented using color, yet designing effective color mappings is a challenging task, requiring one to balance perceptual standards with personal color preference. Current design tools either overwhelm novices with complexity or offer limited customization options. We present ColorMaker, a mixed-initiative approach for creating colormaps. ColorMaker combines fluid user interaction with real-time optimization to generate smooth, continuous color ramps. Users specify their loose color preferences while leaving the algorithm to generate precise color sequences, meeting both designer needs and established guidelines. ColorMaker can create new colormaps, including designs accessible for people with color-vision deficiencies, starting from scratch or with only partial input, thus supporting ideation and iterative refinement. We show that our approach can generate designs with similar or superior perceptual characteristics to standard colormaps. A user study demonstrates how designers of varying skill levels can use this tool to create custom, high-quality colormaps. ColorMaker is available at: colormaker.org
Amey A. Salvi, Kecheng Lu 0002, Michael E. Papka, Yunhai Wang, Khairi Reda
CHI3
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC25
2024 Image Synthesis from a Collection of Depth Enhanced Panoramas: Creating Interactive Extended Reality Experiences from Static Images
abstract
Stereoscopic 360° panoramas are a popular modality for creating cinematic virtual reality experiences. However, media in this format are typically static entities that consumers passively view. This is because objects visible in the scene, and camera properties such as depth of field, must be determined at capture time. We propose a real-time technique for dynamically synthesizing stereoscopic 360° panoramas from a collection of depth enhanced monoscopic panoramas. Using depth information, pixels can be transformed into three-dimensional space and re-projected to a different camera location. Our technique allows for head-motion parallax, dynamic depth of field, and integration of properly occluded virtual objects into the captured scene. Our technique shows minimal discrepancies compared to ground truth captures and reduces error when compared to existing stereoscopic 360° panoramic synthesis techniques. Additionally, our technique makes creating interactive extended reality experiences more accessible since monoscopic 360° cameras are much more common than their stereoscopic counterparts.
Thomas Marrinan, Ethan Honzik, Hal L. N. Brynteson, Michael E. Papka
IMX4
2024 Science in a Blink: Supporting Ensemble Perception in Scalar Fields
abstract
Visualizations support rapid analysis of scientific datasets, allowing viewers to glean aggregate information (e.g., the mean) within split-seconds. While prior research has explored this ability in conventional charts, it is unclear if spatial visualizations used by computational scientists afford a similar ensemble perception capacity. We investigate people’s ability to estimate two summary statistics, mean and variance, from pseudocolor scalar fields. In a crowd- sourced experiment, we find that participants can reliably characterize both statistics, although variance discrimination requires a much stronger signal. Multi-hue and diverging colormaps outperformed monochromatic, luminance ramps in aiding this extraction. Analysis of qualitative responses suggests that participants often estimate the distribution of hotspots and valleys as visual proxies for data statistics. These findings suggest that people’s summary interpretation of spatial datasets is likely driven by the appearance of discrete color segments, rather than assessments of overall luminance. Implicit color segmentation in quantitative displays could thus prove more useful than previously assumed by facilitating quick, gist- level judgments about color-coded visualizations.
Victor A. Mateevitsi, Michael E. Papka, Khairi Reda
IEEE VIS2
2024 MalleTrain: Deep Neural Networks Training on Unfillable Supercomputer Nodes
abstract
First-come first-serve scheduling can result in substantial (up to 10%) of transiently idle nodes on supercomputers. Recognizing that such unfilled nodes are well-suited for deep neural network (DNN) training, due to the flexible nature of DNN training tasks, Liu et al. proposed that the re-scaling DNN training tasks to fit gaps in schedules be formulated as a mixed-integer linear programming (MILP) problem, and demonstrated via simulation the potential benefits of the approach. Here, we introduce MalleTrain, a system that provides the first practical implementation of this approach and that furthermore generalizes it by allowing it to be used even for DNN training applications for which model information is unknown before runtime. Key to this latter innovation is the use of a lightweight online job profiling advisor (JPA) to collect critical scalability information for DNN jobs---information that it then employs to optimize resource allocations dynamically, in real time. We describe the MalleTrain architecture and present the results of a detailed experimental evaluation on a supercomputer GPU cluster and several representative DNN training workloads, including neural architecture search and hyperparameter optimization. Our results not only confirm the practical feasibility of leveraging idle supercomputer nodes for DNN training but improve significantly on prior results, improving training throughput by up to 22.3% without requiring users to provide job scalability information.
Feng Yan 0001, Lei Yang 0001, Ian T. Foster, Michael E. Papka, Zhengchun Liu, Rajkumar Kettimuthu
ICPE5
2023 FreeTrain: A Framework to Utilize Unused Supercomputer Nodes for Training Neural Networks
abstract
Supercomputer scheduling policies commonly result in many transient idle nodes, a phenomenon that is only partially alleviated by backfill scheduling methods that promote small jobs to run before large jobs. Here we describe how to realize a novel use for these otherwise wasted resources, namely, deep neural network (DNN) training. This important workload is easily organized as many small fragments that can be configured dynamically to fit essentially any node × time hole in a supercomputer's schedule. We describe how the task of rescaling suitable DNN training tasks to fit dynamically changing holes can be formulated as a deterministic mixed integer linear programming (MILP)-based resource allocation algorithm, and show that this MILP problem can be solved efficiently at run time. We show further how this MILP problem can be adapted to optimize for administrator- or user-defined metrics. We validate our method with supercomputer scheduler logs and different DNN training scenarios, and demonstrate efficiencies of up to 93% compared with running the same training tasks on dedicated nodes. Our method thus enables substantial supercomputer resources to be allocated to DNN training with no impact on other applications.
Zhengchun Liu, Rajkumar Kettimuthu, Michael E. Papka, Ian T. Foster
CCGrid3
2023 Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
abstract
In the field of high-performance computing (HPC), there has been recent exploration into the use of deep reinforcement learning for cluster scheduling (DRL scheduling), which has demonstrated promising outcomes. However, a significant challenge arises from the lack of interpretability in deep neural networks (DNN), rendering them as black-box models to system managers. This lack of model interpretability hinders the practical deployment of DRL scheduling. In this work, we present a framework called IRL (Interpretable Reinforcement Learning) to address the issue of interpretability of DRL scheduling. The core idea is to interpret DNN (i.e., the DRL policy) as a decision tree by utilizing imitation learning. Unlike DNN, decision tree models are non-parametric and easily comprehensible to humans. To extract an effective and efficient decision tree, IRL incorporates the Dataset Aggregation (DAgger) algorithm and introduces the notion of critical state to prune the derived decision tree. Through trace-based experiments, we demonstrate that IRL is capable of converting a black-box DNN policy into an interpretable rule-based decision tree while maintaining comparable scheduling performance. Additionally, IRL can contribute to the setting of rewards in DRL scheduling.
Boyang Li 0018, Zhiling Lan, Michael E. Papka
MASCOTS3
2023 ChemoGraph: Interactive Visual Exploration of the Chemical Space
abstract
Abstract Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.
Bharat Kale, Austin Clyde, Maoyuan Sun, Arvind Ramanathan, Rick L. Stevens, Michael E. Papka
Comput. Graph. Forum6
2023 The State of the Art in Visualizing Dynamic Multivariate Networks
abstract
Abstract Most real‐world networks are both dynamic and multivariate in nature, meaning that the network is associated with various attributes and both the network structure and attributes evolve over time. Visualizing dynamic multivariate networks is of great significance to the visualization community because of their wide applications across multiple domains. However, it remains challenging because the techniques should focus on representing the network structure, attributes and their evolution concurrently. Many real‐world network analysis tasks require the concurrent usage of the three aspects of the dynamic multivariate networks. In this paper, we analyze current techniques and present a taxonomy to classify the existing visualization techniques based on three aspects: temporal encoding, topology encoding, and attribute encoding. Finally, we survey application areas and evaluation methods; and discuss challenges for future research.
Bharat Kale, Maoyuan Sun, Michael E. Papka
Comput. Graph. Forum3
2022 Toward an In-Depth Analysis of Multifidelity High Performance Computing Systems
abstract
To maintain a robust and reliable supercomputing facility, monitoring it and understanding its hardware system events and behaviors is an essential task. Exascale systems will be increasingly heterogeneous, and the volume of systems data, collected from multiple subsystems and components measured at multiple fidelity levels and temporal resolutions, will continue to grow. In this work, we aim to create an effective solution to analyze diverse and massive datasets gathered from the error logs, job logs, and environment logs of an HPC system, such as a Cray XC40 supercomputer. In this work, we build an end-to-end error log analysis system that analyzes the job logs and gleans insights from their correspondence with hardware error logs and environment logs despite their varying temporal and spatial resolutions. Our machine learning pipeline built in our system is ~92% accurate in predicting the job exit status and does so with sufficient lead time for evasive actions to be taken before the actual failure event occurs.
Shilpika, Bethany Lusch, Murali Emani, Filippo Simini, Venkatram Vishwanath, Michael E. Papka, Kwan-Liu Ma
CCGRID6
2022 MRSch: Multi-Resource Scheduling for HPC
abstract
Emerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU -centric. This advancement forces cluster schedulers to consider multiple schedulable resources during decision-making. Existing scheduling studies rely on heuristic or optimization methods, which are limited by an inability to adapt to new scenarios for ensuring long-term scheduling performance. We present an intelligent scheduling agent named MRSch for multi-resource scheduling in HPC that leverages direct future prediction (DFP), an advanced multi-objective reinforcement learning algorithm. While DFP demonstrated outstanding performance in a gaming competition, it has not been previously explored in the context of HPC scheduling. Several key techniques are developed in this study to tackle the challenges involved in multi-resource scheduling. These techniques enable MRSch to learn an appropriate scheduling pol-icy automatically and dynamically adapt its policy in response to workload changes via dynamic resource prioritizing. We compare MRSch with existing scheduling methods through extensive trace-base simulations. Our results demonstrate that MRSch improves scheduling performance by up to 48 % compared to the existing scheduling methods.
Boyang Li 0018, Yuping Fan, Matthew T. Dearing, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka
CLUSTER7
2022 Hybrid Workload Scheduling on HPC Systems
abstract
Traditionally, on-demand, rigid, and malleable applications have been scheduled and executed on separate systems. The ever-growing workload demands and rapidly developing HPC infrastructure trigger the interest of converging these applications on a single HPC system. Although allocating the hybrid workloads within one system could potentially improve system efficiency, it is difficult to balance the tradeoff between the responsiveness of on-demand requests, incentive for malleable jobs, and the performance of rigid applications. In this study, we present several scheduling mechanisms to address the issues involved in co-scheduling on-demand, rigid, and malleable jobs on a single HPC system. We extensively evaluate and compare their performance under various configurations and workloads. Our experimental results show that our proposed mechanisms are capable of serving on-demand workloads with minimal delay, offering incentives for declaring malleability, and improving system performance.
Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka
IPDPS5
2022 Encoding for Reinforcement Learning Driven Scheduling
Boyang Li 0018, Yuping Fan, Michael E. Papka, Zhiling Lan
JSSPP3
2022 DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing
abstract
Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. The experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.
Yuping Fan, Boyang Li 0018, Dustin Favorite, Naunidh Singh, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan
IEEE Trans. Parallel Distributed Syst.8
2021 Deep Reinforcement Agent for Scheduling in HPC
abstract
Cluster scheduler is crucial in high-performance computing (HPC). It determines when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a novel, hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. A unique training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. The experiments with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 45%.
Yuping Fan, Zhiling Lan, John T. Childers, Paul M. Rich, William E. Allcock, Michael E. Papka
IPDPS6
2021 Color Nameability Predicts Inference Accuracy in Spatial Visualizations
abstract
Abstract Color encoding is foundational to visualizing quantitative data. Guidelines for colormap design have traditionally emphasized perceptual principles, such as order and uniformity. However, colors also evoke cognitive and linguistic associations whose role in data interpretation remains underexplored. We study how two linguistic factors, name salience and name variation, affect people's ability to draw inferences from spatial visualizations. In two experiments, we found that participants are better at interpreting visualizations when viewing colors with more salient names (e.g., prototypical ‘blue’, ‘yellow’, and ‘red’ over ‘teal’, ‘beige’, and ‘maroon’). The effect was robust across four visualization types, but was more pronounced in continuous (e.g., smooth geographical maps) than in similar discrete representations (e.g., choropleths). Participants' accuracy also improved as the number of nameable colors increased, although the latter had a less robust effect. Our findings suggest that color nameability is an important design consideration for quantitative colormaps, and may even outweigh traditional perceptual metrics. In particular, we found that the linguistic associations of color are a better predictor of performance than the perceptual properties of those colors. We discuss the implications and outline research opportunities. The data and materials for this study are available at https://osf.io/asb7n
Khairi Reda, Amey A. Salvi, Jack Gray, Michael E. Papka
Comput. Graph. Forum4
2021 Real-Time Omnidirectional Stereo Rendering: Generating 360° Surround-View Panoramic Images for Comfortable Immersive Viewing
abstract
Surround-view panoramic images and videos have become a popular form of media for interactive viewing on mobile devices and virtual reality headsets. Viewing such media provides a sense of immersion by allowing users to control their view direction and experience an entire environment. When using a virtual reality headset, the level of immersion can be improved by leveraging stereoscopic capabilities. Stereoscopic images are generated in pairs, one for the left eye and one for the right eye, and result in providing an important depth cue for the human visual system. For computer generated imagery, rendering proper stereo pairs is well known for a fixed view. However, it is much more difficult to create omnidirectional stereo pairs for a surround-view projection that work well when looking in any direction. One major drawback of traditional omnidirectional stereo images is that they suffer from binocular misalignment in the peripheral vision as a user's view direction approaches the zenith / nadir (north / south pole) of the projection sphere. This paper presents a real-time geometry-based approach for omnidirectional stereo rendering that fits into the standard rendering pipeline. Our approach includes tunable parameters that enable pole merging - a reduction in the stereo effect near the poles that can minimize binocular misalignment. Results from a user study indicate that pole merging reduces visual fatigue and discomfort associated with binocular misalignment without inhibiting depth perception.
Thomas Marrinan, Michael E. Papka
IEEE Trans. Vis. Comput. Graph.2
2020 Characterization and identification of HPC applications at leadership computing facility
abstract
High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation.
Zhengchun Liu, Ryan Lewis, Rajkumar Kettimuthu, Kevin Harms, Philip H. Carns, Nageswara S. V. Rao, Ian T. Foster, Michael E. Papka
ICS8
2019 Scheduling Beyond CPUs for HPC
abstract
High performance computing (HPC) is undergoing significant changes. The emerging HPC applications comprise both compute- and data-intensive applications. To meet the intense I/O demand from emerging data-intensive applications, burst buffers are deployed in production systems. Existing HPC schedulers are mainly CPU-centric. The extreme heterogeneity of hardware devices, combined with workload changes, forces the schedulers to consider multiple resources (e.g., burst buffers) beyond CPUs, in decision making. In this study, we present a multi-resource scheduling scheme named BBSched that schedules user jobs based on not only their CPU requirements, but also other schedulable resources such as burst buffer. BBSched formulates the scheduling problem into a multi-objective optimization (MOO) problem and rapidly solves the problem using a multi-objective genetic algorithm. The multiple solutions generated by BBSched enables system managers to explore potential tradeoffs among various resources, and therefore obtains better utilization of all the resources. The trace-driven simulations with real system workloads demonstrate that BBSched improves scheduling performance by up to 41% compared to existing methods, indicating that explicitly optimizing multiple resources beyond CPUs is essential for HPC scheduling.
Yuping Fan, Zhiling Lan, Paul M. Rich, William E. Allcock, Michael E. Papka, Brian Austin, David Paul
HPDC5
2018 Topology-aware space-shared co-analysis of large-scale molecular dynamics simulations
Preeti Malakar, Todd S. Munson, Christopher Knight 0001, Venkatram Vishwanath, Michael E. Papka
SC5
2018 Scalable pCT Image Reconstruction Delivered as a Cloud Service
abstract
We describe a cloud-based medical image reconstruction service designed to meet a real-time and daily demand to reconstruct thousands of images from proton cancer treatment facilities worldwide. Rapid reconstruction of a three-dimensional Proton Computed Tomography (pCT) image can require the transfer of 100 GB of data and use of approximately 120 GPU-enabled compute nodes. The nature of proton therapy means that demand for such a service is sporadic and comes from potentially hundreds of clients worldwide. We thus explore the use of a commercial cloud as a scalable and cost-efficient platform for pCT reconstruction. To address the high performance requirements of this application we leverage Amazon Web Services' GPU-enabled cluster resources that are provisioned with high performance networks between nodes. To support episodic demand, we develop an on-demand multi-user provisioning service that can dynamically provision and resize clusters based on image reconstruction requirements, priorities, and wait times. We compare the performance of our pCT reconstruction service running on commercial cloud resources with that of the same application on dedicated local high performance computing resources. We show that we can achieve scalable and on-demand reconstruction of large scale pCT images for simultaneous multi-client requests, processing images in less than 10 minutes for less than $10 per image.
Ryan Chard, Ravi K. Madduri, Nicholas T. Karonis, Kyle Chard, Kirk L. Duffin, Caesar E. Ordoñez, Thomas D. Uram, Justin Fleischauer, Ian T. Foster, Michael E. Papka, John Winans
IEEE Trans. Cloud Comput.10
2018 System-wide trade-off modeling of performance, power, and resilience on petascale systems
Li Yu 0006, Zhou Zhou 0006, Yuping Fan, Michael E. Papka, Zhiling Lan
J. Supercomput.4
2017 Trade-Off Between Prediction Accuracy and Underestimation Rate in Job Runtime Estimates
abstract
Job runtime estimates provided by users are widely acknowledged to be overestimated and runtime overestimation can greatly degrade job scheduling performance. Previous studies focus on improving accuracy of job runtime estimates by reducing runtime overestimation, but fail to address the underestimation problem (i.e., the underestimation of job runtimes). Using an underestimated runtime is catastrophic to a job as the job will be killed by the scheduler before completion. We argue that both the improvement of runtime accuracy and the reduction of underestimation rate are equally important. To address this problem, we propose an online runtime adjustment framework called TRIP. TRIP explores the data censoring capability of the Tobit model to improve prediction accuracy while keeping a low underestimation rate of job runtimes. TRIP can be used as a plugin to job scheduler for improving job runtime estimates and hence boosting job scheduling performance. Preliminary results demonstrate that TRIP is capable of achieving high accuracy of 80% and low underestimation rate of 5%. This is significant as compared to other well-known machine learning methods such as SVM, Random Forest, and Last-2 which result in a high underestimation rate (20%-50%). Our experiments further quantify the amount of scheduling performance gain achieved by the use of TRIP.
Yuping Fan, Paul M. Rich, William E. Allcock, Michael E. Papka, Zhiling Lan
CLUSTER4
2016 Towards optimizing large-scale data transfers with end-to-end integrity verification
abstract
The scale of scientific data generated by experimental facilities and simulations on high-performance computing facilities has been growing rapidly. In many cases, this data needs to be transferred rapidly and reliably to remote facilities for storage, analysis, sharing etc. At the same time, users want to verify the integrity of the data by doing a checksum after the data has been written to disk at the destination, to ensure the file has not been corrupted, for example due to network or storage data corruption, software bugs or human error. This end-to-end integrity verification creates additional overhead (extra disk I/O and more computation) and increases the overall data transfer time. In this paper, we evaluate strategies to maximize the overlap between data transfer and checksum computation. More specifically, we evaluate file-level and block-level (with various block sizes) pipelining to overlap data transfer and checksum computation. We evaluate these pipelining approaches in the context of GridFTP, a widely used protocol for science data transfers. We conducted both theoretical analysis and real experiments to evaluate our methods. The results show that block-level pipelining is an effective method in maximizing the overlap between data transfer and checksum computation and can improve the overall data transfer time with end-to-end integrity verification by up to 70% compared to the sequential execution of transfer and checksum, and by up to 60% compared to file-level pipelining.
Eun-Sung Jung, Rajkumar Kettimuthu, Xian-He Sun, Michael E. Papka
IEEE BigData5
2016 Optimal execution of co-analysis for large-scale molecular dynamics simulations
abstract
The analysis of scientific simulation data enables scientists to derive insights from their simulations. This analysis of the simulation output can be performed at the same execution site as the simulation using the same resources or can be done at a different site. The optimal output frequency is challenging to decide and is often chosen empirically. We propose a mathematical formulation for choosing the optimal frequency of data transfer for analysis and the feasibility of performing the analysis, under the given resource constraints such as network bandwidth, disk space, available memory, and computation time. We propose formulations for two cases of co-analysis - local and remote. We consider various analyses features such as computation time, input data and memory requirement, importance of the analysis and minimum frequency required for performing the analysis. We demonstrate the effectiveness of our approach using molecular dynamics applications on the Mira and Edison supercomputers.
Preeti Malakar, Venkatram Vishwanath, Christopher Knight 0001, Todd S. Munson, Michael E. Papka
SC5
2016 A data driven scheduling approach for power management on HPC systems
abstract
Modern schedulers running on HPC systems traditionally consider the number of resources and the time requested for each job that is to be executed when making scheduling decisions. Until recently this has been sufficient, however as systems get larger, other metrics like power consumption become necessary to ensure system stability. In this paper, we propose a data driven scheduling approach for controlling the power consumption of the entire system under any user defined budget. Here, “data driven” means that our approach actively observes, analyzes, and assesses power behaviors of the system and user jobs to guide scheduling decisions for power management. This design is based on the key observation that HPC jobs have distinct power profiles. Our work contains an empirical analysis of workload power characteristics on a production system, dynamic learner to estimate the job power profile for scheduling, and an online power-aware scheduler for managing the overall system power. Using real workload traces, we demonstrate that our design effectively controls system power consumption while minimizing the impact on system utilization.
Sean Wallace, Xu Yang 0009, Venkatram Vishwanath, William E. Allcock, Susan Coghlan, Michael E. Papka, Zhiling Lan
SC6
2016 Interactive Multi-Modal Display Spaces for Visual Analysis
abstract
Classic visual analysis relies on a single medium for displaying and interacting with data. Large-scale tiled display walls, virtual reality using head-mounted displays or CAVE systems, and collaborative touch screens have all been utilized for data exploration and analysis. We present our initial findings of combining numerous display environments and input modalities to create an interactive multi-modal display space that enables researchers to leverage various pieces of technology that will best suit specific sub-tasks. Our main contributions are 1) the deployment of an input server that interfaces with a wide array of interaction devices to create a single uniform stream of data usable by custom visual applications, and 2) three real-world use cases of leveraging multiple display environments in conjunction with one another to enhance scientific discovery and data dissemination.
Thomas Marrinan, Arthur Nishimoto, Joseph A. Insley, Silvio Rizzi 0001, Andrew E. Johnson 0001, Michael E. Papka
ISS6
2016 Improving sparse data movement performance using multiple paths on the Blue Gene/Q supercomputer
Huy Bui, Eun-Sung Jung, Venkatram Vishwanath, Andrew E. Johnson 0001, Jason Leigh, Michael E. Papka
Parallel Comput.6
2016 Visualizing multiphysics, fluid-structure interaction phenomena in intracranial aneurysms
Paris Perdikaris, Joseph A. Insley, Leopold Grinberg, Yue Yu 0011, Michael E. Papka, George Em Karniadakis
Parallel Comput.5
2016 Application power profiling on IBM Blue Gene/Q
Sean Wallace, Zhou Zhou 0006, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka
Parallel Comput.7
2016 Interstitial and Interlayer Ion Diffusion Geometry Extraction in Graphitic Nanosphere Battery Materials
abstract
Large-scale molecular dynamics (MD) simulations are commonly used for simulating the synthesis and ion diffusion of battery materials. A good battery anode material is determined by its capacity to store ion or other diffusers. However, modeling of ion diffusion dynamics and transport properties at large length and long time scales would be impossible with current MD codes. To analyze the fundamental properties of these materials, therefore, we turn to geometric and topological analysis of their structure. In this paper, we apply a novel technique inspired by discrete Morse theory to the Delaunay triangulation of the simulated geometry of a thermally annealed carbon nanosphere. We utilize our computed structures to drive further geometric analysis to extract the interstitial diffusion structure as a single mesh. Our results provide a new approach to analyze the geometry of the simulated carbon nanosphere, and new insights into the role of carbon defect size and distribution in determining the charge capacity and charge dynamics of these carbon based battery materials.
Attila Gyulassy, Aaron Knoll, Kah Chun Lau, Bei Wang 0001, Peer-Timo Bremer, Michael E. Papka, Larry A. Curtiss, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.6
2015 Effects of Display Size and Resolution on User Behavior and Insight Acquisition in Visual Exploration
abstract
Large high-resolution displays are becoming increasingly common in research settings, providing data scientists with visual interfaces for the analysis of large datasets. Numerous studies have demonstrated unique perceptual and cognitive benefits afforded by these displays in visual analytics and information visualization tasks. However, the effects of these displays on knowledge discovery in exploratory visual analysis are still poorly understood. We present the results of a small-scale study to better understand how display size and resolution affect insight. Analyzing participants' verbal statements, we find preliminary evidence that larger displays with more pixels can significantly increase the number of discoveries reported during visual exploration, while yielding broader, more integrative insights. Furthermore, we find important differences in how participants performed the same visual exploration task using displays of varying sizes. We tie these results to extant work and propose explanations by considering the cognitive and interaction costs associated with visual exploration.
Khairi Reda, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh
CHI3
2015 Multipath Load Balancing for M × N Communication Patterns on the Blue Gene/Q Supercomputer Interconnection Network
abstract
Achievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually have a wide variation from application to application and even from phase to phase in an application, routing algorithms have a limited variation and usually are optimized for typical communication patterns. This results in high networking performance for favored communication patterns but low networking performance for others. In this paper we present approaches for improving networking performance by rebalancing load on physical links on the Blue Gene Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 30% higher throughput on experiment with data and patterns from a real application. The improvement can be up to several times higher throughput than default MPI_Alltoallv used in the Blue Gene Q supercomputer for certain communication patterns.
Huy Bui, Robert L. Jacob, Preeti Malakar, Venkatram Vishwanath, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh
CLUSTER6
2015 Comparison of Vendor Supplied Environmental Data Collection Mechanisms
abstract
The high performance computing landscape is filled with diverse hardware components. A large part of understanding how these components compare to others is by looking at the various environmental aspects of these devices such as power consumption, temperature, etc. Thankfully, vendors of these various pieces of hardware have supported this by providing mechanisms to obtain this data. However, differences not only in the way this data is obtained but also the data which is provided is common between products. In this paper, we take a comprehensive look at the data which is available for the most common pieces of today's HPC landscape, as well as how this data is obtained and how accurate it is. Having surveyed these components, we compare and contrast them noting key differences as well as providing insight into what features future components should have.
Sean Wallace, Venkatram Vishwanath, Susan Coghlan, Zhiling Lan, Michael E. Papka
CLUSTER5
2015 Improving Communication Throughput by Multipath Load Balancing on Blue Gene/Q
abstract
Achievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications, the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually vary from application to application and even from phase to phase in an application, routing algorithms have limited variation and usually are optimized for typical communication patterns. This results in high networking performance for some communication patterns. In this paper we present approaches for improving communication performance by using multiple paths and re-balancing load on physical links on the Blue Gene/Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 43 -- 67% higher throughput on average from 91 experiments, and can achieve higher throughput than default MPI_Alltoallv used for certain communication patterns.
Huy Bui, Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Eun-Sung Jung, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh
HiPC7
2015 Optimal scheduling of in-situ analysis for large-scale scientific simulations
abstract
Today's leadership computing facilities have enabled the execution of transformative simulations at unprecedented scales. However, analyzing the huge amount of output from these simulations remains a challenge. Most analyses of this output is performed in post-processing mode at the end of the simulation. The time to read the output for the analysis can be significantly high due to poor I/O bandwidth, which increases the end-to-end simulation-analysis time. Simulation-time analysis can reduce this end-to-end time. In this work, we present the scheduling of in-situ analysis as a numerical optimization problem to maximize the number of online analyses subject to resource constraints such as I/O bandwidth, network bandwidth, rate of computation and available memory. We demonstrate the effectiveness of our approach through two application case studies on the IBM Blue Gene/Q system.
Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Christopher Knight 0001, Mark Hereld, Sven Leyffer, Michael E. Papka
SC7
2015 Quantitative modeling of power performance tradeoffs on extreme scale systems
Li Yu 0006, Zhou Zhou 0006, Sean Wallace, Michael E. Papka, Zhiling Lan
J. Parallel Distributed Comput.4
2014 Exploring void search for fault detection on extreme scale systems
abstract
Mean Time Between Failures (MTBF), now calculated in days or hours, is expected to drop to minutes on exascale machines. The advancement of resilience technologies greatly depends on a deeper understanding of faults arising from hardware and software components. This understanding has the potential to help us build better fault tolerance technologies. For instance, it has been proved that combining checkpointing and failure prediction leads to longer checkpoint intervals, which in turn leads to fewer total checkpoints. In this paper we present a new approach for fault detection based on the Void Search (VS) algorithm. VS is used primarily in astrophysics for finding areas of space that have a very low density of galaxies. We evaluate our algorithm using real environmental logs from Mira Blue Gene/Q supercomputer at Argonne National Laboratory. Our experiments show that our approach can detect almost all faults (i.e., sensitivity close to 1) with a low false positive rate (i.e., specificity values above 0.7). We also compare our algorithm with a number of existing detection algorithms, and find that ours outperforms all of them.
Eduardo Berrocal, Li Yu 0006, Sean Wallace, Michael E. Papka, Zhiling Lan
CLUSTER4
2014 Scalable Parallel I/O on a Blue Gene/Q Supercomputer Using Compression, Topology-Aware Data Aggregation, and Subfiling
abstract
In this paper, we propose an approach to improving the I/O performance of an IBM Blue Gene/Q supercomputing system using a novel framework that can be integrated into high performance applications. We take advantage of the system's tremendous computing resources and high interconnection bandwidth among compute nodes to efficiently exploit I/O bandwidth. This approach focuses on lossless data compression, topology-aware data movement, and subfiling. The efficacy of this solution is demonstrated using microbenchmarks and an application-level benchmark.
Huy Bui, Hal Finkel, Venkatram Vishwanath, Salman Habib 0002, Katrin Heitmann, Jason Leigh, Michael E. Papka, Kevin Harms
PDP7
2014 RBF Volume Ray Casting on Multicore and Manycore CPUs
abstract
Abstract Modern supercomputers enable increasingly large N‐body simulations using unstructured point data. The structures implied by these points can be reconstructed implicitly. Direct volume rendering of radial basis function (RBF) kernels in domain‐space offers flexible classification and robust feature reconstruction, but achieving performant RBF volume rendering remains a challenge for existing methods on both CPUs and accelerators. In this paper, we present a fast CPU method for direct volume rendering of particle data with RBF kernels. We propose a novel two‐pass algorithm: first sampling the RBF field using coherent bounding hierarchy traversal, then subsequently integrating samples along ray segments. Our approach performs interactively for a range of data sets from molecular dynamics and astrophysics up to 82 million particles. It does not rely on level of detail or subsampling, and offers better reconstruction quality than structured volume rendering of the same data, exhibiting comparable performance and requiring no additional preprocessing or memory footprint other than the BVH. Lastly, our technique enables multi‐field, multi‐material classification of particle data, providing better insight and analysis.
Aaron Knoll, Ingo Wald, Paul A. Navrátil, Anne Bowen, Khairi Reda, Michael E. Papka, Kelly P. Gaither
Comput. Graph. Forum6
2013 Application power profiling on IBM Blue Gene/Q
abstract
The power consumption of state of the art supercomputers, because of their complexity and unpredictable workloads, is extremely difficult to estimate. Accurate and precise results, as are now possible with the latest generation of supercomputers, are therefore a welcome addition to the landscape. Only recently have end users been afforded the ability to access the power consumption of their applications. However, just because it's possible for end users to obtain this data does not mean it's a trivial task. This emergence of new data is therefore not only understudied, but also not fully understood. In this paper, we provide detailed power consumption analysis of microbenchmarks running on Argonne's latest generation of IBM Blue Gene supercomputers, Mira, a Blue Gene/Q system. The analysis is done utilizing our power monitoring library, MonEQ, built on the IBM provided Environmental Monitoring (EMON) API. We describe the importance of sub-second polling of various power domains and the implications they present. To this end, previously well understood applications will now have new facets of potential analysis.
Sean Wallace, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka
CLUSTER6
2013 A Generic High-Performance Method for Deinterleaving Scientific Data
Eric R. Schendel, Steve Harenberg, Houjun Tang, Venkatram Vishwanath, Michael E. Papka, Nagiza F. Samatova
Euro-Par5
2013 Scalable in situ scientific data encoding for analytical query processing
Sriram Lakshminarasimhan, David A. Boyuka II, Saurabh V. Pendse, Xiaocheng Zou, John Jenkins, Venkatram Vishwanath, Michael E. Papka, Nagiza F. Samatova
HPDC7
2013 Early Experience on the Blue Gene/Q Supercomputing System
abstract
The Argonne Leadership Computing Facility (ALCF) is home to Mira, a 10 PF Blue Gene/Q (BG/Q) system. The BG/Q system is the third generation in Blue Gene architecture from IBM and like its predecessors combines system-onchip technology with a proprietary interconnect (5-D torus). Each compute node has 16 augmented PowerPC A2 processor cores with support for simultaneous multithreading, 4-wide double precision SIMD, and different data prefetching mechanisms. Mira offers several new opportunities for tuning and scaling scientific applications. This paper discusses our early experience with a subset of micro-benchmarks, MPI benchmarks, and a variety of science and engineering applications running at ALCF. Both performance and power are studied and results on BG/Q is compared with its predecessor BG/P. Several lessons gleaned from tuning applications on the BG/Q architecture for better performance and scalability are shared.
Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Jiayuan Meng, Michael E. Papka
IPDPS5
2013 Characterization and modeling of PIDX parallel I/O for performance optimization
abstract
Parallel I/O library performance can vary greatly in response to user-tunable parameter values such as aggregator count, file count, and aggregation strategy. Unfortunately, manual selection of these values is time consuming and dependent on characteristics of the target machine, the underlying file system, and the dataset itself. Some characteristics, such as the amount of memory per core, can also impose hard constraints on the range of viable parameter values. In this work we address these problems by using machine learning techniques to model the performance of the PIDX parallel I/O library and select appropriate tunable parameter values. We characterize both the network and I/O phases of PIDX on a Cray XE6 as well as an IBM Blue Gene/P system. We use the results of this study to develop a machine learning model for parameter space exploration and performance prediction.
Sidharth Kumar, Avishek Saha, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert Latham, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci
SC11
2013 Integrating dynamic pricing of electricity into energy aware scheduling for HPC systems
abstract
The research literature to date mainly aimed at reducing energy consumption in HPC environments. In this paper we propose a job power aware scheduling mechanism to reduce HPC's electricity bill without degrading the system utilization. The novelty of our job scheduling mechanism is its ability to take the variation of electricity price into consideration as a means to make better decisions of the timing of scheduling jobs with diverse power profiles. We verified the effectiveness of our design by conducting trace-based experiments on an IBM Blue Gene/P and a cluster system as well as a case study on Argonne's 48-rack IBM Blue Gene/Q system. Our preliminary results show that our power aware algorithm can reduce electricity bill of HPC systems as much as 23%.
Xu Yang 0009, Zhou Zhou 0006, Sean Wallace, Zhiling Lan, Wei Tang 0001, Susan Coghlan, Michael E. Papka
SC7
2013 Distributed and hardware accelerated computing for clinical medical imaging using proton computed tomography (pCT)
Nicholas T. Karonis, Kirk L. Duffin, Caesar E. Ordoñez, Béla Erdélyi, Thomas D. Uram, Eric C. Olson, George Coutrakon, Michael E. Papka
J. Parallel Distributed Comput.8
2012 Volume rendering with multidimensional peak finding
abstract
Peak finding provides more accurate classification for direct volume rendering by sampling directly at local maxima in a transfer function, allowing for better reproduction of high-frequency features. However, the 1D peak finding technique does not extend to higherdimensional classification. In this work, we develop a new method for peak finding with multidimensional transfer functions, which looks for peaks along the image of the ray. We use piecewise approximations to dynamically sample in transfer function space between world-space samples. As with unidimensional peak finding, this approach is useful for specifying transfer functions with greater precision, and for accurately rendering noisy volume data at lower sampling rates. Multidimensional peak finding produces comparable image quality with order-of-magnitude better performance, and can reproduce features omitted entirely by standard classification. With no precomputation or storage requirements, it is an attractive alternative to preintegration for multidimensional transfer functions.
Natallia Kotava, Aaron Knoll, Mathias Schott, Christoph Garth, Xavier Tricoche, Christoph Kessler, Elaine Cohen, Charles D. Hansen, Michael E. Papka, Hans Hagen
PacificVis9
2012 Efficient data restructuring and aggregation for I/O acceleration in PIDX
abstract
Hierarchical, multiresolution data representations enable interactive analysis and visualization of large-scale simulations. One promising application of these techniques is to store high performance computing simulation output in a hierarchical Z (HZ) ordering that translates data from a Cartesian coordinate scheme to a one-dimensional array ordered by locality at different resolution levels. However, when the dimensions of the simulation data are not an even power of 2, parallel HZ ordering produces sparse memory and network access patterns that inhibit I/O performance. This work presents a new technique for parallel HZ ordering of simulation datasets that restructures simulation data into large (power of 2) blocks to facilitate efficient I/O aggregation. We perform both weak and strong scaling experiments using the S3D combustion application on both Cray-XE6 (65,536 cores) and IBM Blue Gene/P (131,072 cores) platforms. We demonstrate that data can be written in hierarchical, multiresolution format with performance competitive to that of native data-ordering methods.
Sidharth Kumar, Venkatram Vishwanath, Philip H. Carns, Joshua A. Levine, Robert Latham, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci
SC10
2011 Full-resolution interactive CPU volume rendering with coherent BVH traversal
abstract
We present an efficient method for volume rendering by ray casting on the CPU. We employ coherent packet traversal of an implicit bounding volume hierarchy, heuristically pruned using preintegrated transfer functions, to exploit empty or homogeneous space. We also detail SIMD optimizations for volumetric integration, trilinear interpolation, and gradient lighting. The resulting system performs well on low-end and laptop hardware, and can outperform out-of-core GPU methods by orders of magnitude when rendering large volumes without level-of-detail (LOD) on a workstation. We show that, while slower than GPU methods for low-resolution volumes, an optimized CPU renderer does not require LOD to achieve interactive performance on large data sets.
Aaron Knoll, Sebastian Thelen, Ingo Wald, Charles D. Hansen, Hans Hagen, Michael E. Papka
PacificVis6
2011 Experiences using smaash to manage data-intensive simulations
abstract
High performance scientific computer simulations created with such systems as the University of Chicago's FLASH code generate enormous amounts of data that must be captured, cataloged, and analyzed. Unless this is formally done, monitoring such simulations, tracking and reproducing old ones, and analyzing and archiving their output, can be haphazard and idiosyncratic. Smaash, a simulation management and analysis system that has been developed at the University of Chicago and Argonne National Laboratory, seeks to solve some of these problems by offering what approaches a single point of control and analysis, a metadata-base, and a set of tools that automate some of what scientists have been doing by hand.
Randy Hudson, John Norris, Lynn B. Reid, Klaus Weide, George Cal Jordan IV, Michael E. Papka
HPDC6
2011 A new computational paradigm in multiscale simulations: application to brain blood flow
abstract
Interfacing atomistic-based with continuum-based simulation codes is now required in many multiscale physical and biological systems. We present the computational advances that have enabled the first multiscale simulation on 190,740 processors by coupling a high-order (spectral element) Navier-Stokes solver with a stochastic (coarse-grained) Molecular Dynamics solver based on Dissipative Particle Dynamics (DPD). The key contributions are proper interface conditions for overlapped domains, topology-aware communication, SIMDization, multiscale visualization and a new domain partitioning for atomistic solvers. We study blood flow in a patient-specific cerebrovasculature with a brain aneurysm, and analyze the interaction of blood cells with the arterial walls endowed with a glycocalyx causing thrombus formation and eventual aneurysm rupture. The macro-scale dynamics (about 3 billion unknowns) are resolved by NεκTαr - a spectral element solver; the micro-scale flow and cell dynamics within the aneurysm are resolved by an in-house version of DPD-LAMMPS (for an equivalent of about 100 billions molecules).
Leopold Grinberg, Joseph A. Insley, Vitali A. Morozov, Michael E. Papka, George Em Karniadakis, Dmitry A. Fedosov, Kalyan Kumaran
SC4
2011 Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems
abstract
There is growing concern that I/O systems will be hard pressed to satisfy the requirements of future leadership-class machines. Even current machines are found to be I/O bound for some applications. In this paper, we identify existing performance bottlenecks in data movement for I/O on the IBM Blue Gene/P (BG/P) supercomputer currently deployed at several leadership computing facilities. We improve the I/O performance by exploiting the network topology of BG/P for collective I/O, leveraging data semantics of applications and incorporating asynchronous data staging. We demonstrate the efficacy of our approaches for synthetic benchmark experiments and for application-level benchmarks at scale on leadership computing systems.
Venkatram Vishwanath, Mark Hereld, Vitali A. Morozov, Michael E. Papka
SC4
2010 A Web 2.0-Based Scientific Application Framework
abstract
A significant obstacle to building usable, web-based interfaces for computational science in a Grid environment is how to deploy scientific applications on computational resources and expose these applications as web services. To streamline the development of these interfaces, we propose a new application framework that can deliver user-defined scientific workflows as both web services and OpenSocial gadgets. Through this application framework, scientists can focus on defining computational workflows using domain-specific applications and can use the software tools in the framework to quickly generate gadgets for running the applications and visualizing the output from workflow executions. By assembling these domain-specific gadgets and some common gadgets predefined in the framework for workflow management, scientists can easily set up a customized computational workspace to meet their requirements.
Wenjun Wu 0001, Thomas D. Uram, Michael Wilde, Mark Hereld, Michael E. Papka
ICWS5
2010 Accelerating I/O Forwarding in IBM Blue Gene/P Systems
abstract
Current leadership-class machines suffer from a significant imbalance between their computational power and their I/O bandwidth. I/O forwarding is a paradigm that attempts to bridge the increasing performance and scalability gap between the compute and I/O components of leadership-class machines to meet the requirements of data-intensive applications by shipping I/O calls from compute nodes to dedicated I/O nodes. I/O forwarding is a critical component of the I/O subsystem of the IBM Blue Gene/P supercomputer currently deployed at several leadership computing facilities. In this paper, we evaluate the performance of the existing I/O forwarding mechanisms for BG/P and identify the performance bottlenecks in the current design. We augment the I/O forwarding with two approaches: I/O scheduling using a work-queue model and asynchronous data staging. We evaluate the efficacy of our approaches using microbenchmarks and application-level benchmarks on leadership class systems.
Venkatram Vishwanath, Mark Hereld, Kamil Iskra, Dries Kimpe, Vitali A. Morozov, Michael E. Papka, Robert B. Ross, Kazutomo Yoshii
SC6
2007 Enabling community access to TeraGrid visualization resources
abstract
Abstract Visualization is an important part of the data analysis process. Many researchers, however, do not have access to the resources required to do visualization effectively for large datasets. This problem is illustrated through several user scenarios. To remedy this problem, we propose a Visualization Gateway that provides simplified access to such resources to a broad population of users. The current implementation of this gateway is described, including the technology used and the services made available. In particular, a detailed description of a ParaView portlet is included. A proposed design for enabling access to community users is discussed. Technology as well as policy issues that were raised, including security and data management, are covered, as are methods for providing additional services, scaling to include additional resources, and other areas of future development. The paper concludes with a summary of the topics covered. Copyright © 2006 John Wiley & Sons, Ltd.
Justin Binns, Jonathan DiCarlo, Joseph A. Insley, Ti Leggett, Cory Lueninghoener, John-Paul Navarro, Michael E. Papka
Concurr. Comput. Pract. Exp.7
2007 Runtime Visualization of the Human Arterial Tree
abstract
Large-scale simulation codes typically execute for extended periods of time and often on distributed computational resources. Because these simulations can run for hours, or even days, scientists like to get feedback about the state of the computation and the validity of its results as it runs. It is also important that these capabilities be made available with little impact on the performance and stability of the simulation. Visualizing and exploring data in the early stages of the simulation can help scientists identify problems early, potentially avoiding a situation where a simulation runs for several days, only to discover that an error with an input parameter caused both time and resources to be wasted. We describe an application that aids in the monitoring and analysis of a simulation of the human arterial tree. The application provides researchers with high-level feedback about the state of the ongoing simulation and enables them to investigate particular areas of interest in greater detail. The application also offers monitoring information about the amount of data produced and data transfer performance among the various components of the application.
Joseph A. Insley, Michael E. Papka, Suchuan Dong, George Em Karniadakis, Nicholas T. Karonis
IEEE Trans. Vis. Comput. Graph.2
2006 Extending Multicast Communications by Hybrid Overlay Network
abstract
Recently scalable streaming applications have been developed by the new emerging networking technologies. One interesting challenge is how to build an efficient and reliable network topology over heterogeneous network systems. In this paper, we argue that hybrid network overlay, via both IP multicast and overlay mesh multicast, is an important technology for group communications. In this paper, we propose a hybrid network overlay construction with three performance goals: high performance (low end-to-end transmission latency, high bandwidth mesh overlay links), low end-to-end hop count and high reliability. This paper compares the performance penalties from IP multicast, overlay multicast and hybrid overlay multicast via various topology simulations.
Michael E. Papka, Rick L. Stevens
ICC2
2006 Ultra-scale visualization - Workshop on ultra-scale visualization
abstract
The output from the massively parallel scientific simulations is so voluminous and complex that advanced visualization technologies are necessary to interpret the calculated results. Even though visualization technology has progressed significantly in recent years, we are barely capable of visualizing and analyzing terascale data to its full extent, and petascale datasets are on the horizon. This workshop aims at addressing this pressing issue by fostering communication between visualization researchers and practitioners. The workshop attendees will be introduced to the latest and greatest research innovations in large data visualization and also help direct further research direction through an open discussion session.
James P. Ahrens, Hank Childs, John P. Clyne, E. Wes Bethel, Jian Huang 0007, Scott Klasky, Kwan-Liu Ma, Kenneth Moreland, Michael E. Papka, Valerio Pascucci, Han-Wei Shen, Deborah Silver
SC9
2006 CupHolder: A Multi-Person Interactive High-Resolution Workstation
abstract
Demand for high resolution visualization, large pixel real estate collaborative workspaces, and interactive computer interfaces continue to drive researchers to develop new physical portals connecting them to their computational tools, to their data, and to their colleagues around the world. In this paper we describe CupHolder, a high performance workstation designed to support interactive collaboration and research activities. It is configured to enable display of high resolution imagery while enabling a comfortable interactive environment for one to several co-located researchers. Moreover it is driven by a high performance commodity cluster that provides substantial local rendering muscle as well as a high performance interface to Grid-based computational tasks. CupHolder is comprised of commodity components. It is primarily novel because it represents an integration of these components into a new form factor that we believe is a useful precursor and testbed for future integrated workspaces.
Mark Hereld, Michael E. Papka, Justin Binns, Rick L. Stevens
VR2
2006 Simulating and visualizing the human arterial system on the TeraGrid
Suchuan Dong, Joseph A. Insley, Nicholas T. Karonis, Michael E. Papka, Justin Binns, George Em Karniadakis
Future Gener. Comput. Syst.4
2005 An Infrastructure of Network Services for Seamless Integration in Advanced Collaborative Computing Environments
abstract
Advanced collaborative computing environments are one of the most important tools for integrating high-performance computers and computations and for interacting with colleagues around the world. However, heterogeneous characteristics such as network transfer rates, computational abilities, and hierarchical systems make the seamless integration of distributed resources a challenge. In this paper, we argue that advanced collaborative computing environments need an infrastructure of network services to support distributed and quality guaranteed multimedia applications. Accordingly, we propose the design of network services for high-performance collaborative computing. We present a collaborative environment network service infrastructure (CENSI) to embed network services into various systems intelligently and elastically. We also discuss three management modules: a three-party matching module (resources, requests, and network services), a module for performance monitoring and evaluation of group communications, and a module for distribution topology analysis
Ivan R. Judson, Thomas D. Uram, S. Lefvert, Terry Disz, Michael E. Papka, Rick L. Stevens
CLUSTER6
2005 Performance Metrics of IP Multicast Sessions
abstract
Most research on design and implementation of multicast network systems has concentrated mainly on the capacities of network devices, without completely considering the impact from multicast sessions on network systems. This impact heavily depends on the activities during the life span of multicast sessions, such as the session initialization, communication, termination, and membership management, and on the influence within the overall architecture context of multicast sessions, such as routing devices and protocols. Based on the fundamental rules and principles of multicast sessions, this paper defines performance metrics for IP multicast sessions. The paper summarizes the metrics in four categories: latency, quality of service, group characteristics, and link bandwidth consumptions. Also presented are results from our simulations.
Michael E. Papka, Rick L. Stevens
ISM2
2005 Real Time Change Detection and Alerts from Highway Traffic Data
abstract
We developed a testbed containing: real time data from over 830 highway traffic sensors in the Chicago region, data about weather, and text data about events that might affect traffic. The goal was to detect in real time interesting changes in traffic conditions. Given the size and complexity of the data, we choose to build a large number of separate baseline models. We built a separate baseline for each hour in the day, for each day in the week, and for every 2 or 3 traffic sensors, resulting in over 42,000 separate baseline models. We also built a baseline engine to build the necessary baselines automatically. We modified an open source scoring engine to process in real time each new sensor reading, update the appropriate feature vectors, score the updated feature vectors using the baseline models, and send out real time alerts when deviations from the baselines were detected.
Robert L. Grossman, Michal Sabala, Anushka Anand, Steve Eick, Leland Wilkinson, John Chaves, Steve Vejcik, John F. Dillenburg, Peter C. Nelson, Doug Rorem, Javid Alimohideen, Jason Leigh, Michael E. Papka, Rick L. Stevens
SC14
2004 Capability matching of data streams with network services
abstract
Distributed computing middleware needs to support a wide range of resources, such as diverse software components, various hardware devices, and heterogeneous operating systems and architectures. Current technologies are unable to implement a maintenance-free platform to be compatible with such different computing environments. This situation is presenting an increasing challenge as Grid computing becomes more widespread. The infrastructure of network services (CENSA and CENSI) has been proposed to address this challenge. A seamless Grid computing environment, supported by network services, is composed of various streams such as data, video, audio, and text. We define a mathematical model of capability matching for three-party agreements: requests from users, resources, and network services. Based on the mathematical model, we provide a general approach for capability matching. We also present a new language schema for capability description. As an example, we embed the general matchmaker in the architecture of the access Grid. Several tests of accuracy and performance are discussed.
Ivan R. Judson, Thomas D. Uram, Terry Disz, Michael E. Papka, Rick L. Stevens
CCGRID5
2004 Simulation of neocortical epileptiform activity using parallel computing
Wim van Drongelen, Hyong Lee, Mark Hereld, Matthew Cohoon, Frank Elsen, Michael E. Papka, Rick L. Stevens
Neurocomputing7
2003 High-resolution remote rendering of large datasets in a collaborative environment
Nicholas T. Karonis, Michael E. Papka, Justin Binns, John Bresnahan, Joseph A. Insley, Joseph M. Link
Future Gener. Comput. Syst.2
2002 GridMapper: A Tool for Visualizing the Behavior of Large-Scale Distributed Systems
abstract
Grid applications can combine the use of computation, storage, network, and other resources. These resources are often geographically distributed, adding to application complexity and thus the difficulty of understanding application performance. We present GridMapper, a tool for monitoring and visualizing the behavior of such distributed systems. GridMapper builds on basic mechanisms for registering, discovering, and accessing performance information sources, as well as for mapping from domain names to physical locations. The visualization system itself then supports the automatic layout of distributed sets of such sources and animation of their activities. We use a set of examples to illustrate how the system can provide valuable insights into the behavior and performance of a range of different applications.
William E. Allcock, Joseph Bester, John Bresnahan, Ian T. Foster, Jarek Gawor, Joseph A. Insley, Joseph M. Link, Michael E. Papka
HPDC8
1999 Numerical Simulation and Immersive Visualization of Hairpin Vortices
abstract
To better understand the vortex dynamics of coherent structures in turbulent and transitional boundary layers, we consider direct numerical simulation of the interaction between a flat-plate-boundary-layer flow and an isolated hemispherical roughness element. Of principal interest is the evolution of hairpin vortices that form an interlacing pattern in the wake of the hemisphere, lift away from the wall, and are stretched by the shearing action of the boundary layer. Using animations of unsteady three-dimensional representations of this flow, produced by the vtk toolkit and enhanced to operate in a CAVE virtual environment, we identify and study several key features in the evolution of this complex vortex topology not previously observed in other visualization formats. 1 Introduction We present visualization results of numerically generated hairpin vortex formation and evolution in incompressible boundary layer flows. At moderate flow speeds, hairpin vortices provide an exampl...
Henry M. Tufo, Paul F. Fischer, Michael E. Papka, Kristopher J. Blom
SC3
1996 Tools for Distributed Collaborative Environments: A Research Agenda
abstract
Argues that future computing environments will be collaboration-oriented, globally distributed and computation/information-rich. These environments will be accessed via multiple interface devices. As we move from a desktop-centric computing model to a network-centric model, new approaches in the way software and data are handled will need to be developed. In this article, we outline the requirements for enabling the technological infrastructure, describe some first steps that we have taken toward building this infrastructure, and sketch directions for future development.
Ian T. Foster, Michael E. Papka, Rick L. Stevens
HPDC2
1996 UbiWorld: An Environment Integrating Virtual Reality, Supercomputing and Design
abstract
Summary form only given. UbiWorld is a concept that ties together the notion of ubiquitous computing (Ubicomp) with that of using virtual reality for rapid prototyping. The goal is to develop an environment where one can explore Ubicomp-type concepts without having to build real Ubicomp hardware. The basic notion is to extend object models in a virtual world using distributed wide-area heterogeneous computing technology to provide complex networking and processing capabilities to virtual reality objects. Starting with the CAVE/sup TM/ family of display devices, we integrate tools for the construction of 30 objects into the existing library. Then, using these objects as models, we can embed new information technology within them. The plan is then to couple the virtual objects to remote computers via fine-grain heterogenous computing technology to provide Ubicomp behavior and functionality to the modeled objects. We tightly couple the process-defined behavior with the 30 objects and place these objects into rooms, creating a shared virtual world where users can experiment with using the virtual devices. Each object in the world has its behavior controlled by a program running some place on the network. This behavior could be one that in the real object would be provided by a local computer or by a combination of local computer and network connection to remote processors or databases. These "behavior" processes are able to communicate with each other using a shared protocol (UbiWorldcomm). These object also react and are influenced directly by interactions with the virtual world and users.
Michael E. Papka, Rick L. Stevens
HPDC1