Venkatram Vishwanath

dblp:30/6369 · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-7248-6116ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 56 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Queue wait time prediction in high performance computing (HPC) systems
abstract
Abstract High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.
Nwamaka Okafor, Bethany Lusch, Venkatram Vishwanath
J. Supercomput.3
2025 A Deep Probabilistic Framework for Continuous Time Dynamic Graph Generation
abstract
Recent advancements in graph representation learning have shifted attention towards dynamic graphs, which exhibit evolving topologies and features over time. The increased use of such graphs creates a paramount need for generative models suitable for applications such as data augmentation, obfuscation, and anomaly detection. However, there are few generative techniques that handle continuously changing temporal graph data; existing work largely relies on augmenting static graphs with additional temporal information to model dynamic interactions between nodes. In this work, we propose a fundamentally different approach: We instead directly model interactions as a joint probability of an edge forming between two nodes at a given time. This allows us to autoregressively generate new synthetic dynamic graphs in a largely assumption free, scalable, and inductive manner. We formalize this approach as DG-Gen, a generative framework for continuous time dynamic graphs, and demonstrate its effectiveness over five datasets. Our experiments demonstrate that DG-Gen not only generates higher fidelity graphs compared to traditional methods but also significantly advances link prediction tasks.
Ryien Hosseini, Filippo Simini, Venkatram Vishwanath, Henry Hoffmann
AAAI3
2025 Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf Benchmarks
abstract
The significantly increasing use of artificial intelligence (AI) has led to the availability of specialized AI accelerators, aiming to enhance the performance and energy efficiency of AI workloads. In this paper, we conduct an initial study to evaluate the energy requirements of four AI accelerators: Nvidia A100 GPUs, Intel Habana Gaudi Processing Units (HPUs), Graphcore Bow-Pod64 Intelligence Processing Units (IPUs), and GroqRack Language Processing Units (LPUs) using two popular MLPerf benchmarks: BERT-Large and ResNet50. We report the energy requirements for the two benchmarks to achieve a common MLPerfspecified target accuracy. The benchmarks and AI accelerators were chosen based on the following criteria: publicly available tools or libraries from the vendors to monitor power consumption, publicly available optimized models from the vendors, and access to the AI accelerators. Our experimental results indicate that for ResNet50, Intel Gaudi2 HPUs delivered the highest throughput, the lowest energy consumption for both training and inference, and the highest inference energy efficiency, while Graphcore demonstrated the highest training energy efficiency. For BERTLarge pre-training, Intel Gaudi2 outperformed both Nvidia A100 and Graphcore in terms of time, energy consumption, and training energy efficiency. However, for BERT-Large inference, Nvidia A100 achieved the shortest time and lowest energy consumption; Graphcore exhibited the highest throughput in both pre-training and inference, along with the highest inference energy efficiency. We discuss our observations and findings while exploring the associated tradeoffs.
Farah Ferdaus, Xingfu Wu, Valerie Taylor 0001, Zhiling Lan, Sanjif Shanmugavelu, Venkatram Vishwanath, Michael E. Papka
CCGrid6
2025 Machine Learning Driven Auto-Tuning for Non-Uniform All-to-All Collectives
abstract
Non-uniform all-to-all communication patterns present optimization challenges in parallel computing due to their irregular data distribution and dynamic behavior. While MPI_Alltoallv provides the standard interface for such exchanges, achieving optimal performance requires careful selection among multiple implementation variants and tuning of algorithm-specific parameters. This paper presents a data-driven autotuning framework that combines machine learning-based runtime prediction with a lookup-table mechanism for fast configuration selection. The ML model estimates the communication time of each algorithm configuration under a given system setup, allowing the framework to identify the optimal implementation and parameter set based on predicted performance. We validate our approach through comprehensive benchmarking of MPI_Alltoallv and two specialized algorithms across varying process counts, message sizes, and tunable parameters. Applied to a real MPI-based transitive closure application on the Fugaku supercomputer, our framework achieves up to$6.03 \times$reduction in communication time over the vendor implementation, providing a detailed understanding of non-uniform collective communication behavior and a practical framework for automatic performance optimization in HPC applications.
Kunting Qi, Jens Domke, Seydou Ba, Venkatram Vishwanath, Michael E. Papka, Sidharth Kumar
HiPC5
2025 Langvision-Lora-Nas: Neural Architecture Search for Variable Lora Rank In Vision Language Models
abstract
Vision Language Models (VLMs) integrate visual and text modalities to enable multimodal understanding and generation. These models typically combine a Vision Transformer (ViT) as an image encoder and a Large Language Model (LLM) for text generation. LoRA (Low-Rank Adaptation) is an efficient fine-tuning method to adapt pre-trained models to new tasks by introducing low-rank updates to their weights. While LoRA has emerged as a powerful technique for fine-tuning large models by introducing low-rank updates, current implementations assume a fixed rank, potentially limiting flexibility and efficiency across diverse tasks. This paper introduces LangVision-LoRA-NAS, a novel framework that integrates Neural Architecture Search (NAS) with LoRA to optimize VLMs for variable-rank adaptation. Our approach leverages NAS to dynamically search for the optimal LoRA rank configuration tailored to specific multimodal tasks, balancing performance and computational efficiency. Through extensive experiments using the LLaMA-3.2-11B model on several datasets, LangVision-LoRA-NAS demonstrates notable improvement in model performance while reducing fine-tuning costs. Our Base and searched fine-tuned models on LLaMA-3.2-11B-Vision-Instruct can be found here and the code for LangVision-LoRA-NAS can be found here.
Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath
ICIP3
2025 Quality Measures for Dynamic Graph Generative Models
abstract
Deep generative models have recently achieved significant success in modeling graph data, including dynamic graphs, where topology and features evolve over time. However, unlike in vision and natural language domains, evaluating generative models for dynamic graphs is challenging due to the difficulty of visualizing their output, making quantitative metrics essential. In this work, we develop a new quality metric for evaluating generative models of dynamic graphs. Current metrics for dynamic graphs typically involve discretizing the continuous-evolution of graphs into static snapshots and then applying conventional graph similarity measures. This approach has several limitations: (a) it models temporally related events as i.i.d. samples, failing to capture the non-uniform evolution of dynamic graphs; (b) it lacks a unified measure that is sensitive to both features and topology; (c) it fails to provide a scalar metric, requiring multiple metrics without clear superiority; and (d) it requires explicitly instantiating each static snapshot, leading to impractical runtime demands that hinder evaluation at scale. We propose a novel metric based on the Johnson-Lindenstrauss lemma, applying random projections directly to dynamic graph data. This results in an expressive, scalar, and application-agnostic measure of dynamic graph similarity that overcomes the limitations of traditional methods. We also provide a comprehensive empirical evaluation of metrics for continuous-time dynamic graphs, demonstrating the effectiveness of our approach compared to existing methods. Our implementation is available at https://github.com/ryienh/jl-metric.
Ryien Hosseini, Filippo Simini, Venkatram Vishwanath, Rebecca Willett, Henry Hoffmann
ICLR3
2025 Sketch-Augmented Features Improve Learning Long-Range Dependencies in Graph Neural Networks
abstract
Graph Neural Networks learn on graph-structured data by iteratively aggregating local neighborhood information. While this local message passing paradigm imparts a powerful inductive bias and exploits graph sparsity, it also yields three key challenges: (i) oversquashing of long-range information, (ii) oversmoothing of node representations, and (iii) limited expressive power. In this work we inject randomized global embeddings of node features, which we term Sketched Random Features, into standard GNNs, enabling them to efficiently capture long-range dependencies. The embeddings are unique, distance-sensitive, and topology-agnostic---properties which we analytically and empirically show alleviate the aforementioned limitations when injected into GNNs. Experimental results on real-world graph learning tasks confirm that this strategy consistently improves performance over baseline GNNs, offering both a standalone solution and a complementary enhancement to existing techniques such as graph positional encodings.
Ryien Hosseini, Filippo Simini, Venkatram Vishwanath, Rebecca Willett, Henry Hoffmann
NeurIPS3
2025 AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
abstract
Generative machine learning offers new opportunities to better understand complex Earth system dynamics. Recent diffusion-based methods address spectral biases and improve ensemble calibration in weather forecasting compared to deterministic methods, yet have so far proven difficult to scale stably at high resolutions. We introduce AERIS, a 1.3 to 80B parameter pixel-level Swin diffusion transformer to address this gap, and SWiPe, a generalizable technique that composes window parallelism with sequence and pipeline parallelism to shard window-based transformers without added communication cost or increased global batch size. On Aurora (10,080 nodes), AERIS sustains 10.21 ExaFLOPS (mixed precision) and a peak performance of 11.21 ExaFLOPS with 1 × 1 patch size on the 0.25° ERA5 dataset, achieving 95.5% weak scaling efficiency, and 81.6% strong scaling efficiency. AERIS outperforms the IFS ENS and remains stable on seasonal scales to 90 days, highlighting the potential of billion-parameter diffusion models for weather and climate prediction.
Väinö Hatanpää, Eugene Ku, Jason Stock, Murali Emani, Sam Foreman, Chunyong Jung, Sandeep Madireddy, Varuni Sastry 0001, Ray A. O. Sinurat, Huihuo Zheng, Sam Wheeler, Troy Arcomano, Venkatram Vishwanath, Rao Kotamarthi
SC14
2024 A Multi-Level, Multi-Scale Visual Analytics Approach to Assessment of Multifidelity HPC Systems
abstract
The ability to monitor and interpret hardware system events and behaviors is crucial to improving the robustness and reliability of these systems, especially in a supercomputing facility. The growing complexity and scale of these systems demand an increase in monitoring data collected at multiple fidelity levels and varying temporal resolutions. In this work, we aim to build a holistic analytical system that helps make sense of such massive data, mainly the hardware logs, job logs, and environment logs collected from disparate subsystems and components of a supercomputer system. This end-to-end log analysis system, coupled with visual analytics support, allows users to glean and promptly extract supercomputer usage and error patterns at varying temporal and spatial resolutions. We use multi-resolution dynamic mode decomposition (mrDMD), a technique that depicts high-dimensional data as correlated spatial-temporal variations patterns or modes, to extract variation patterns isolated at specified frequencies. Our improvements to the mrDMD algorithm help promptly reveal useful information in the massive environment log dataset, which is then associated with the processed hardware and job log datasets using our visual analytics system. Furthermore, our system can identify the usage and error patterns filtered at user, project, and subcomponent levels. We exemplify the effectiveness of our approach with two use scenarios with the Cray XC40 supercomputer.
Shilpika, Bethany Lusch, Murali Emani, Filippo Simini, Venkatram Vishwanath, Michael E. Papka, Kwan-Liu Ma
CCGrid5
2024 WActiGrad: Structured Pruning for Efficient Finetuning and Inference of Large Language Models on AI Accelerators
Krishna Teja Chitty-Venkata, Varuni Katti Sastry, Murali Emani, Venkatram Vishwanath, Sanjif Shanmugavelu, Sylvia Howland
Euro-Par (2)4
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC29
2024 Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance Predictor
Murali Emani, Xiaodong Yu 0001, Dingwen Tao, Xin He 0054, Pengfei Su 0001, Keren Zhou 0001, Venkatram Vishwanath
USENIX ATC8
2024 Adding topology and memory awareness in data aggregation algorithms
abstract
With the growing gap between computing power and the ability of large-scale systems to ingest data, I/O is becoming the bottleneck for many scientific applications. Improving read and write performance thus becomes decisive, and requires consideration of the complexity of architectures. In this paper, we introduce TAPIOCA, an architecture-aware data aggregation library. TAPIOCA offers an optimized implementation of the two-phase I/O scheme for collective I/O operations, taking advantage of the many levels of memory and storage that populate modern HPC systems, and leveraging network topology. We show that TAPIOCA can significantly improve the I/O bandwidth of synthetic benchmarks and I/O kernels of scientific applications running on leading supercomputers. For example, on HACC-IO, a cosmology code, TAPIOCA improves data writing by a factor of 13 on nearly a third of the target supercomputer.
Francois Tessier, Venkatram Vishwanath, Emmanuel Jeannot
Future Gener. Comput. Syst.2
2023 Asynchronous Decentralized Bayesian Optimization for Large Scale Hyperparameter Optimization
abstract
Bayesian optimization (BO) is a promising approach for hyperparameter optimization of deep neural networks (DNNs), where each model training can take minutes to hours. In BO, a computationally cheap surrogate model is employed to learn the relationship between parameter configurations and their performance such as accuracy. Parallel BO methods often adopt single manager/multiple workers strategies to evaluate multiple hyperparameter configurations simultaneously. Despite significant hyperparameter evaluation time, the overhead in such centralized schemes prevents these methods to scale on a large number of workers. We present an asynchronous-decentralized BO, wherein each worker runs a sequential BO and asynchronously communicates its results through shared storage. We scale our method without loss of computational efficiency with above 95% of worker's utilization to 1,920 parallel workers (full production queue of the Polaris supercomputer) and demonstrate improvement in model accuracy as well as faster convergence on the CANDLE benchmark from the Exascale computing project.
Romain Egele, Isabelle Guyon, Venkatram Vishwanath, Prasanna Balaprakash
e-Science3
2023 TrainBF: High-Performance DNN Training Engine Using BFloat16 on AI Accelerators
Siddhisanket Raskar, Murali Emani, Venkatram Vishwanath
Euro-Par4
2023 A survey of techniques for optimizing transformer inference
Krishna Teja Chitty-Venkata, Sparsh Mittal, Murali Emani, Venkatram Vishwanath, Arun K. Somani
J. Syst. Archit.4
2022 Stimulus: Accelerate Data Management for Scientific AI applications in HPC
abstract
Modern scientific workflows couple simulations with AI-powered analytics by frequently exchanging data to accelerate time-to-science to reduce the complexity of the simulation planes. However, this data exchange is limited in performance and portability due to a lack of support for scientific data formats in AI frameworks. We need a cohesive mechanism to effectively integrate at scale complex scientific data formats such as HDF5, PnetCDF, ADIOS2, GNCF, and Silo into popular AI frameworks such as TensorFlow, PyTorch, and Caffe. To this end, we designed Stimulus, a data management library for ingesting scientific data effectively into the popular AI frameworks. We utilize the StimOps functions along with StimPack abstraction to enable the integration of scientific data formats with any AI framework. The evaluations show that Stimulus outperforms several large-scale applications with different use-cases such as Cosmic Tagger (consuming HDF5 dataset in PyTorch), Distributed FFN (consuming HDF5 dataset in TensorFlow), and CosmoFlow (converting HDF5 into TFRecord and then consuming that in TensorFlow) by 5.3 x, 2.9 x, and 1.9 x respectively with ideal I/O scalability up to 768 GPUs on the Summit supercomputer. Through Stimulus, we can portably extend existing popular AI frameworks to cohesively support any complex scientific data format and efficiently scale the applications on large-scale supercomputers.
Hariharan Devarajan, Antonios Kougkas, Huihuo Zheng, Venkatram Vishwanath, Xian-He Sun
CCGRID4
2022 Toward an In-Depth Analysis of Multifidelity High Performance Computing Systems
abstract
To maintain a robust and reliable supercomputing facility, monitoring it and understanding its hardware system events and behaviors is an essential task. Exascale systems will be increasingly heterogeneous, and the volume of systems data, collected from multiple subsystems and components measured at multiple fidelity levels and temporal resolutions, will continue to grow. In this work, we aim to create an effective solution to analyze diverse and massive datasets gathered from the error logs, job logs, and environment logs of an HPC system, such as a Cray XC40 supercomputer. In this work, we build an end-to-end error log analysis system that analyzes the job logs and gleans insights from their correspondence with hardware error logs and environment logs despite their varying temporal and spatial resolutions. Our machine learning pipeline built in our system is ~92% accurate in predicting the job exit status and does so with sufficient lead time for evasive actions to be taken before the actual failure event occurs.
Shilpika, Bethany Lusch, Murali Emani, Filippo Simini, Venkatram Vishwanath, Michael E. Papka, Kwan-Liu Ma
CCGRID5
2022 HDF5 Cache VOL: Efficient and Scalable Parallel I/O through Caching Data on Node-local Storage
abstract
Modern-era high performance computing (HPC) systems are providing multiple levels of memory and storage layers to bridge the performance gap between fast memory and slow disk-based storage system managed by Lustre or GPFS. Several of the recent HPC systems are equipped with SSD and NVMe-based storage that is attached locally to compute nodes. A few systems are providing an SSD-based “burst buffer” intermediate storage layer that is accessible by all compute nodes as a single file system. Although these hardware layers are intended to reduce the latency gap between memory and disk-based long-term storage, how to utilize them has been left to the users. High-level I/O libraries, such as HDF5 and netCDF, can potentially take advantage of the node-local storage as a cache for reducing I/O latency from capacity storage. However, it is challenging to use node-local storage in parallel I/O especially for a single shared file. In this paper, we present an approach to integrate node-local storage as transparent caching or staging layers in a high-level parallel I/O library without placing the burden of managing these layers on users. We designed this to move data asynchronously between the caching storage layer and a parallel file system to overlap the data movement overhead in performing I/O with compute phases. We implement this approach as an external HDF5 Virtual Object Layer (VOL) connector, named Cache VOL. HDF5 VOL is a layer of abstraction in HDF5 that allows intercepting the public HDF5 application programming interface (API) and performing various optimizations to data movement after the interception. Existing HDF5 applications can use Cache VOL with minimal code modifications. We evaluated the performance of Cache VOL in HPC applications such as VPIC-10, and deep learning applications such as ImageNet and CosmoFlow. We show that using Cache VOL, one can achieve higher observed I/O performance, more scalable and stable I/O compared to direct I/O to the parallel file system, thus achieving faster time-to-solution in scientific simulations. While the caching approach is implemented in HDF5, the methods are applicable in other high-level I/O libraries.
Huihuo Zheng, Venkatram Vishwanath, Quincey Koziol, Houjun Tang, John Ravi, John Mainzer, Surendra Byna
CCGRID2
2022 Efficient Design Space Exploration for Sparse Mixed Precision Neural Architectures
abstract
Pruning and Quantization are two effective Deep Neural Network (DNN) compression methods for efficient inference on various hardware platforms. Pruning refers to removing unimportant weights or nodes, whereas Quantization converts the floating-point parameters to low-bit fixed integer representation. The pruned and low precision models result in smaller and faster inference models on hardware platforms with almost the same accuracy as the unoptimized network. Tensor Cores in Nvidia Ampere 100 (A100) GPU supports (1) 2:4 fine-grained sparse pruning where 2 out of every 4 elements are pruned, and (2) traditional dense multiplication to achieve a good accuracy and performance trade-off. The A100 Tensor Core also takes advantage of 1-bit, 4-bit, and 8-bit multiplication to speed up the inference of a model. Hence, finding the right matrix type (dense or 2:4 sparse) along with the precision for each layer becomes a combinatorial problem. Neural Architecture Search (NAS) can alleviate such problems by automating the architecture design process instead of a brute-force search. In this paper, we propose (i) Mixed Sparse and Precision Search (MSPS), a NAS framework to search for efficient sparse and mixed-precision quantized model within the predefined search space and fixed backbone neural network (Eg. ResNet50), and (ii) Architecture, Sparse and Precision Search (ASPS) to jointly search for kernel size and number of filters, and sparse-precision combination of each layer. We illustrate the effectiveness of our methods targeting A100 Tensor Core on Nvidia GPUs by searching efficient sparse-mixed precision networks on ResNet50 and achieving better accuracy-latency trade-off models compared to the manually designed Uniform Sparse Int8 networks.
Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath, Arun K. Somani
HPDC3
2021 DLIO: A Data-Centric Benchmark for Scientific Deep Learning Applications
abstract
Deep learning has been shown as a successful method for various tasks, and its popularity results in numerous open-source deep learning software tools. Deep learning has been applied to a broad spectrum of scientific domains such as cosmology, particle physics, computer vision, fusion, and astrophysics. Scientists have performed a great deal of work to optimize the computational performance of deep learning frameworks. However, the same cannot be said for I/O performance. As deep learning algorithms rely on big-data volume and variety to effectively train neural networks accurately, I/O is a significant bottleneck on large-scale distributed deep learning training. This study aims to provide a detailed investigation of the I/O behavior of various scientific deep learning workloads running on the Theta supercomputer at Argonne Leadership Computing Facility. In this paper, we present DLIO, a novel representative benchmark suite built based on the I/O profiling of the selected workloads. DLIO can be utilized to accurately emulate the I/O behavior of modern scientific deep learning applications. Using DLIO, application developers and system software solution architects can identify potential I/O bottlenecks in their applications and guide optimizations to boost the I/O performance leading to lower training times by up to 6.7x.
Hariharan Devarajan, Huihuo Zheng, Antonios Kougkas, Xian-He Sun, Venkatram Vishwanath
CCGRID5
2021 AgEBO-tabular: joint neural architecture and hyperparameter search with autotuned data-parallel training for tabular data
abstract
Developing high-performing predictive models for large tabular data sets is a challenging task. Neural architecture search (NAS) is an AutoML approach that generates and evaluates multiple neural networks with different architectures concurrently to automatically discover an high performing model. A key issue in NAS, particularly for large data sets, is the large computation time required to evaluate each generated architecture. While data-parallel training has the potential to address this issue, a straightforward approach can result in significant loss of accuracy. To that end, we develop AgEBO-Tabular, which combines Aging Evolution (AE) to search over neural architectures and asynchronous Bayesian optimization (BO) to search over hyperparameters to adapt data-parallel training. We evaluate the efficacy of our approach on two large predictive modeling tabular data sets from the Exascale Computing Project-CANcer Distributed Learning Environment (ECP-CANDLE).
Romain Egele, Prasanna Balaprakash, Isabelle Guyon, Venkatram Vishwanath, Fangfang Xia, Rick L. Stevens, Zhengying Liu
SC4
2020 SeeSAw: Optimizing Performance of In-Situ Analytics Applications under Power Constraints
abstract
Future supercomputers will need to operate under a power budget. At the same time, in-situ analysis-where a set of analysis tasks are concurrently executed and periodically communicate with a scientific simulation-is expected to be a primary HPC workload to overcome the increasing gap between the performance of the storage system relative to the computational capabilities of these machines. Ongoing research focuses on efficient coupling of simulation and analysis considering memory or I/O constraints, but power poses a new constraint that has not yet been addressed for these workflows. There are two state-of-the-art HPC power management approaches: 1) a power-aware scheme that measures and reallocates power based on observed usage and 2) a time-aware scheme that measures the relative time between communicating software modules and reallocates power based on timing differences. We find that considering only one feedback metric has two major drawbacks: 1) both approaches miss opportunities to improve performance and 2) they often make incorrect decisions when facing the unique requirements of in-situ analysis. We therefore propose SeeSAw-an application-aware power management approach, which uses both time and power feedback to balance a power budget and maximize performance for in-situ analysis workloads. We evaluate SeeSAw using the molecular dynamics simulation LAMMPS with a set of built-in analyses running on the Theta supercomputer on up to 1024 nodes. We find that the strictly power-aware approach slows down LAMMPS as much as ~25%. The strictly time-aware approach shows improvements of up to ~13% and slowdowns as much as ~60%. In contrast, SeeSAw achieves ~4-30% performance improvements.
Ivana Marincic, Venkatram Vishwanath, Henry Hoffmann
IPDPS2
2020 ExaHDF5: Delivering Efficient Parallel I/O on Exascale Computing Systems
Surendra Byna, M. Scot Breitenfeld, Bin Dong 0002, Quincey Koziol, Elena Pourmal, Dana Robinson, Jérome Soumagne, Houjun Tang, Venkatram Vishwanath, Richard Warren
J. Comput. Sci. Technol.9
2019 Scalable reinforcement-learning-based neural architecture search for cancer deep learning research
abstract
Cancer is a complex disease, the understanding and treatment of which are being aided through increases in the volume of collected data and in the scale of deployed computing power. Consequently, there is a growing need for the development of data-driven and, in particular, deep learning methods for various tasks such as cancer diagnosis, detection, prognosis, and prediction. Despite recent successes, however, designing high-performing deep learning models for nonimage and nontext cancer data is a time-consuming, trial-and-error, manual task that requires both cancer domain and deep learning expertise. To that end, we develop a reinforcement-learning-based neural architecture search to automate deep-learning-based predictive model development for a class of representative cancer data. We develop custom building blocks that allow domain experts to incorporate the cancer-data-specific characteristics. We show that our approach discovers deep neural network architectures that have significantly fewer trainable parameters, shorter training time, and accuracy similar to or higher than those of manually designed architectures. We study and demonstrate the scalability of our approach on up to 1,024 Intel Knights Landing nodes of the Theta supercomputer at the Argonne Leadership Computing Facility.
Prasanna Balaprakash, Romain Egele, Misha Salim, Stefan M. Wild, Venkatram Vishwanath, Fangfang Xia, Thomas S. Brettin, Rick L. Stevens
SC5
2018 Toward Scalable and Asynchronous Object-Centric Data Management for HPC
abstract
Emerging high performance computing (HPC) systems are expected to be deployed with an unprecedented level of complexity due to a deep system memory and storage hierarchy. Efficient and scalable methods of data management and movement through this hierarchy is critical for scientific applications using exascale systems. Moving toward new paradigms for scalable I/O in the extreme-scale era, we introduce novel object-centric data abstractions and storage mechanisms that take advantage of the deep storage hierarchy, named Proactive Data Containers (PDC). In this paper, we formulate object-centric PDCs and their mappings in different levels of the storage hierarchy. PDC adopts a client-server architecture with a set of servers managing data movement across storage layers. To demonstrate the effectiveness of the proposed PDC system, we have measured performance of benchmarks and I/O kernels from scientific simulation and analysis applications using PDC programming interface, and compared the results with existing highly tuned I/O libraries. Using asynchronous I/O along with data and metadata optimizations, PDC demonstrates up to 23× speedup over HDF5 and PLFS in writing and reading data from a plasma physics simulation. PDC achieves comparable performance with HDF5 and PLFS in reading and writing data of a single timestep at small scale, and outperforms them at a scale of larger than 10K cores. In contrast to existing storage systems, PDC offers user-space data management with the flexibility to allocate the number of PDC servers depending on the workload.
Houjun Tang, Surendra Byna, Francois Tessier, Bin Dong 0002, Jingqing Mu, Quincey Koziol, Jérome Soumagne, Venkatram Vishwanath, Jialin Liu 0002, Richard Warren
CCGrid9
2018 Optimizing Data Aggregation by Leveraging the Deep Memory Hierarchy on Large-scale Systems
abstract
Effective data aggregation is of paramount importance for data-centric applications in order to improve data movement for I/O or to facilitate complex workflows, such as in-situ analysis, as well as coupling models and data for multi-physics. A key challenge for data aggregation in current and upcoming architectures is the heterogeneity of memory and storage systems (including DRAM, MCDRAM, NVRAM or parallel file system). One has to take advantage of this hierarchy and the characteristics of each tier to achieve improved performance at scale. In this paper, we present a topology and memory-aware data movement library performing data aggregation on large-scale systems. We first detail our hardware abstraction layer to accomplish code and performance portability on various platforms. Next, we present a cost model taking into account the system interconnect and the memory properties to determine an appropriate location for aggregating data. We also describe how we have implemented a data aggregation mechanism through the read algorithm. Finally, we show how we can improve data movement on a visualization cluster and a leadership-class supercomputer up to 16K processes with a benchmark and two typical I/O kernels. Particularly, we demonstrate how our approach can decrease the I/O time of a classic workflow by 26%.
Francois Tessier, Paul Gressier, Venkatram Vishwanath
ICS3
2018 Topology-aware space-shared co-analysis of large-scale molecular dynamics simulations
Preeti Malakar, Todd S. Munson, Christopher Knight 0001, Venkatram Vishwanath, Michael E. Papka
SC4
2017 TAPIOCA: An I/O Library for Optimized Topology-Aware Data Aggregation on Large-Scale Supercomputers
abstract
Reading and writing data efficiently from storage system is necessary for most scientific simulations to achieve good performance at scale. Many software solutions have been developed to decrease the I/O bottleneck. One well-known strategy, in the context of collective I/O operations, is the two-phase I/O scheme. This strategy consists of selecting a subset of processes to aggregate contiguous pieces of data before performing reads/writes. In this paper, we present TAPIOCA, an MPI-based library implementing an efficient topology-aware two-phase I/O algorithm. We show how TAPIOCA can take advantage of double-buffering and one-sided communication to reduce as much as possible the idle time during data aggregation. We also introduce our cost model leading to a topology-aware aggregator placement optimizing the movements of data. We validate our approach at large scale on two leadership-class supercomputers: Mira (IBM BG/Q) and Theta (Cray XC40). We present the results obtained with TAPIOCA on a micro-benchmark and the I/O kernel of a large-scale simulation. On both architectures, we show a substantial improvement of I/O performance compared with the default MPI I/O implementation. On BG/Q+GPFS, for instance, our algorithm leads to a performance improvement by a factor of twelve while on the Cray XC40 system associated with a Lustre filesystem, we achieve an improvement of four.
Francois Tessier, Venkatram Vishwanath, Emmanuel Jeannot
CLUSTER2
2017 Data movement optimizations for independent MPI I/O on the Blue Gene/Q
Preeti Malakar, Venkatram Vishwanath
Parallel Comput.2
2016 Performance analysis, design considerations, and applications of extreme-scale in situ infrastructures
abstract
A key trend facing extreme-scale computational science is the widening gap between computational and I/O rates, and the challenge that follows is how to best gain insight from simulation data when it is increasingly impractical to save it to persistent storage for subsequent visual exploration and analysis. One approach to this challenge is centered around the idea of in situ processing, where visualization and analysis processing is performed while data is still resident in memory. This paper examines several key design and performance issues related to the idea of in situ processing at extreme scale on modern platforms: scalability, overhead, performance measurement and analysis, comparison and contrast with a traditional post hoc approach, and interfacing with simulation codes. We illustrate these principles in practice with studies, conducted on large-scale HPC platforms, that include a miniapplication and multiple science application codes, one of which demonstrates in situ methods in use at greater than 1M-way concurrency.
Utkarsh Ayachit, Andrew C. Bauer, Earl P. N. Duque, Greg Eisenhauer, Nicola J. Ferrier, Junmin Gu, Kenneth E. Jansen, Burlen Loring, Zarija Lukic, Suresh Menon, Dmitriy Morozov, Patrick O'Leary, Reetesh Ranjan, Michel E. Rasquin, Christopher P. Stone, Venkatram Vishwanath, Gunther H. Weber, Brad Whitlock, Matthew Wolf, Kesheng Wu, E. Wes Bethel
SC16
2016 Optimal execution of co-analysis for large-scale molecular dynamics simulations
abstract
The analysis of scientific simulation data enables scientists to derive insights from their simulations. This analysis of the simulation output can be performed at the same execution site as the simulation using the same resources or can be done at a different site. The optimal output frequency is challenging to decide and is often chosen empirically. We propose a mathematical formulation for choosing the optimal frequency of data transfer for analysis and the feasibility of performing the analysis, under the given resource constraints such as network bandwidth, disk space, available memory, and computation time. We propose formulations for two cases of co-analysis - local and remote. We consider various analyses features such as computation time, input data and memory requirement, importance of the analysis and minimum frequency required for performing the analysis. We demonstrate the effectiveness of our approach using molecular dynamics applications on the Mira and Edison supercomputers.
Preeti Malakar, Venkatram Vishwanath, Christopher Knight 0001, Todd S. Munson, Michael E. Papka
SC2
2016 A data driven scheduling approach for power management on HPC systems
abstract
Modern schedulers running on HPC systems traditionally consider the number of resources and the time requested for each job that is to be executed when making scheduling decisions. Until recently this has been sufficient, however as systems get larger, other metrics like power consumption become necessary to ensure system stability. In this paper, we propose a data driven scheduling approach for controlling the power consumption of the entire system under any user defined budget. Here, “data driven” means that our approach actively observes, analyzes, and assesses power behaviors of the system and user jobs to guide scheduling decisions for power management. This design is based on the key observation that HPC jobs have distinct power profiles. Our work contains an empirical analysis of workload power characteristics on a production system, dynamic learner to estimate the job power profile for scheduling, and an online power-aware scheduler for managing the overall system power. Using real workload traces, we demonstrate that our design effectively controls system power consumption while minimizing the impact on system utilization.
Sean Wallace, Xu Yang 0009, Venkatram Vishwanath, William E. Allcock, Susan Coghlan, Michael E. Papka, Zhiling Lan
SC3
2016 In Situ Methods, Infrastructures, and Applications on High Performance Computing Platforms
abstract
Abstract The considerable interest in the high performance computing (HPC) community regarding analyzing and visualization data without first writing to disk, i. e., in situ processing, is due to several factors. First is an I/O cost savings, where data is analyzed/visualized while being generated, without first storing to a filesystem. Second is the potential for increased accuracy, where fine temporal sampling of transient analysis might expose some complex behavior missed in coarse temporal sampling. Third is the ability to use all available resources, CPU's and accelerators, in the computation of analysis products. This STAR paper brings together researchers, developers and practitioners using in situ methods in extreme‐scale HPC with the goal to present existing methods, infrastructures, and a range of computational science and engineering applications using in situ analysis and visualization.
Andrew C. Bauer, Hasan Abbasi, James P. Ahrens, Hank Childs, Berk Geveci, Scott Klasky, Kenneth Moreland, Patrick O'Leary, Venkatram Vishwanath, Brad Whitlock, E. Wes Bethel
Comput. Graph. Forum9
2016 Workflow performance improvement using model-based scheduling over multiple clusters and clouds
Ketan Maheshwari, Eun-Sung Jung, Jiayuan Meng, Vitali A. Morozov, Venkatram Vishwanath, Rajkumar Kettimuthu
Future Gener. Comput. Syst.5
2016 Improving sparse data movement performance using multiple paths on the Blue Gene/Q supercomputer
Huy Bui, Eun-Sung Jung, Venkatram Vishwanath, Andrew E. Johnson 0001, Jason Leigh, Michael E. Papka
Parallel Comput.3
2016 Application power profiling on IBM Blue Gene/Q
Sean Wallace, Zhou Zhou 0006, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka
Parallel Comput.3
2015 TECA: Petascale Pattern Recognition for Climate Science
Prabhat, Surendra Byna, Venkatram Vishwanath, Eli Dart, Michael F. Wehner, William D. Collins
CAIP (2)3
2015 Multipath Load Balancing for M × N Communication Patterns on the Blue Gene/Q Supercomputer Interconnection Network
abstract
Achievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually have a wide variation from application to application and even from phase to phase in an application, routing algorithms have a limited variation and usually are optimized for typical communication patterns. This results in high networking performance for favored communication patterns but low networking performance for others. In this paper we present approaches for improving networking performance by rebalancing load on physical links on the Blue Gene Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 30% higher throughput on experiment with data and patterns from a real application. The improvement can be up to several times higher throughput than default MPI_Alltoallv used in the Blue Gene Q supercomputer for certain communication patterns.
Huy Bui, Robert L. Jacob, Preeti Malakar, Venkatram Vishwanath, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh
CLUSTER4
2015 Comparison of Vendor Supplied Environmental Data Collection Mechanisms
abstract
The high performance computing landscape is filled with diverse hardware components. A large part of understanding how these components compare to others is by looking at the various environmental aspects of these devices such as power consumption, temperature, etc. Thankfully, vendors of these various pieces of hardware have supported this by providing mechanisms to obtain this data. However, differences not only in the way this data is obtained but also the data which is provided is common between products. In this paper, we take a comprehensive look at the data which is available for the most common pieces of today's HPC landscape, as well as how this data is obtained and how accurate it is. Having surveyed these components, we compare and contrast them noting key differences as well as providing insight into what features future components should have.
Sean Wallace, Venkatram Vishwanath, Susan Coghlan, Zhiling Lan, Michael E. Papka
CLUSTER2
2015 Improving Communication Throughput by Multipath Load Balancing on Blue Gene/Q
abstract
Achievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications, the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually vary from application to application and even from phase to phase in an application, routing algorithms have limited variation and usually are optimized for typical communication patterns. This results in high networking performance for some communication patterns. In this paper we present approaches for improving communication performance by using multiple paths and re-balancing load on physical links on the Blue Gene/Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 43 -- 67% higher throughput on average from 91 experiments, and can achieve higher throughput than default MPI_Alltoallv used for certain communication patterns.
Huy Bui, Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Eun-Sung Jung, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh
HiPC3
2015 Optimal scheduling of in-situ analysis for large-scale scientific simulations
abstract
Today's leadership computing facilities have enabled the execution of transformative simulations at unprecedented scales. However, analyzing the huge amount of output from these simulations remains a challenge. Most analyses of this output is performed in post-processing mode at the end of the simulation. The time to read the output for the analysis can be significantly high due to poor I/O bandwidth, which increases the end-to-end simulation-analysis time. Simulation-time analysis can reduce this end-to-end time. In this work, we present the scheduling of in-situ analysis as a numerical optimization problem to maximize the number of online analyses subject to resource constraints such as I/O bandwidth, network bandwidth, rate of computation and available memory. We demonstrate the effectiveness of our approach through two application case studies on the IBM Blue Gene/Q system.
Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Christopher Knight 0001, Mark Hereld, Sven Leyffer, Michael E. Papka
SC2
2015 Cluster-to-cluster data transfer with data compression over wide-area networks
Eun-Sung Jung, Rajkumar Kettimuthu, Venkatram Vishwanath
J. Parallel Distributed Comput.3
2014 Improving Multisite Workflow Performance Using Model-Based Scheduling
abstract
Workflows play an important role in expressing and executing scientific applications. In recent years, a variety of computational sites and resources have emerged, and users often have access to multiple resources that are geographically distributed. These computational sites are heterogeneous in nature and performance of different tasks in a workflow varies from one site to another. Additionally, users typically have a limited resource allocation at each site. In such cases, judicious scheduling strategy is required in order to map tasks in the workflow to resources so that the workload is balanced among sites and the overhead is minimized in data transfer. Most existing systems either run the entire workflow in a single site or use naive approaches to distribute the tasks across sites or leave it to the user to optimize the allocation of tasks to distributed resources. This results in a significant loss in productivity for a scientist. In this paper, we propose a multi-site workflow scheduling technique that uses performance models to predict the execution time on different resources and dynamic probes to identify the achievable network throughput between sites. We evaluate our approach using real world applications in a distributed environment using the Swift distributed execution framework and show that our approach improves the execution time by up to 60% compared to the default schedule.
Ketan Maheshwari, Eun-Sung Jung, Jiayuan Meng, Venkatram Vishwanath, Rajkumar Kettimuthu
ICPP4
2014 Scalable Parallel I/O on a Blue Gene/Q Supercomputer Using Compression, Topology-Aware Data Aggregation, and Subfiling
abstract
In this paper, we propose an approach to improving the I/O performance of an IBM Blue Gene/Q supercomputing system using a novel framework that can be integrated into high performance applications. We take advantage of the system's tremendous computing resources and high interconnection bandwidth among compute nodes to efficiently exploit I/O bandwidth. This approach focuses on lossless data compression, topology-aware data movement, and subfiling. The efficacy of this solution is demonstrated using microbenchmarks and an application-level benchmark.
Huy Bui, Hal Finkel, Venkatram Vishwanath, Salman Habib 0002, Katrin Heitmann, Jason Leigh, Michael E. Papka, Kevin Harms
PDP3
2014 Efficient I/O and Storage of Adaptive-Resolution Data
abstract
We present an efficient, flexible, adaptive-resolution I/O framework that is suitable for both uniform and Adaptive Mesh Refinement (AMR) simulations. In an AMR setting, current solutions typically represent each resolution level as an independent grid which often results in inefficient storage and performance. Our technique coalesces domain data into a unified, multiresolution representation with fast, spatially aggregated I/O. Furthermore, our framework easily extends to importance-driven storage of uniform grids, for example, by storing regions of interest at full resolution and nonessential regions at lower resolution for visualization or analysis. Our framework, which is an extension of the PIDX framework, achieves state of the art disk usage and I/O performance regardless of resolution of the data, regions of interest, and the number of processes that generated the data. We demonstrate the scalability and efficiency of our framework using the Uintah and S3D large-scale combustion codes on the Mira and Edison supercomputers.
Sidharth Kumar, John Edwards 0002, Peer-Timo Bremer, Aaron Knoll, Cameron Christensen, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Valerio Pascucci
SC6
2013 Model-driven multisite workflow scheduling
abstract
Workflows continue to play an important role in expressing and deploying scientific applications. In recent years, a wide variety of computational sites have emerged with shared access to users. A user may not be able to complete a complex workflow at a single site. It is thus beneficial to run different tasks of a workflow on different sites. For such cases, judicious scheduling strategy is required in order to map tasks in the workflow to resources at multiple sites so that the workload is balanced among sites and the overhead is minimized in data transfer. The key challenge is that the data transfer rate among sites varies based on the network capacity and load. We propose a workflow scheduling technique that tackles the multi-site task distribution challenge by using data movement performance modeling. We applied this technique to schedule an earth observation science workflow over three sites. Executed via the Swift parallel scripting paradigm, we augmented its default schedule and improved the time-to-completion by up to 52%.
Ketan Maheshwari, Eun-Sung Jung, Jiayuan Meng, Venkatram Vishwanath, Rajkumar Kettimuthu
CLUSTER4
2013 Application power profiling on IBM Blue Gene/Q
abstract
The power consumption of state of the art supercomputers, because of their complexity and unpredictable workloads, is extremely difficult to estimate. Accurate and precise results, as are now possible with the latest generation of supercomputers, are therefore a welcome addition to the landscape. Only recently have end users been afforded the ability to access the power consumption of their applications. However, just because it's possible for end users to obtain this data does not mean it's a trivial task. This emergence of new data is therefore not only understudied, but also not fully understood. In this paper, we provide detailed power consumption analysis of microbenchmarks running on Argonne's latest generation of IBM Blue Gene supercomputers, Mira, a Blue Gene/Q system. The analysis is done utilizing our power monitoring library, MonEQ, built on the IBM provided Environmental Monitoring (EMON) API. We describe the importance of sub-second polling of various power domains and the implications they present. To this end, previously well understood applications will now have new facets of potential analysis.
Sean Wallace, Venkatram Vishwanath, Susan Coghlan, John R. Tramm, Zhiling Lan, Michael E. Papka
CLUSTER2
2013 A Generic High-Performance Method for Deinterleaving Scientific Data
Eric R. Schendel, Steve Harenberg, Houjun Tang, Venkatram Vishwanath, Michael E. Papka, Nagiza F. Samatova
Euro-Par4
2013 Scalable in situ scientific data encoding for analytical query processing
Sriram Lakshminarasimhan, David A. Boyuka II, Saurabh V. Pendse, Xiaocheng Zou, John Jenkins, Venkatram Vishwanath, Michael E. Papka, Nagiza F. Samatova
HPDC6
2013 Early Experience on the Blue Gene/Q Supercomputing System
abstract
The Argonne Leadership Computing Facility (ALCF) is home to Mira, a 10 PF Blue Gene/Q (BG/Q) system. The BG/Q system is the third generation in Blue Gene architecture from IBM and like its predecessors combines system-onchip technology with a proprietary interconnect (5-D torus). Each compute node has 16 augmented PowerPC A2 processor cores with support for simultaneous multithreading, 4-wide double precision SIMD, and different data prefetching mechanisms. Mira offers several new opportunities for tuning and scaling scientific applications. This paper discusses our early experience with a subset of micro-benchmarks, MPI benchmarks, and a variety of science and engineering applications running at ALCF. Both performance and power are studied and results on BG/Q is compared with its predecessor BG/P. Several lessons gleaned from tuning applications on the BG/Q architecture for better performance and scalability are shared.
Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Jiayuan Meng, Michael E. Papka
IPDPS3
2013 Characterization and modeling of PIDX parallel I/O for performance optimization
abstract
Parallel I/O library performance can vary greatly in response to user-tunable parameter values such as aggregator count, file count, and aggregation strategy. Unfortunately, manual selection of these values is time consuming and dependent on characteristics of the target machine, the underlying file system, and the dataset itself. Some characteristics, such as the amount of memory per core, can also impose hard constraints on the range of viable parameter values. In this work we address these problems by using machine learning techniques to model the performance of the PIDX parallel I/O library and select appropriate tunable parameter values. We characterize both the network and I/O phases of PIDX on a Cray XE6 as well as an IBM Blue Gene/P system. We use the results of this study to develop a machine learning model for parameter space exploration and performance prediction.
Sidharth Kumar, Avishek Saha, Venkatram Vishwanath, Philip H. Carns, John A. Schmidt, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert Latham, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci
SC3
2013 Multi-domain job coscheduling for leadership computing systems
Wei Tang 0001, Narayan Desai, Venkatram Vishwanath, Daniel Buettner, Zhiling Lan
J. Supercomput.3
2012 Evaluating Power-Monitoring Capabilities on IBM Blue Gene/P and Blue Gene/Q
abstract
Power consumption is becoming a critical factor as we continue our quest toward exascale computing. Yet, actual power utilization of a complete system is an insufficiently studied research area. Estimating the power consumption of a large scale system is a nontrivial task because a large number of components are involved and because power requirements are affected by the (unpredictable) workloads. Clearly needed is a power-monitoring infrastructure that can provide timely and accurate feedback to system developers and application writers so that they can optimize the use of this precious resource. Many existing large-scale installations do feature power-monitoring sensors, however, those are part of environmental- and health monitoring sub systems and were not designed with application level power consumption measurements in mind. In this paper, we evaluate the existing power monitoring of IBM Blue Gene systems, with the goal of understanding what capabilities are available and how they fare with respect to spatial and temporal resolution, accuracy, latency, and other characteristics. We find that with a careful choice of dedicated micro benchmarks, we can obtain meaningful power consumption data even on Blue Gene/P, where the interval between available data points is measured in minutes. We next evaluate the monitoring subsystem on Blue Gene/Q, and are able to study the power characteristics of FPU and memory subsystems of Blue Gene/Q. We find the monitoring subsystem capable of providing second-scale resolution of power data conveniently separated between node components with seven seconds latency. This represents a significant improvement in power monitoring infrastructure, and hope future systems will enable real-time power measurement in order to better understand application behavior at a finer granularity.
Kazutomo Yoshii, Kamil Iskra, Rinku Gupta, Pete Beckman, Venkatram Vishwanath, Chenjie Yu, Susan Coghlan
CLUSTER5
2012 Efficient data restructuring and aggregation for I/O acceleration in PIDX
abstract
Hierarchical, multiresolution data representations enable interactive analysis and visualization of large-scale simulations. One promising application of these techniques is to store high performance computing simulation output in a hierarchical Z (HZ) ordering that translates data from a Cartesian coordinate scheme to a one-dimensional array ordered by locality at different resolution levels. However, when the dimensions of the simulation data are not an even power of 2, parallel HZ ordering produces sparse memory and network access patterns that inhibit I/O performance. This work presents a new technique for parallel HZ ordering of simulation datasets that restructures simulation data into large (power of 2) blocks to facilitate efficient I/O aggregation. We perform both weak and strong scaling experiments using the S3D combustion application on both Cray-XE6 (65,536 cores) and IBM Blue Gene/P (131,072 cores) platforms. We demonstrate that data can be written in hierarchical, multiresolution format with performance competitive to that of native data-ordering methods.
Sidharth Kumar, Venkatram Vishwanath, Philip H. Carns, Joshua A. Levine, Robert Latham, Giorgio Scorzelli, Hemanth Kolla, Ray W. Grout, Robert B. Ross, Michael E. Papka, Jacqueline Chen, Valerio Pascucci
SC2
2012 Dataflow-driven GPU performance projection for multi-kernel transformations
abstract
Applications often have a sequence of parallel operations to be offloaded to graphics processors; each operation can become an individual GPU kernel. Developers typically explore a variety of transformations for each kernel. Furthermore, it is well known that efficient data management is critical in achieving high GPU performance and that "fusing" multiple kernels into one may greatly improve data locality. Doing so, however, requires transformations across multiple, potentially nested, parallel loops; at the same time, the original code semantics and data dependency must be preserved. Since each kernel may have distinct data access patterns, their combined dataflow can be nontrivial. As a result, the complexity of multi-kernel transformations often leads to significant effort with no guarantee of performance benefits. This paper proposes a dataflow-driven analytical framework to project GPU performance for a sequence of parallel operations. Users need only provide CPU code skeletons for a sequence of parallel loops. The framework can then automatically identify opportunities for multi-kernel transformations and data management. It is also able to project the overall performance without implementing GPU code or using physical hardware.
Jiayuan Meng, Vitali A. Morozov, Venkatram Vishwanath, Kalyan Kumaran
SC3
2011 PIDX: Efficient Parallel I/O for Multi-resolution Multi-dimensional Scientific Datasets
abstract
The IDX data format provides efficient, cache oblivious, and progressive access to large-scale scientific datasets by storing the data in a hierarchical Z (HZ) order. Data stored in IDX format can be visualized in an interactive environment allowing for meaningful explorations with minimal resources. This technology enables real-time, interactive visualization and analysis of large datasets on a variety of systems ranging from desktops and laptop computers to portable devices such as iPhones/iPads and over the web. While the existing ViSUS API for writing IDX data is serial, there are obvious advantages of applying the IDX format to the output of large scale scientific simulations. We have therefore developed PIDX - a parallel API for writing data in an IDX format. With PIDX it is now possible to generate IDX datasets directly from large scale scientific simulations with the added advantage of real-time monitoring and visualization of the generated data. In this paper, we provide an overview of the IDX file format and how it is generated using PIDX. We then present a data model description and a novel aggregation strategy to enhance the scalability of the PIDX library. The S3D combustion application is used as an example to demonstrate the efficacy of PIDX for a real-world scientific simulation. S3D is used for fundamental studies of turbulent combustion requiring exceptionally high fidelity simulations. PIDX achieves up to 18 GiB/s I/O throughput at 8,192 processes for S3D to write data out in the IDX format. This allows for interactive analysis and visualization of S3D data, thus, enabling in situ analysis of S3D simulation.
Sidharth Kumar, Venkatram Vishwanath, Philip H. Carns, Brian Summa, Giorgio Scorzelli, Valerio Pascucci, Robert B. Ross, Jacqueline Chen, Hemanth Kolla, Ray W. Grout
CLUSTER2
2011 GROPHECY: GPU performance projection from CPU code skeletons
abstract
We propose GROPHECY, a GPU performance projection framework that can estimate the performance benefit of GPU acceleration without actual GPU programming or hardware. Users need only to skeletonize pieces of CPU code that are targets for GPU acceleration. Code skeletons are automatically transformed in various ways to mimic tuned GPU codes with characteristics resembling real implementations. The synthesized characteristics are used by an existing analytical model to project GPU performance. The cost and benefit of GPU development can then be estimated according to the transformed code skeleton that yields the best projected performance. With GROPHECY, users can leap toward GPU acceleration only when the cost-benefit makes sense. The framework is validated using kernel benchmarks and data-parallel codes in legacy scientific applications. The measured performance of manually tuned codes deviates from the projected performance by 17% in geometric mean.
Jiayuan Meng, Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Thomas D. Uram
SC4
2011 Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems
abstract
There is growing concern that I/O systems will be hard pressed to satisfy the requirements of future leadership-class machines. Even current machines are found to be I/O bound for some applications. In this paper, we identify existing performance bottlenecks in data movement for I/O on the IBM Blue Gene/P (BG/P) supercomputer currently deployed at several leadership computing facilities. We improve the I/O performance by exploiting the network topology of BG/P for collective I/O, leveraging data semantics of applications and incorporating asynchronous data staging. We demonstrate the efficacy of our approaches for synthetic benchmark experiments and for application-level benchmarks at scale on leadership computing systems.
Venkatram Vishwanath, Mark Hereld, Vitali A. Morozov, Michael E. Papka
SC1
2010 Multi-application inter-tile synchronization on ultra-high-resolution display walls
abstract
Ultra-high-resolution tiled-display walls are typically driven by a cluster of computers. Each computer may drive one or more displays. Synchronization between the computers is necessary to ensure that animated imagery displayed on the wall appears seamless. Most tiled-display middleware systems are designed around the assumption that only a single application instance is running in the tiled display at a time. Therefore synchronization can be achieved with a simple solution such as a networked barrier. When a tiled display has to support multiple applications at the same time, however, the simple networked barrier approach does not scale. In this paper we propose and experimentally validate two synchronization algorithms to achieve low-latency, intertile synchronization for multiple applications with independently varying frame rates. The two-phase algorithm is more generally applicable to various highresolution tiled display systems. The one-phase algorithm provides superior results but requires support for the Network Time Protocol and is more CPU-intensive.
Sungwon Nam, Sachin Deshpande, Venkatram Vishwanath, Byungil Jeong, Luc Renambot, Jason Leigh
MMSys3
2010 Accelerating I/O Forwarding in IBM Blue Gene/P Systems
abstract
Current leadership-class machines suffer from a significant imbalance between their computational power and their I/O bandwidth. I/O forwarding is a paradigm that attempts to bridge the increasing performance and scalability gap between the compute and I/O components of leadership-class machines to meet the requirements of data-intensive applications by shipping I/O calls from compute nodes to dedicated I/O nodes. I/O forwarding is a critical component of the I/O subsystem of the IBM Blue Gene/P supercomputer currently deployed at several leadership computing facilities. In this paper, we evaluate the performance of the existing I/O forwarding mechanisms for BG/P and identify the performance bottlenecks in the current design. We augment the I/O forwarding with two approaches: I/O scheduling using a work-queue model and asynchronous data staging. We evaluate the efficacy of our approaches using microbenchmarks and application-level benchmarks on leadership class systems.
Venkatram Vishwanath, Mark Hereld, Kamil Iskra, Dries Kimpe, Vitali A. Morozov, Michael E. Papka, Robert B. Ross, Kazutomo Yoshii
SC1
2009 The OptIPortal, a scalable visualization, storage, and computing interface device for the OptiPuter
Thomas A. DeFanti, Jason Leigh, Luc Renambot, Byungil Jeong, Alan Verlo, Lance Long, Maxine D. Brown, Dan Sandin, Venkatram Vishwanath, Mason J. Katz, Philip M. Papadopoulos, Joseph P. Keefe, Gregory R. Hidley, Gregory Dawe, Ian Kaufman, Bryan Glogowski, Kai-Uwe Doerr, Rajvikram Singh, Javier Girado
Future Gener. Comput. Syst.9
2009 Accelerating tropical cyclone analysis using LambdaRAM, a distributed data cache over wide-area ultra-fast networks
Venkatram Vishwanath, Robert Burns, Jason Leigh, Michael S. Seablom
Future Gener. Comput. Syst.1
2008 The Rails Toolkit - Enabling End-System Topology-Aware High End Computing
abstract
We present a novel rails approach so that future e-Science applications can effectively exploit future system architectures, including multi-core and many-core architectures, multiple network cards, multiple graphical processing units and hybrid hierarchical memory architectures. We define "rail" as the co-scheduling of two or more of these resources. This approach enables creation of parallel multi-rails through every aspect of an end system: from processing on the multi- and many cores, to generation of multiple data flows and streaming over multi-lane network interface card (NIC) connected via a parallel interconnect. We describe a novel open-source multi-rail toolkit and the evaluation of end-system parameters that impact the efficiency of such multi-rail systems, including Interrupt, Memory, Thread and Core Affinities -- key properties for achieving scalable performance.
Venkatram Vishwanath, Jason Leigh, Sungwon Nam, Luc Renambot, Takashi Shimizu, Hirokazu Takahashi, Makoto Takizawa 0004, Osamu Kamatani
eScience1
2008 Specification and Verification of LambdaRAM: A Wide-area Distributed Cache for High Performance Computing
abstract
LambdaRAM is a high-performance, multidimensional, wide-area, distributed cache that takes advantage of massively available memory from multiple clusters interconnected by ultra high-speed networking to provide data-intensive scientific applications with rapid access to both local and remote data without suffering the latency bottlenecks often associated with large storage systems and wide-area data access. LambdaRAM has been demonstrated to yield significant performance speed-ups for geophysical and Bioscience applications accessing extremely large datasets. Currently, LambdaRAM is being integrated by NASA for the modelling, analysis and prediction (MAP) program applications to study tropical cyclones. Formal verification o/LambdaRAM is important to NASA to ensure that LambdaRAM operates reliably in real-time mission critical deployments. We present our preliminary steps towards full formal verification of LambdaRAM. We first give an abstract description of the system and then verify several of its properties. Most of the proofs are accomplished by automatic techniques, while some require deductive steps.
Venkatram Vishwanath, Lenore D. Zuck, Jason Leigh
MEMOCODE1
2006 LambdaBridge: A Scalable Architecture for Future Generation Terabit Applications
abstract
LambdaGrid applications as typified by data-intensive collaborative visualization are likely to be the first users of terabit-level networking. This paper features a main enabler of collaborative visualization over LambdaGrid, the Scalable Adaptive Graphics Environment (SAGE) in particular, and anticipates the future needs of LambdaGrid applications in general. We present a scalable architecture called LambdaBridge to 'bridge' LambdaGrid applications with the terabit optical core networks. LambdaBridge will provision and control predictable- performance networks for end-systems (i.e. Grid clusters) using on-demand lambda/VLAN provisioning and end-system traffic shaping. The chief contribution of LambdaBridge is a deep understanding of how to synergistically bridge provisionable networks, end-systems, and future generation distributed terabit applications. This paper also introduces Scalable Visualcasting - a multicasting service for LambdaGrid that incorporates IP multicasting and optical multicasting.
Venkatram Vishwanath, Byungil Jeong, Ratko Jagodic, Eric He, Luc Renambot, Andrew E. Johnson 0001, Jason Leigh
BROADNETS2
2006 AR-PIN/PDC: Flexible Advance Reservation of Intradomain and Interdomain Lightpaths
abstract
A collection of Grid computing resources interconnected by an application-configurable network of lightpaths is called a LambdaGrid. It provides data-intensive applications with the needed deterministic network bandwidth to transport data between grid instruments, high-performance storage systems, compute clusters and visualization systems, that is often needed for real-time interactive scientific exploration. Advance reservation is needed to guarantee the availability of network resources. Flexible scheduling affords users greater convenience while also improving resource utilization and acceptance rate. In this paper we propose a unified flexible advance reservation model called FARM and apply this model to the cross-domain Routing and Wavelength Assignment problem. We will present the architecture and implementation of a coordinated Interdomain and Intradomain optical control plane, which is capable of flexible advance reservations. Our simulation results show that by just relaxing the reservation constraint and providing flexibility on the starting time, the network can carry 39% more load, and resource utilization can improve by 41%.
Eric He, Venkatram Vishwanath, Jason Leigh
GLOBECOM3
2006 The first functional demonstration of optical virtual concatenation as a technique for achieving Terabit networking
Akira Hirano, Luc Renambot, Byungil Jeong, Jason Leigh, Alan Verlo, Venkatram Vishwanath, Rajvikram Singh, Julieta Aguilera, Andrew E. Johnson 0001, Thomas A. DeFanti, Lance Long, Nicholas Schwarz, Maxine D. Brown, Naohide Nagatsu, Yukio Tsukishima, Masahito Tomizawa, Yutaka Miyamoto, Masahiko Jinno, Yoshihiro Takigawa, Osamu Ishida
Future Gener. Comput. Syst.6
2006 The global lambda visualization facility: An international ultra-high-definition wide-area visualization collaboratory
Jason Leigh, Luc Renambot, Andrew E. Johnson 0001, Byungil Jeong, Ratko Jagodic, Nicholas Schwarz, Dmitry Svistula, Rajvikram Singh, Julieta Aguilera, Venkatram Vishwanath, Brenda Lopez, Dan Sandin, Tom Peterka, Javier Girado, Robert Kooima, Jinghua Ge, Lance Long, Alan Verlo, Thomas A. DeFanti, Maxine D. Brown, Donna J. Cox, Robert Patterson, Patrick Dorn, Paul Wefel, Stuart Levy, Jonas Talandis, Joe Reitzer, Tom Prudhomme, Tom Coffin, Paul Wielinga, Bram Stolk, Gee Bum Koo, Jaeyoun Kim, Sangwoo Han, Jongwon Kim 0001, Brian Corrie II, Todd Zimmerman, Pierre Boulanger, Manuel Garcia
Future Gener. Comput. Syst.11
2004 JuxtaView - a tool for interactive visualization of large imagery on scalable tiled displays
abstract
JuxtaView is a cluster-based application for viewing ultra-high-resolution images on scalable tiled displays. We present in JuxtaView, a new parallel computing and distributed memory approach for out-of-core montage visualization, using LambdaRAM, a software-based network-level cache system. The ultimate goal of JuxtaView is to enable a user to interactively roam through potentially terabytes of distributed, spatially referenced image data such as those from electron microscopes, satellites and aerial photographs. In working towards this goal, we describe our first prototype implemented over a local area network, where the image is distributed using LambdaRAM, on the memory of all nodes of a PC cluster driving a tiled display wall. Aggressive prefetching schemes employed by LambdaRAM help to reduce latency involved in remote memory access. We compare LambdaRAM with a more traditional memory-mapped file approach for out-of-core visualization.
Naveen K. Krishnaprasad, Venkatram Vishwanath, Shalini Venkataraman, A. G. Rao, Luc Renambot, Jason Leigh, Andrew E. Johnson 0001
CLUSTER2
2004 Vol-a-Tile - A Tool for Interactive Exploration of Large Volumetric Data on Scalable Tiled Displays
abstract
We present the current state of Vol-a-Tile, an interactive tool for exploring large volumetric data on scalable tiled displays. Vol-a-Tile presents a variety of features employed by scientists at the Scripps Institution of Oceanography on data collected from the Anatomy of a Ridge-Axis Discontinuity seismic experiment. Hardware texture mapping and level-of-detail techniques provide interactivity. A high-performance network protocol is used to connect remote data sources over high-bandwidth photonic networks.
Nicholas Schwarz, Shalini Venkataraman, Luc Renambot, Naveen K. Krishnaprasad, Venkatram Vishwanath, Jason Leigh, Andrew E. Johnson 0001, Graham Kent, Atul Nayak
IEEE Visualization5