Wei-keng Liao

dblp:45/1370 · also Wei-Keng Liao · DBLP profile ↗
← Back
103ranked-venue papers
12as first author
20since 2021 · last 2024
0009-0008-9411-2543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 60 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 26 · 10 since 2021Databases, data management, data science and information retrieval · 22 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Automated Nanoparticle Image Processing Pipeline for AI-Driven Materials Characterization
abstract
Recent innovations have made it possible to produce millions of distinct nanoparticles on a chip. These vast volumes of data are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image preprocessing step before model training can produce significantly higher-performing models in a fraction of the time and make the model more robust to different image noise levels and microscope acquisition settings. The proposed image processing pipeline effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures, resulting in higher performance and significant cost savings. Experiments demonstrate superior performance relative to our baseline, including a 15% improvement in recall and more than a 10% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives in our application, and our best-performing model obtains a precision of 97.3% and weighted F-score of 95.9% on an unseen test set. Additionally, model training time is reduced from 15.5 hours to 32 seconds. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will offer a considerable speedup in the laboratory, allowing researchers to rapidly and accurately analyze much greater volumes of data and accelerate materials discovery.
Alexandra L. Day, Carolin B. Wahl, Roberto dos Reis, Wei-keng Liao, Vinayak P. Dravid, Alok N. Choudhary, Ankit Agrawal 0001
CIKM4
2024 Combining Transfer Learning and Representation Learning to Improve Predictive Analytics on Small Materials Data
abstract
Modern data mining methods have seen a widespread and growing application in the field of materials science for regression-based predictive modeling due to their effectiveness in extracting and utilizing the hidden information from the materials datasets. However, due to the costly and time-consuming nature of the methods involved in obtaining the experimental and computational data, the majority of the materials datasets are small in size. Moreover, limited hand-engineered representations available from the raw materials data make it harder to improve the accuracy of predictive models on such small and specialized training datasets. In this paper, we introduce a novel technique that combines transfer learning (TL) and representation learning (RL) using a pre-trained deep neural network to maximize accuracy without additional computational costs on inorganic material properties. The performance of the proposed method is compared against traditional machine learning (ML), and deep neural network models trained from scratch (SC) with elemental fraction (EF) as input, more informative physical attributes (PA) as input (for a stringent comparison), as well as conventional TL and RL techniques using deep neural networks. The results demonstrate that the proposed method can improve the accuracy as compared to SC models and conventional TL and RL techniques.
Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
ICMLA2
2024 Enhancing Deep Neural Network Classification Performance Through Novel Weight Initialization: t-SNE Supported Walsh Matrix Approach
abstract
Deep Neural Networks, as a subset of AI, outper-form in understanding complex relationships. The key to this success lies in the network's ability to adapt to problem-specific nuances. During model training, the network dynamically optimizes its weights by updating them during backpropagation while trying to minimize the value of the loss function. Throughout this process, the shaping of model weights is crucially linked to how they were initialized. In this study, we introduce the auxiliary network model, called Sup-Walsh (Support Walsh), which reorganizes weights to enhance class boundaries. We tested our approach on three publicly available datasets using popular classification models. For instance, when using AlexNet [1] on the MNIST dataset [2], integrating Sup-Walsh led to a significant increase in accuracy after first epoch from 14.61% to 78.99%. Similarly, GoogleNet [3] on the FashionMNIST dataset [4] showed a notable 31.61% accuracy difference between configurations without and with Sup-Walsh after first epoch. Across nearly all experiments, our proposed method consistently outperformed existing approaches, demonstrating its potential to improve classification accuracy. Code availability: Code is available at Efficient-Weight-Initializer.
Muhammed Nur Talha Kilic, Vishu Gupta, Yuwei Mao, Kewei Wang 0002, Alec Peltekian, Alok N. Choudhary, Wei-keng Liao, Ankit Agrawal 0001
ICMLA7
2024 Tackling the Nonlinearity Problem in Inverse Modeling: Mixture Density Network-Backed Quantized AutoEncoder
abstract
Generative models have been widely used in the field of computer vision due to their ability to produce unseen data points. Its application has proven to be useful in various scientific domains such as materials science for generating new microstructure images that require learning nonlinear and one-to-many property-microstructure relationships. However, existing simulation-based solutions for this application are inefficient and time-consuming. Moreover, nonlinearity from the lower to higher dimensions poses considerable challenges. In this work, we propose a novel Mixture Density Network (MDN) based Quan-tized Autoencoder Network designed to generate microstructure images from only a single data point by establishing one-to-many nonlinear relationships from property to microstructure. Once the autoencoder effectively compresses spatial information into the property domain, we use the combination of produced latent vectors and property values to create a supplementary dataset for MDN. Upon completion of the model training, generative structures are extracted and merged to create a framework that generates images based on target property values (i.e., absorption values in this study). The trained MDN demonstrates proficiency in generating latent vectors within distribution, while the proposed Vector Quantized Variational Autoencoder (VQ-VAE) efficiently maps the embedding table to the latent space, generating images from property values within the range of properties observed during training. We demonstrate that our proposed model consistently outperforms the baselines with respect to generating new microstructure images having target properties and overcoming the above-mentioned challenges.
Muhammed Nur Talha Kilic, Yuwei Mao, Vishu Gupta, Alok N. Choudhary, Wei-keng Liao, Ankit Agrawal 0001
ICMLA5
2024 Deep Learning Based Inverse Modeling for Materials Design: From Microstructure and Property to Processing
abstract
Polycrystalline materials are crucial in various industries, necessitating a comprehensive understanding of the processing-structure-property-performance (PSPP) relationships. Traditional experimental methods are laborious and slow, while computational approaches predominantly address forward problems, deriving structures and properties from processing conditions. Conversely, inferring processing parameters from desired microstructures and properties remains a crucial yet challenging inverse problem due to the complex and nonlinear mappings involved. In this work, we propose a deep learning-based framework exploring non-sequential and sequential models to address two key inverse problems: predicting processing parameters from microstructures and from properties. Focusing on microstructural texture defined by the orientation distribution function (ODF), we apply our framework to copper, generating a dataset of 31,588 unique processing strain rates ($s$-1) in [0, 1] with corresponding ODFs and homogenized properties through simulations. Our inverse prediction results on processing parameters demonstrate high accuracy, with average test RMSEs of 0.0152 from microstructures and 0.0295 from properties. These findings validate the framework's efficacy as a tool for polycrystalline materials process design, enabling the precise determination of processing methods to achieve desired microstructures and properties.
Kewei Wang 0002, Yuwei Mao, Mahmudul Hasan 0016, Md Maruf Billah, Muhammed Nur Talha Kilic, Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Pinar Acar, Ankit Agrawal 0001
ICMLA7
2024 h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre-exascale platforms
abstract
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.
Jean Luca Bez, Houjun Tang, M. Scot Breitenfeld, Huihuo Zheng, Wei-keng Liao, Kaiyuan Hou, Zanhua Huang, Surendra Byna
Concurr. Comput. Pract. Exp.5
2023 A Case Study of Data Management Challenges Presented in Large-Scale Machine Learning Workflows
abstract
Running scientific workflow applications on high-performance computing systems provides promising results in terms of accuracy and scalability. An example is the particle track reconstruction research in high-energy physics that consists of multiple machine-learning tasks. However, as the modern HPC system scales up, researchers spend more effort on coordinating the individual workflow tasks due to their increasing demands on computational power, large memory footprint, and data movement among various storage devices. These issues are further exacerbated when intermediate result data must be shared among different tasks and each is optimized to fulfill its own design goals, such as the shortest time or minimal memory footprint. In this paper, we investigate the data management challenges presented in scientific workflows. We observe that individual tasks, such as data generation, data curation, model training, and inference, often use data layouts only best for one's I/O performance but orthogonal to its successive tasks. We propose various solutions by employing alternative data structures and layouts in consideration of two tasks running consecutively in the workflow. Our experimental results show up to a 16.46x and 3.42x speedup for initialization time and I/O time respectively, compared to previous approaches.
Claire Songhyun Lee, V. Hewes, Giuseppe Cerati, Jim Kowalkowski, Adam Aurisano, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
CCGrid8
2023 A Deep Learning Framework for Time-Series Processing-Microstructure-Property Prediction
abstract
A process simulator provides valuable insights into the evolution of microstructures under various elementary processes, employing the Orientation Distribution Function (ODF) as a representation of the microstructure's texture. However, such simulations often involve complex physical computations, making them time-consuming. To address this, our study introduces an artificial intelligence (AI)-based framework to predict the microstructural texture of polycrystalline materials using a specified deformation process. As a case study, we apply our framework to copper. The dataset includes 3,125 unique processing parameter combinations and their corresponding ODF vectors generated using a process simulator. The resulting predictions enable the calculation of homogenized properties. As opposed to traditional material processing simulations, our AI-driven framework offers faster results with minimal error rates (less than 0.5%). This indicates that our approach is a promising tool for rapidly predicting processing-specific microstructures and properties, thereby offering significant improvements over conventional simulation techniques.
Yuwei Mao, Mahmudul Hasan 0016, Claire Songhyun Lee, Muhammed Nur Talha Kilic, Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Pinar Acar, Ankit Agrawal 0001
ICMLA6
2023 Pre-Activation based Representation Learning to Enhance Predictive Analytics on Small Materials Data
abstract
Artificial intelligence based predictive modeling has become increasingly sought-after in the field of materials science for training property prediction models due to their promising ability to extract and utilize data-driven information from materials data. However, current methods typically use limited hand-engineered fixed-length representations obtained from available composition-based information only, making model inputs a stumbling block when handling small and specialized training datasets. In this paper, we study and propose a method to perform representation learning (RL) that is both applicable and adaptive for generalized use across various domains. We introduce a RL technique that utilizes pre-activation based representations extracted from a model pre-trained using a deep neural network to maximize the accuracy. We perform model training for inorganic material properties using composition-based numerical vectors representing the elemental fractions (EF) of the materials by leveraging source models trained on large datasets to build target models on small datasets and then compare its performance against traditional machine learning (ML), deep neural network and RL-based graph neural network (GNN) models trained from scratch (SC) with EF as input, more informative physical attributes (PA) as input, as well as conventional TL/RL techniques. Using large$(\sim 345K)$datasets for source model training and small computational$(\sim 28K)$and experimental$(\sim 2K)$datasets for target model training and testing, we show that the proposed RL methods help significantly improve the accuracy of the model as compared to the SC models and conventional TL/RL techniques for all data sizes and properties by using only EF as input. We also perform a statistical significance analysis by calculating the p-value to find that the observed improvement in the accuracy of proposed RL model over SC, RL-based GNN, and conventional TL/RL models is indeed significant.
Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
IJCNN2
2023 AI for Learning Deformation Behavior of a Material: Predicting Stress-Strain Curves 4000x Faster Than Simulations
abstract
Stress-strain curves are important representations of a given material's mechanical properties, which depend primarily on the orientation of the individual crystals in the microstructure. Generating stress-strain curves from numerical methods such as the crystal plasticity finite element (CPFE) simulations is computationally intensive. As a result, it is difficult to generate complete stress-strain curves for all possible orientations of a material. In this work, we propose a bilinear stress-strain curve prediction framework for metallic alloys by integrating supervised and unsupervised deep learning methods via transfer learning principles. As a specific case-study, we focus on predicting stress-strain curves of Nickel (Ni)-based superalloys that have important applications in aerospace industry. Using a small training set of just 100 complete stress-strain curves (4,000 strain steps each) of different orientations generated by CPFE simulation code, we were able to build a model that could accurately predict stress-strain curves (<2 % error) using simple features that could be obtained by running the CPFE simulation for just a single strain step. The proposed model can thus predict the complete stress-strain curve for a given orientation of Ni-based superalloys in a fraction of a second, which amounts to a speedup of over 4000x as compared to the simulation.
Yuwei Mao, Shahriyar Keshavarz, Vishu Gupta, Andrew C. E. Reid, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
IJCNN5
2023 I/O in WRF: A Case Study in Modern Parallel I/O Techniques
abstract
Large-scale parallel applications can face significant I/O performance bottlenecks, making efficient I/O crucial. This work presents a comparative study of several parallel I/O implementations in the Weather Research and Forecasting model, including PnetCDF blocking and non-blocking I/O options, netCDF4, HDF5 Log VOL, and ADIOS. For I/O methods creating files in a canonical data layout, PnetCDF's non-blocking option offers up to 2x improvement over its blocking option and up to 4.5x over HDF5 via netCDF4, demonstrating the effectiveness of the write request aggregation technique. The HDF5 Log VOL outperforms ADIOS with a 4x improvement in write performance when creating files in the log layout, although both require non-negligible time to convert the file back to canonical order for post-run analysis. From these results we extract some observations that can guide I/O strategies for modern parallel codes.
Zanhua Huang, Kaiyuan Hou, Ankit Agrawal 0001, Alok N. Choudhary, Robert B. Ross, Wei-keng Liao
SC6
2022 Using Multi-Resolution Data to Accelerate Neural Network Training in Scientific Applications
abstract
Neural networks are powerful solutions to many scientific applications; however, they usually require long model training time due to large training data sets or large model size. Research has been focused on developing numerical optimization algorithms and parallel processing to reduce the training time. In this work, we propose a multi-resolution strategy that can reduce the training time by training the model with the reduced-resolution data samples at the beginning and later switching to the original resolution data samples. This strategy is motivated by the observation that coarser versions of many applications can be solved faster than their denser counterparts, and the solution to a coarser problem could be used to initialize the solution to the denser problem. When applying the idea to neural network training, coarse data can have a similar effect on the learning curves at the early stage as the dense data but requires less time. Once the curves no longer improve significantly, our strategy switches to using the data in original resolution. The key in this process is the ability to generate multiple resolutions of a problem automatically, which could usually be done with scientific applications with spatial and temporal continuity. We use two real-world scientific applications, CosmoFlow and DeepCAM, to evaluate the proposed mixed-resolution training strategy. Our experiment results demonstrate that the proposed training strategy effectively reduces the end-to-end training time while achieving a comparable accuracy to that of the training only with the original data. While maintaining the same model accuracy, our multi-resolution training strategy reduces the end-to-end training time up to 30% and 23% for CosmoFlow and DeepCAM, respectively.
Kewei Wang 0002, Sunwoo Lee 0001, Jan Balewski, Alex Sim, Peter Nugent, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu, Wei-keng Liao
CCGRID9
2022 BRNet: Branched Residual Network for Fast and Accurate Predictive Modeling of Materials Properties
abstract
Machine Learning (ML) and Deep Learning (DL) have become increasingly popular in the field of materials science for building property prediction models owing to their ability to efficiently extract and understand data-driven relationships between materials composition, structure, and properties. In general, materials property prediction are regression problems with a vector-based input material representation. While fully connected layers have been widely used in deep neural networks to predict materials properties, simply adding more and more layers to create a deep model often degrades their performance due to the vanishing gradient problem, thereby limiting usage. In this paper, we study and propose architectural principles for building deep regression neural networks comprising fully connected layers with numerical vectors that bypass manual feature engineering. We introduce a novel deep regression neural network with branched residual learning, BRNet, consisting of branching of layers to maximize variation of features learned from the input or previous layer and places skip connections after each layer to minimize the information loss due to vanishing gradient. We perform BRNet model training for inorganic material properties using numerical vectors representing the elemental fractions of the compositions of the respective materials and compare its performance against other traditional ML and DL techniques, including ElemNet and IRNet. Using multiple datasets (such as OQMD, MP, JARVIS) for training and testing, we show that BRNet models are significantly more accurate than the state-of-the-art ML methods and DL models for all data sizes by using only raw elemental fractions as input. We also show that BRNet's branched residual learning requires fewer parameters and leads to better convergence during the training phase than other neural networks, thus resulting in faster model training.
Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
SDM2
2022 Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency
Sunwoo Lee 0001, Qiao Kang, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
J. Parallel Distributed Comput.6
2022 A case study on parallel HDF5 dataset concatenation for high energy physics data analysis
Sunwoo Lee 0001, Kaiyuan Hou, Kewei Wang 0002, Saba Sehrish, Marc F. Paterno, Jim Kowalkowski, Quincey Koziol, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
Parallel Comput.11
2021 Supporting Data Compression in PnetCDF
abstract
Recently, the dramatic increase of the data amounts drives up the demand for data compression among HPC applications. Although many file systems and I/O middlewares have incorporated compression features, few high-level parallel I/O libraries support data compression due to the challenges of achieving scalable performance on HPC systems. This paper presents the design and implementation of the variable compression feature in the Parallel NetCDF library. Our design employs the same concept of chunking used by the HDF5 library, but we focus on enabling I/O aggregation across multiple requests to address the challenges on performance and scalability. We evaluate our solution using the I/O kernel of real-world scientific applications and analyze the impacts of data compression on parallel I/O performance. Our result suggests that handling multiple requests at once can significantly improve the parallel I/O performance on chunked and compressed data.
Kaiyuan Hou, Qiao Kang, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE BigData6
2021 Optimizing Performance of Parallel I/O Accesses to Non-contiguous Blocks in Multiple Array Variables
abstract
Accessing non-contiguous blocks in multiple array variables is a challenging I/O pattern for parallel applications to obtain good I/O performance. High-level I/O libraries such as HDF5 allow users to implement this pattern conveniently, but users have observed significant performance bottlenecks in the two-phase I/O implementation of MPI-IO. Recent studies have advanced the two-phase I/O performance by novel communication algorithms, but such improvements still have limitations. Two-phase I/O has to faithfully process inputs from high-level I/O libraries, so that implementation overheads can accumulate for improper usage of high-level I/O libraries. In this paper, we propose approaches for efficient usage of high-level I/O libraries that can circumvent major collective I/O overheads. We adopt a multi-dataset implementation of HDF5 dataset I/O to aggregate non-contiguous requests for array blocks and provide corresponding parameter assignment strategies. These approaches reduce the overheads caused by communication straggler effects in two-phase I/O. We show that our proposed methods can improve the parallel I/O performance up to 8× on two supercomputing systems for the HDF5 implementations of an I/O kernel extracted from climate simulation code compared with its baseline implementations.
Qiao Kang, M. Scot Breitenfeld, Kaiyuan Hou, Wei-keng Liao, Robert B. Ross, Surendra Byna
IEEE BigData4
2021 Asynchronous I/O Strategy for Large-Scale Deep Learning Applications
abstract
Many scientific applications have started using deep learning methods for their classification or regression problems. However, for data-intensive scientific applications, I/O performance can be the major performance bottleneck. In order to effectively solve important real-world problems using deep learning methods on High-Performance Computing (HPC) systems, it is essential to address the poor I/O performance issue in large-scale neural network training. In this paper, we propose an asynchronous I/O strategy that can be generally applied to deep learning applications. Our I/O strategy employs an I/O -dedicated thread per process, that performs I/O operations independently of the training progress. The I/O thread reads many training samples at once to reduce the total number of I/O operations per epoch. Given the fixed amount of training data, the fewer the I/O operations per epoch, the shorter the overall I/O time. The I/O operations are also overlapped with the computations using the double-buffering method. We evaluate our I/O strategy using two real-world scientific applications, CosmoFlow and Neuron-Inverter. Our experimental results demonstrate that the proposed I/O strategy significantly improves the scaling performance without affecting the regression performance.
Sunwoo Lee 0001, Qiao Kang, Kewei Wang 0002, Jan Balewski, Alex Sim, Ankit Agrawal 0001, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao
HiPC10
2021 SIGRNN: Synthetic Minority Instances Generation in Imbalanced Datasets using a Recurrent Neural Network
Reda Al-Bahrani, Dipendra Jha, Qiao Kang, Sunwoo Lee 0001, Zijiang Yang 0008, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary
ICPRAM6
2021 Enhancing Phase Mapping for High-throughput X-ray Diffraction Experiments using Fuzzy Clustering
Dipendra Jha, K. V. L. V. Narayanachari, Denis T. Keane, Wei-keng Liao, Alok N. Choudhary, Yip-Wah Chung, Michael J. Bedzyk, Ankit Agrawal 0001
ICPRAM5
2020 Communication-Efficient Local Stochastic Gradient Descent for Scalable Deep Learning
abstract
Synchronous Stochastic Gradient Descent (SGD) with data parallelism, the most popular parallel training strategy for deep learning, suffers from expensive gradient communications. Local SGD with periodic model averaging is a promising alternative to synchronous SGD. The algorithm allows each worker to locally update its own model, and periodically averages the model parameters across all the workers. While this algorithm enjoys less frequent communications, the convergence rate is strongly affected by the number of workers. In order to scale up the local SGD training without losing accuracy, the number of workers should be sufficiently small so that the model converges reasonably fast. In this paper, we discuss how to exploit the degree of parallelism in local SGD while maintaining model accuracy. Our training strategy employs multiple groups of processes and each group trains a local model based on data parallelism. The local models are periodically averaged across all the groups. Based on this hierarchical parallelism, we design a model averaging algorithm that has a cheaper communication cost than allreduce-based approach. We also propose a practical metric for finding the maximum number of workers that does not cause a significant accuracy loss. Our experimental results demonstrate that our proposed training strategy provides a significantly improved scalability while achieving a comparable model accuracy to synchronous SGD.
Sunwoo Lee 0001, Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE BigData5
2020 Predicting Resource Requirement in Intermediate Palomar Transient Factory Workflow
abstract
Quickly identifying astronomical transients from synoptic surveys is critical to many recent astrophysical discoveries. However, each of the data processing pipelines in these surveys contains dozens of stages with highly varying time and space requirements. Properly predicting the resources required to run these pipelines is critical for the allocation of computing resources and reducing the discovery response time. We propose a machine learning strategy for this prediction task and demonstrate its effectiveness using a set of timing measurements from the intermediate Palomar Transient Factory (iPTF) workflow. The proposed model utilizes the spatiotemporal correlation of astronomical images, where nearby patches of the sky (space) are likely to have a similar number of objects of interest and workflows executed in the recent past (time) are likely to use a similar amount of time because the machines and data storage systems are likely to be in similar states. We capture the relationship among these spatial and temporal features in a Bayesian network and study how they impact the prediction accuracy. This Bayesian network helps us to identify the most influential features for predictions. With proper features, our models achieve errors close to the random variance boundary within batches of images taken at the same time, which can be regarded as the intrinsic limit of prediction accuracy.
Qiao Kang, Alex Sim, Peter Nugent, Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu
CCGRID5
2020 Improving all-to-many personalized communication in two-phase I/O
abstract
As modern parallel computers enter the exascale era, the communication cost for redistributing requests becomes a significant bottleneck in MPIIO routines. The communication kernel for request redistribution, which has an all-to-many personalized communication pattern for application programs with a large number of noncontiguous requests, plays an essential role in the overall performance. This paper explores the available communication kernels for two-phase I/O communication. We generalize the spread-out algorithm to adapt to the all-to-many communication pattern of two-phase I/O by reducing the communication straggler effect. Communication throttling methods that reduce communication contention for asynchronous MPI implementation are adopted to improve communication performance further. Experimental results are presented using different communication kernels running on Cray XC40 Cori and IBM AC922 Summit supercomputers with different I/O patterns. Our study shows that adjusting communication kernel algorithms for different I/O patterns can improve the end-to-end performance up to 10 times compared with default MPI-IO implementations.
Qiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
SC7
2020 Improving MPI Collective I/O for High Volume Non-Contiguous Requests With Intra-Node Aggregation
abstract
Two-phase I/O is a well-known strategy for implementing collective MPI-IO functions. It redistributes I/O requests among the calling processes into a form that minimizes the file access costs. As modern parallel computers continue to grow into the exascale era, the communication cost of such request redistribution can quickly overwhelm collective I/O performance. This effect has been observed from parallel jobs that run on multiple compute nodes with a high count of MPI processes on each node. To reduce the communication cost, we present a new design for collective I/O by adding an extra communication layer that performs request aggregation among processes within the same compute nodes. This approach can significantly reduce inter-node communication contention when redistributing the I/O requests. We evaluate the performance and compare it with the original two-phase I/O on Cray XC40 parallel computers (Theta and Cori) with Intel KNL and Haswell processors. Using I/O patterns from two large-scale production applications and an I/O benchmark, we show our proposed method effectively reduces the communication cost and hence maintains the scalability for a large number of processes.
Qiao Kang, Sunwoo Lee 0001, Kaiyuan Hou, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
IEEE Trans. Parallel Distributed Syst.7
2019 Spatiotemporal Real-Time Anomaly Detection for Supercomputing Systems
abstract
The demands of increasingly large scientific application workflows lead to the need for more powerful supercomputers. As the scale of supercomputing systems have grown, the prediction of fault tolerance has become an increasingly critical area of study, since the prediction of system failures can improve performance by saving checkpoints in advance. We propose a real-time failure detection algorithm that adopts an event-based prediction model. The prediction model is a convolutional neural network that utilizes both traditional event attributes and additional spatio-temporal features. We present a case study using our proposed method with six years of reliability, availability, and serviceability event logs recorded by Mira, a Blue Gene/Q supercomputer at Argonne National Laboratory. In the case study, we have shown that our failure prediction model is not limited to predict the occurrence of failures in general. It is capable of accurately detecting specific types of critical failures such as coolant and power problems within reasonable lead time ranges. Our case study shows that the proposed method can achieve a F1score of 0.56 for general failures, 0.97 for coolant failures, and 0.86 for power failures.
Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Alex Sim, Kesheng Wu, Rajkumar Kettimuthu, Pete Beckman, Zhengchun Liu, Wei-keng Liao
IEEE BigData9
2019 Improving Scalability of Parallel CNN Training by Adjusting Mini-Batch Size at Run-Time
abstract
Training Convolutional Neural Network (CNN) is a computationally intensive task, requiring efficient parallelization to shorten the execution time. Considering the ever-increasing size of available training data, the parallelization of CNN training becomes more important. Data-parallelism, a popular parallelization strategy that distributes the input data among compute processes, requires the mini-batch size to be sufficiently large to achieve a high degree of parallelism. However, training with large batch size is known to produce a low convergence accuracy. In image restoration problems, for example, the batch size is typically tuned to a small value between 16 ~ 64, making it challenging to scale up the training. In this paper, we propose a parallel CNN training strategy that gradually increases the mini-batch size and learning rate at run-time. While improving the scalability, this strategy also maintains the accuracy close to that of the training with a fixed small batch size. We evaluate the performance of the proposed parallel CNN training algorithm with image regression and classification applications using various models and datasets.
Sunwoo Lee 0001, Qiao Kang, Sandeep Madireddy, Prasanna Balaprakash, Ankit Agrawal 0001, Alok N. Choudhary, Rick Archibald, Wei-keng Liao
IEEE BigData8
2019 A Real-Time Iterative Machine Learning Approach for Temperature Profile Prediction in Additive Manufacturing Processes
abstract
Additive Manufacturing (AM) is a manufacturing paradigm that builds three-dimensional objects from a computer-aided design model by successively adding material layer by layer. AM has become very popular in the past decade due to its utility for fast prototyping such as 3D printing as well as manufacturing functional parts with complex geometries using processes such as laser metal deposition that would be difficult to create using traditional machining. As the process for creating an intricate part for an expensive metal such as Titanium is prohibitive with respect to cost, computational models are used to simulate the behavior of AM processes before the experimental run. However, as the simulations are computationally costly and time-consuming for predicting multiscale multi-physics phenomena in AM, physics-informed data-driven machine-learning systems for predicting the behavior of AM processes are immensely beneficial. Such models accelerate not only multiscale simulation tools but also empower real-time control systems using in-situ data. In this paper, we design and develop essential components of a scientific framework for developing a data-driven model-based real-time control system. Finite element methods are employed for solving time-dependent heat equations and developing the database. The proposed framework uses extremely randomized trees - an ensemble of bagged decision trees as the regression algorithm iteratively using temperatures of prior voxels and laser information as inputs to predict temperatures of subsequent voxels. The models achieve mean absolute percentage errors below 1% for predicting temperature profiles for AM processes. The code is made available for the research community at https://github.com/paularindam/ml-iter-additive.
Arindam Paul, Mojtaba Mozaffar, Zijiang Yang 0008, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
DSAA4
2019 Peak Area Detection Network for Directly Learning Phase Regions from Raw X-ray Diffraction Patterns
abstract
X-ray diffraction (XRD) is a well-known technique used by scientists and engineers to determine the atomic-scale structures as a basis for understanding the composition-structure-property relationship of materials. The current approach for the analysis of XRD data is a multi-stage process requiring several intensive computations such as integration along 2θ for conversion to 1D patterns (intensity-2θ), background removal by polynomial fitting, and indexing against a large database of reference peaks. It impacts the decisions about the subsequent experiments of the materials under investigation and delays the overall process. In this paper, we focus on eliminating such multi-stage XRD analysis by directly learning the phase regions from the raw (2D) XRD image. We introduce a peak area detection network (PADNet) that directly learns to predict the phase regions using the raw XRD patterns without any need for explicit preprocessing and background removal. PADNet contains specially designed large symmetrical convolutional filters at the first layer to capture the peaks and automatically remove the background by computing the difference in intensity counts across different symmetries. We evaluate PADNet using two sets of XRD patterns collected from SLAC and Bruker D-8 for the Sn-Ti-Zn-O composition space; each set contains 177 experimental XRD patterns with their phase regions. We find that PADNet can successfully classify the XRD patterns independent of the presence of background noise and perform better than the current approach of extrapolating phase region labels based on 1D XRD patterns.
Dipendra Jha, Aaron Gilad Kusne, Reda Al-Bahrani, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
IJCNN5
2019 Transfer Learning Using Ensemble Neural Networks for Organic Solar Cell Screening
abstract
Organic Solar Cells are a promising technology for solving the clean energy crisis in the world. However, generating candidate chemical compounds for solar cells is a time-consuming process requiring thousands of hours of laboratory analysis. For a solar cell, the most important property is the power conversion efficiency which is dependent on the highest occupied molecular orbitals (HOMO) values of the donor molecules. Recently, machine learning techniques have proved to be very useful in building predictive models for HOMO values of donor structures of Organic Photovoltaic Cells (OPVs). Since experimental datasets are limited in size, current machine learning models are trained on data derived from calculations based on density functional theory (DFT). Molecular line notations such as SMILES or InChI are popular input representations for describing the molecular structure of donor molecules. The two types of line representations encode different information, such as SMILES defines the bond types while InChi defines protonation. In this work, we present an ensemble deep neural network architecture, called SINet, which harnesses both the SMILES and InChI molecular representations to predict HOMO values and leverage the potential of transfer learning from a sizeable DFT-computed dataset- Harvard CEP to build more robust predictive models for relatively smaller HOPV datasets. Harvard CEP dataset contains molecular structures and properties for 2.3 million candidate donor structures for OPV while HOPV contains DFT-computed and experimental values of 350 and 243 molecules respectively. Our results demonstrate significant performance improvement from the use of transfer learning and leveraging both molecular representations.
Arindam Paul, Dipendra Jha, Reda Al-Bahrani, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
IJCNN4
2019 Deep learning based domain knowledge integration for small datasets: Illustrative applications in materials informatics
abstract
Deep learning has shown its superiority to traditional machine learning methods in various fields, and in general, its success depends on the availability of large amounts of reliable data. However, in some scientific fields such as materials science, such big data is often expensive or even impossible to collect. Thus given relatively small datasets, most of data-driven methods are based on traditional machine learning methods, and it is challenging to apply deep learning for many tasks in these fields. In order to take the advantage of deep learning even for small datasets, a domain knowledge integration approach is proposed in this work. The efficacy of the proposed approach is tested on two materials science datasets with different types of inputs and outputs, for which domain knowledge-aware convolutional neural networks (CNNs) are developed and evaluated against traditional machine learning methods and standard CNN-based approaches. Experiment results demonstrate that integrating domain knowledge into deep learning can not only improve the model's performance for small datasets, but also make the prediction results more explainable based on domain knowledge.
Zijiang Yang 0008, Reda Al-Bahrani, Andrew C. E. Reid, Stefanos Papanikolaou, Surya R. Kalidindi, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
IJCNN6
2019 IRNet: A General Purpose Deep Residual Regression Framework for Materials Discovery
abstract
Materials discovery is crucial for making scientific advances in many domains. Collections of data from experiments and first-principle computations have spurred interest in applying machine learning methods to create predictive models capable of mapping from composition and crystal structures to materials properties. Generally, these are regression problems with the input being a 1D vector composed of numerical attributes representing the material composition and/or crystal structure. While neural networks consisting of fully connected layers have been applied to such problems, their performance often suffers from the vanishing gradient problem when network depth is increased. Hence, predictive modeling for such tasks has been mainly limited to traditional machine learning techniques such as Random Forest. In this paper, we study and propose design principles for building deep regression networks composed of fully connected layers with numerical vectors as input. We introduce a novel deep regression network with individual residual learning, IRNet, that places shortcut connections after each layer so that each layer learns the residual mapping between its output and input. We use the problem of learning properties of inorganic materials from numerical attributes derived from material composition and/or crystal structure to compare IRNet's performance against that of other machine learning techniques. Using multiple datasets from the Open Quantum Materials Database (OQMD) and Materials Project for training and evaluation, we show that IRNet provides significantly better prediction performance than the state-of-the-art machine learning approaches currently used by domain scientists. We also show that IRNet's use of individual residual learning leads to better convergence during the training phase than when shortcut connections are between multi-layer stacks while maintaining the same number of parameters.
Dipendra Jha, Logan T. Ward, Zijiang Yang 0008, Christopher Wolverton, Ian T. Foster, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001
KDD6
2019 Scalable Algorithms for MPI Intergroup Allgather and Allgatherv
Qiao Kang, Jesper Larsson Träff, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
Parallel Comput.6
2018 Parallel DBSCAN Algorithm Using a Data Partitioning Strategy with Spark Implementation
abstract
DBSCAN is a well-known clustering algorithm which is based on density and is able to identify arbitrary shaped clusters and eliminate noise data. However, existing parallel implementation strategies based on MPI lack fault tolerance and there is no guarantee that their workload is balanced. Although some of Hadoop-based approaches have been proposed, they do not perform well in terms of scalability since the merge process is not efficient.We propose a scalable parallel DBSCAN algorithm by applying a partitioning strategy. It is implemented in Apache Spark. In order to reduce search time, kdtree is used in our algorithm. To achieve better performance and scalability based on kdtree, we adopt an effective partitioning technique aimed at producing balanced sub-domains which can be computed within Spark executors. Moreover, we came up with a new merging technique: through mapping the relationship between the local points and their bordering neighbors, all the partial clusters which are generated in executors are merged to form the final complete clusters. We have observed and verified (through experiments) that this merging approach is very effective in reducing the time taken for the merge phase and very scalable with increasing the number of processing cores and the generated partial clusters.We implemented the algorithm in Java, evaluated its scalability by using different number of processing cores, and using real and synthetic datasets containing up to several hundred million high-dimensional points. We used three scales of datasets to evaluate our implementation. For small scale, we use 50k, 100k, and 500k data points, obtaining up to a factor of 14.9 speedup when using 16 cores. For medium scale, we use 1.0m, 1.5m, and 1.9m data points, obtaining a factor of 109.2 speedup when using 128 cores. For large scale, we use 61.0m, 91.5m, and 115.9m data points, obtaining a factor of 8344.5 speedup when using 16384 cores.
Dianwei Han, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
IEEE BigData3
2018 Full-Duplex Inter-Group All-to-All Broadcast Algorithms with Optimal Bandwidth
abstract
MPI inter-group collective communication patterns can be viewed as bipartite graphs that divide processes into two disjoint groups in which messages are transferred between but not within the groups. Such communication patterns can serve as basic operations for scientific application workflows. In this paper, we present parallel algorithms for inter-group all-to-all broadcast (Allgather) communication with optimal bandwidth for any message size and process number under single-port communication constraints. We implement the algorithms using MPI point-to-point and intra-group collective communication functions and evaluate their performance on the Cori supercomputer at NERSC. Using message sizes ranging from 256B to 64MB, the experiments show a significant performance improvement achieved by our algorithm, which is up to 9.27 times faster than production MPI libraries that adopt the so called root-gathering algorithm.
Qiao Kang, Jesper Larsson Träff, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
EuroMPI6
2017 Parallel Deep Convolutional Neural Network Training by Exploiting the Overlapping of Computation and Communication
abstract
Training Convolutional Neural Network (CNN) is a computationally intensive task whose parallelization has become critical in order to complete the training in an acceptable time. However, there are two obstacles to developing a scalable parallel CNN in a distributed-memory computing environment. One is the high degree of data dependency exhibited in the model parameters across every two adjacent minibatches and the other is the large amount of data to be transferred across the communication channel. In this paper, we present a parallelization strategy that maximizes the overlap of inter-process communication with the computation. The overlapping is achieved by using a thread per compute node to initiate communication after the gradients are available. The output data of backpropagation stage is generated at each model layer, and the communication for the data can run concurrently with the computation of other layers. To study the effectiveness of the overlapping and its impact on the scalability, we evaluated various model architectures and hyperparameter settings. When training VGG-A model using ImageNet data sets, we achieve speedups of 62.97× and 77.97× on 128 compute nodes using mini-batch sizes of 256 and 512, respectively.
Sunwoo Lee 0001, Dipendra Jha, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
HiPC5
2017 Building Halo Merger Trees from the Q Continuum Simulation
abstract
Cosmological N-body simulations rank among the most computationally intensive efforts today. A key challenge is the analysis of structure, substructure, and the merger history for many billions of compact particle clusters, called halos. Effectively representing the merging history of halos is essential for many galaxy formation models used to generate synthetic sky catalogs, an important application of modern cosmological simulations. Generating realistic mock catalogs requires computing the halo formation history from simulations with large volumes and billions of halos over many time steps, taking hundreds of terabytes of analysis data. We present fast parallel algorithms for producing halo merger trees and tracking halo substructure from a single-level, density-based clustering algorithm. Merger trees are created from analyzing the halo-particle membership function in adjacent snapshots, and substructure is identified by tracking the "cores" of merging halos – sets of particles near the halo center. Core tracking is performed after creating merger trees and uses the relationships found during tree construction to associate substructures with hosts. The algorithms are implemented with MPI and evaluated on a Cray XK7 supercomputer using up to 16,384 processes on data from HACC, a modern cosmological simulation framework. We present results for creating merger trees from 101 analysis snapshots taken from the Q Continuum, a large volume, high mass resolution, cosmological simulation evolving half a trillion particles.
Esteban Rangel, Nicholas Frontiere, Salman Habib 0002, Katrin Heitmann, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary
HiPC5
2017 A flexible I/O arbitration framework for netCDF-based big data processing workflows on high-end supercomputers
abstract
Summary On the verge of the convergence between high‐performance computing and Big Data processing, it has become increasingly prevalent to deploy large‐scale data analytics workloads on high‐end supercomputers. Such applications often come in the form of complex workflows with various different components, assimilating data from scientific simulations as well as from measurements streamed from sensor networks, such as radars and satellites. For example, as part of the Flagship 2020 (post‐K) supercomputer project of Japan, RIKEN is investigating the feasibility of a highly accurate weather forecasting system that would provide a real‐time outlook for severe guerrilla rainstorms. One of the main performance bottlenecks of this application is the lack of efficient communication among workflow components, which currently takes place over the parallel file system.In this paper, we present an initial study of a direct communication framework designed for complex workflows that eliminates unnecessary file I/O among components. Specifically, we propose an I/O arbitration layer that provides direct parallel data transfer (both synchronous and asynchronous) among job components that rely on the netCDF interface for performing I/O operations. Our solution requires only minimal modifications to application code. Moreover, we propose a configuration file–based approach that allows users to specify the desired data transfer pattern among workflow components, offering a general solution for different application contexts. We present a preliminary evaluation of the proposed framework on the K Computer (running on up to 4800 compute nodes) using RIKEN's experimental weather forecasting workflow as a case study.
Jianwei Liao 0001, Balazs Gerofi, Guo-Yuan Lien, Takemasa Miyoshi, Seiya Nishizawa, Hirofumi Tomita, Wei-keng Liao, Alok N. Choudhary, Yutaka Ishikawa
Concurr. Comput. Pract. Exp.7
2017 Reducing I/O variability using dynamic I/O path characterization in petascale storage systems
Seung Woo Son 0001, Saba Sehrish, Wei-keng Liao, Ron A. Oldfield, Alok N. Choudhary
J. Supercomput.3
2016 Evaluation of K-means data clustering algorithm on Intel Xeon Phi
abstract
Intel Xeon Phi is a processor based on MIC architecture that contains a large number of compute cores with a high local memory bandwidth and 512-bit vector processing units. To achieve high performance on Xeon Phi, it is important for programmers to explore all the software features provided by the Intel compiler and libraries to fully utilize the new hardware resources. In this paper, we use the K-Means algorithm to study the performance of various Intel software settings available for Xeon Phi and their impacts to the performance of K-means. At first we examine different memory layouts for storing data points using Intel compiler-intrinsic functions. During distance calculation, the computational kernel of K-means, when the size of individual input data points is not vector-friendly, we pad the data points to align with the VPU width. At last, we implement a parallel reduction to increase memory access parallelism and cache hits. These techniques enable us to successfully take advantage of thread-level parallelism and data-level parallelism on Xeon Phi. Experimental results demonstrate large performance gains over the default auto-vectorization approach. The K-Means implemented with the proposed techniques achieves up to 68.65% and 56.14% performance improvements for aligned datasets and unaligned datasets, respectively. For high-dimensional aligned datasets, we achieved up to 53.49% performance improvement on a large-scale parallel computer.
Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Nikos Hardavellas, Alok N. Choudhary
IEEE BigData2
2016 Materials discovery: Understanding polycrystals from large-scale electron patterns
abstract
This paper explores the idea of modeling a large image data collection of polycrystal electron patterns, in order to detect insights in understanding materials discovery. There is an emerging interest in applying big data processing, management and modeling methods to scientific images, which often come in a form and with patterns only interpretable to domain experts. While large-scale machine learning approaches have demonstrated certain superiority in analyzing, summarizing, and providing an understandable route to data types like natural images, speeches and texts, scientific images is still a relatively unexplored area. Deep convolutional neural networks, despite their recent triumph in natural image understanding, are still rarely seen adapted to experimental microscopic images, especially in a large scale. To the best of our knowledge, we present the first deep learning solution towards a scientific image indexing problem using a collection of over 300K microscopic images. The result obtained is 54% better than a dictionary lookup method which is state-of-the-art in the materials science society.
Rosanne Liu, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary, Marc De Graef
IEEE BigData3
2016 PinterNet: A thematic label curation tool for large image datasets
abstract
Recent progress in big data and computer vision with deep learning models has gained a lot of attention. Deep learning has been performed on tasks such as image classification, object detection, image segmentation, image captioning, visual question and answering, using large collections of annotated images. This calls for more curated large image datasets with clearer descriptions, cleaner contents, and diversified usability. However, the curation and labeling of such datasets can be labor-intensive. In this paper, we present PinterNet, an algorithm for automatic curation and label generation from noisy textual descriptions, and also publish a big image dataset containing over 110K images automatically labeled with their themes. Our dataset is hierarchical in nature, it has high level category information which we refer as verticals with fine-grained thematic labels at lower level. This advocates a new type of hierarchical theme classification problem closer to human cognition and of business value. We provide benchmark performances using deep learning models based on AlexNet architecture with different pre-training schemes for this novel task and new data.
Rosanne Liu, Diana Palsetia, Arindam Paul, Reda Al-Bahrani, Dipendra Jha, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary
IEEE BigData6
2016 A Filtering-based Clustering Algorithm for Improving Spatio-temporal Kriging Interpolation Accuracy
abstract
Geostatistical interpolation is the process that uses existing data and statistical models as inputs to predict data in unobserved spatio-temporal contexts as output. Kriging is a well-known geostatistical interpolation method that minimizes mean square error of prediction. The result interpolated by Kriging is accurate when consistency of statistical properties in data is assumed. However, without this assumption, Kriging interpolation has poor accuracy. To address this problem, this paper presents a new filtering-based clustering algorithm that partitions data into clusters such that the interpolation error within each cluster is significantly reduced, which in turn improves the overall accuracy. Comparisons to traditional Kriging are made with two real-world datasets using two error criteria: normalized mean square error(NMSE) and χ2 test statistics for normalized deviation measurement. Our method has reduced NMSE by more than 50% for both datasets over traditional Kriging. Moreover, χ2 tests have also shown significant improvements of our approach over traditional Kriging.
Qiao Kang, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary
CIKM2
2016 Parallel DTFE Surface Density Field Reconstruction
abstract
We improve the interpolation accuracy and efficiency of the Delaunay tessellation field estimator (DTFE) for surface density field reconstruction by proposing an algorithm that takes advantage of the adaptive triangular mesh for line-of-sight integration. The costly computation of an intermediate 3D grid is completely avoided by our method and only optimally chosen interpolation points are computed, thus, the overall computational cost is significantly reduced. The algorithm is implemented as a parallel shared-memory kernel for large-scale grid rendered field reconstructions in our distributed-memory framework designed for N-body gravitational lensing simulations in large volumes. We also introduce a load balancing scheme to optimize the efficiency of processing a large number of field reconstructions. Our results show our kernel outperforms existing software packages for volume weighted density field reconstruction, achieving~10x speedup, and our load balancing algorithm gains an additional~3.6x speedup at scales with~16k processes.
Esteban Rangel, Nan Li 0022, Salman Habib 0002, Tom Peterka, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
CLUSTER6
2016 Parallel Implementation of Lossy Data Compression for Temporal Data Sets
abstract
Many scientific data sets contain temporal dimensions. These are the data storing information at the same spatial location but different time stamps. Some of the biggest temporal datasets are produced by parallel computing applications such as simulations of climate change and fluid dynamics. Temporal datasets can be very large and cost a huge amount of time to transfer among storage locations. Using data compression techniques, files can be transferred faster and save storage space. NUMARCK is a lossy data compression algorithm for temporal data sets that can learn emerging distributions of element-wise change ratios along the temporal dimension and encodes them into an index table to be concisely represented. This paper presents a parallel implementation of NUMARCK. Evaluated with six data sets obtained from climate and astrophysics simulations, parallel NUMARCK achieved scalable speedups of up to 8788 when running 12800 MPI processes on a parallel computer. We also compare the compression ratios against two lossy data compression algorithms, ISABELA and ZFP. The results show that NUMARCK achieved higher compression ratio than ISABELA and ZFP.
William Hendrix, Seung Woo Son 0001, Christoph Federrath, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
HiPC6
2015 IOPro: a parallel I/O profiling and visualization framework for high-performance storage systems
Seong Jo Kim, Seung Woo Son 0001, Mahmut T. Kandemir, Wei-keng Liao, Rajeev Thakur, Alok N. Choudhary
J. Supercomput.5
2014 NUMARCK: Machine Learning Algorithm for Resiliency and Checkpointing
abstract
Data check pointing is an important fault tolerance technique in High Performance Computing (HPC) systems. As the HPC systems move towards exascale, the storage space and time costs of check pointing threaten to overwhelm not only the simulation but also the post-simulation data analysis. One common practice to address this problem is to apply compression algorithms to reduce the data size. However, traditional lossless compression techniques that look for repeated patterns are ineffective for scientific data in which high-precision data is used and hence common patterns are rare to find. This paper exploits the fact that in many scientific applications, the relative changes in data values from one simulation iteration to the next are not very significantly different from each other. Thus, capturing the distribution of relative changes in data instead of storing the data itself allows us to incorporate the temporal dimension of the data and learn the evolving distribution of the changes. We show that an order of magnitude data reduction becomes achievable within guaranteed user-defined error bounds for each data point. We propose NUMARCK, North western University Machine learning Algorithm for Resiliency and Check pointing, that makes use of the emerging distributions of data changes between consecutive simulation iterations and encodes them into an indexing space that can be concisely represented. We evaluate NUMARCK using two production scientific simulations, FLASH and CMIP5, and demonstrate a superior performance in terms of compression ratio and compression accuracy. More importantly, our algorithm allows users to specify the maximum tolerable error on a per point basis, while compressing the data by an order of magnitude.
Zhengzhang Chen, Seung Woo Son 0001, William Hendrix, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
SC5
2014 High performance data clustering: a comparative analysis of performance for GPU, RASC, MPI, and OpenMP implementations
Luobin Yang, Steve C. Chiu, Wei-keng Liao, Michael A. Thomas
J. Supercomput.3
2013 A probabilistic graphical model for brand reputation assessment in social networks
abstract
Social media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this paper, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and -0.71, respectively.
Kunpeng Zhang 0001, Doug Downey, Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
ASONAM7
2013 Mining diabetes complication and treatment patterns for clinical decision support
abstract
The fast development of hospital information systems (HIS) produces a large volume of electronic medical records, which provides a comprehensive source for exploratory analysis and statistics to support clinical decision-making. In this paper, we investigate how to utilize the heterogeneous medical records to aid the clinical treatments of diabetes mellitus. Diabetes mellitus, simply diabetes, is a group of metabolic diseases, which is often accompanied with many complications. We propose a Symptom-Diagnosis-Treatment model to mine the diabetes complication patterns and to unveil the latent association mechanism between treatments and symptoms from large volume of electronic medical records. Furthermore, we study the demographic statistics of patient population w.r.t. complication patterns in real data and observe several interesting phenomena. The discovered complication and treatment patterns can help physicians better understand their specialty and learn previous experiences. Our experiments on a collection of one-year diabetes clinical records from a famous geriatric hospital demonstrate the effectiveness of our approaches.
Lu Liu 0005, Jie Tang 0001, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
CIKM5
2013 Dynamic file striping and data layout transformation on parallel system with fluctuating I/O workload
abstract
As the number of compute cores on modern parallel machines increases to more than hundreds of thousands, scalable and consistent I/O performance is becoming hard to obtain due to fluctuating file system performance. This fluctuation is often caused by rebuilding RAID disk from hardware failures or concurrent jobs competing for I/O. We present a mechanism that stripes across a dynamically-selected subset of I/O servers with the lightest workload to achieve the best I/O bandwidth available from the system. We implement this mechanism into an I/O software layer that enables memory-to-file data layout transformation and allows transparent file partitioning. File partitioning is a technique that divides data among a set of files and manages file access, making data appear as a single file to users. Experimental results on NERSC's Hopper indicate that our approach effectively isolates I/O variation on shared systems and improves overall I/O performance significantly.
Seung Woo Son 0001, Saba Sehrish, Wei-keng Liao, Ron A. Oldfield, Alok N. Choudhary
CLUSTER3
2013 Forecast Oriented Classification of Spatio-Temporal Extreme Events
Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Kunpeng Zhang 0001, Ankit Agrawal 0001, Wei-keng Liao, Nagiza F. Samatova, Alok N. Choudhary
IJCAI6
2013 Improving collective I/O performance by pipelining request aggregation and file access
abstract
In this paper, we propose a multi-buffer pipelining approach to improve collective I/O performance by overlapping the dominant request aggregation phases with the I/O phase in the two-phase I/O implementation. Our pipelining method first divides the collective buffer into a group of small size buffers for an individual collective I/O call and then pipelines the asynchronous communication to exchange the I/O requests with the I/O requests sent to the file system. Our performance evaluation of a representative I/O benchmark and a production application shows 20% improvement in the I/O time, given theoretical upper bound of 50% when both phases completely overlap.
Saba Sehrish, Seung Woo Son 0001, Wei-keng Liao, Alok N. Choudhary, Karen Schuchardt
EuroMPI3
2013 Scalable parallel OPTICS data clustering using graph algorithmic techniques
abstract
OPTICS is a hierarchical density-based data clustering algorithm that discovers arbitrary-shaped clusters and eliminates noise using adjustable reachability distance thresholds. Parallelizing OPTICS is considered challenging as the algorithm exhibits a strongly sequential data access order. We present a scalable parallel OPTICS algorithm (Poptics) designed using graph algorithmic concepts. To break the data access sequentiality, POPTICS exploits the similarities between the OPTICS algorithm and Prim's Minimum Spanning Tree algorithm. Additionally, we use the disjoint-set data structure to achieve a high parallelism for distributed cluster extraction. Using high dimensional datasets containing up to a billion floating point numbers, we show scalable speedups of up to 27.5 for our OpenMP implementation on a 40-core shared-memory machine, and up to 3,008 for our MPI implementation on a 4,096-core distributed-memory machine. We also show that the quality of the results given by POPTICS is comparable to those given by the classical OPTICS algorithm.
Md. Mostofa Ali Patwary, Diana Palsetia, Ankit Agrawal 0001, Wei-keng Liao, Fredrik Manne, Alok N. Choudhary
SC4
2013 Fast Algorithms for the Maximum Clique Problem on Massive Sparse Graphs
Bharath Pattabiraman, Md. Mostofa Ali Patwary, Assefaw Hadish Gebremedhin, Wei-keng Liao, Alok N. Choudhary
WAW4
2012 Parallel hierarchical clustering on shared memory platforms
abstract
Hierarchical clustering has many advantages over traditional clustering algorithms like k-means, but it suffers from higher computational costs and a less obvious parallel structure. Thus, in order to scale this technique up to larger datasets, we present SHRINK, a novel shared-memory algorithm for single-linkage hierarchical clustering based on merging the solutions from overlapping sub-problems. In our experiments, we find that SHRINK provides a speedup of 18–20 on 36 cores on both real and synthetic datasets of up to 250,000 points. Source code for SHRINK is available for download on our website, http://cucis.ece.northwestern.edu.
William Hendrix, Md. Mostofa Ali Patwary, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
HiPC4
2012 A new scalable parallel DBSCAN algorithm using the disjoint-set data structure
abstract
DBSCAN is a well-known density based clustering algorithm capable of discovering arbitrary shaped clusters and eliminating noise data. However, parallelization of DBSCAN is challenging as it exhibits an inherent sequential data access order. Moreover, existing parallel implementations adopt a master-slave strategy which can easily cause an unbalanced workload and hence result in low parallel efficiency. We present a new parallel DBSCAN algorithm (PDSDBSCAN) using graph algorithmic concepts. More specifically, we employ the disjoint-set data structure to break the access sequentiality of DBSCAN. In addition, we use a tree-based bottom-up approach to construct the clusters. This yields a better-balanced workload distribution. We implement the algorithm both for shared and for distributed memory. Using data sets containing up to several hundred million high-dimensional points, we show that PDSDBSCAN significantly outperforms the master-slave approach, achieving speedups up to 25.97 using 40 cores on shared memory architecture, and speedups up to 5,765 using 8,192 cores on distributed memory architecture.
Md. Mostofa Ali Patwary, Diana Palsetia, Ankit Agrawal 0001, Wei-keng Liao, Fredrik Manne, Alok N. Choudhary
SC4
2012 Sentiment identification by incorporating syntax, semantics and context information
abstract
This paper proposes a method based on conditional random fields to incorporate sentence structure (syntax and semantics) and context information to identify sentiments of sentences within a document. It also proposes and evaluates two different active learning strategies for labeling sentiment data. The experiments with the proposed approach demonstrate a 5-15% improvement in accuracy on Amazon customer reviews compared to existing supervised learning and rule-based methods.
Kunpeng Zhang 0001, Yusheng Xie, Yu Cheng 0001, Daniel Honbo, Doug Downey, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
SIGIR7
2012 Accelerating pairwise statistical significance estimation for local alignment by harvesting GPU's power
abstract
BACKGROUND: Pairwise statistical significance has been recognized to be able to accurately identify related sequences, which is a very important cornerstone procedure in numerous bioinformatics applications. However, it is both computationally and data intensive, which poses a big challenge in terms of performance and scalability. RESULTS: We present a GPU implementation to accelerate pairwise statistical significance estimation of local sequence alignment using standard substitution matrices. By carefully studying the algorithm's data access characteristics, we developed a tile-based scheme that can produce a contiguous data access in the GPU global memory and sustain a large number of threads to achieve a high GPU occupancy. We further extend the parallelization technique to estimate pairwise statistical significance using position-specific substitution matrices, which has earlier demonstrated significantly better sequence comparison accuracy than using standard substitution matrices. The implementation is also extended to take advantage of dual-GPUs. We observe end-to-end speedups of nearly 250 (370) × using single-GPU Tesla C2050 GPU (dual-Tesla C2050) over the CPU implementation using Intel Corei7 CPU 920 processor. CONCLUSIONS: Harvesting the high performance of modern GPUs is a promising approach to accelerate pairwise statistical significance estimation for local sequence alignment.
Sanchit Misra, Ankit Agrawal 0001, Md. Mostofa Ali Patwary, Wei-keng Liao, Zhiguang Qin, Alok N. Choudhary
BMC Bioinform.5
2012 Delegation-Based I/O Mechanism for High Performance Computing Systems
abstract
Massively parallel applications often require periodic data checkpointing for program restart and post-run data analysis. Although high performance computing systems provide massive parallelism and computing power to fulfill the crucial requirements of the scientific applications, the I/O tasks of high-end applications do not scale. Strict data consistency semantics adopted from traditional file systems are inadequate for homogeneous parallel computing platforms. For high performance parallel applications independent I/O is critical, particularly if checkpointing data are dynamically created or irregularly partitioned. In particular, parallel programs generating a large number of unrelated I/O accesses on large-scale systems often face serious I/O serializations introduced by lock contention and conflicts at file system layer. As these applications may not be able to utilize the I/O optimizations requiring process synchronization, they pose a great challenge for parallel I/O architecture and software designs. We propose an I/O mechanism to bridge the gap between scientific applications and parallel storage systems. A static file domain partitioning method is developed to align the I/O requests and produce a client-server mapping that minimizes the file lock acquisition costs and eliminates the lock contention. Our performance evaluations of production application I/O kernels demonstrate scalable performance and achieve high I/O bandwidths.
Arifa Nisar, Wei-keng Liao, Alok N. Choudhary
IEEE Trans. Parallel Distributed Syst.2
2011 Mining millions of reviews: a technique to rank products based on importance of reviews
abstract
As online shopping becomes increasingly more popular, many shopping web sites encourage existing customers to add reviews of products purchased. These reviews make an impact on the purchasing decisions of potential customers. At Amazon.com for instance, some products receive hundreds of reviews. It is overwhelming and time restrictive for most customers to read, comprehend and make decisions based on all of these reviews. Customers most likely end up reading only a small fraction of the reviews usually in the order which they are presented on the product page. Incorporating various product review factors, such as: content related to product quality, time of the review, content related to product durability and historically older positive customer reviews will have different impacts on the products rankings. Thus, the automated mining of product reviews and opinions to produce a re-calculated product ranking score is a valuable tool which would allow potential customers to make more informed decisions. In this paper, we present a product ranking model that applies weights to product review factors to calculate a products ranking score. Our experiments use the customer reviews from Amazon.com as input to our product ranking model which produces product ranking results that closely relate to the products sales ranking as reported by the retailer.
Kunpeng Zhang 0001, Yu Cheng 0001, Wei-keng Liao, Alok N. Choudhary
ICEC3
2011 Supporting computational data model representation with high-performance I/O in parallel netCDF
abstract
Parallel computational scientific applications have been described by their computation and communication patterns. From a storage and I/O perspective, these applications can also be grouped into separate data models based on the way data is organized and accessed during simulation, analysis, and visualization. Parallel netCDF is a popular library used in many scientific applications to store scientific datasets and provides high-performance parallel I/O. Although the metadata-rich netCDF file format can effectively store and describe regular multi-dimensional array datasets, it does not address the full range of current and future computational science data models. In this paper, we present a new storage scheme in Parallel netCDF to represent a broad variety of data models used in modern computational scientific applications. This scheme also allows concurrent metadata construction for different data objects from multiple groups of application processes, an important feature in obtaining a high degree of I/O parallelism for data models exhibiting irregular data distribution. Furthermore, we employ non-blocking I/O functions to aggregate irregularly distributed data requests into large, contiguous data requests, to achieve high-performance I/O. Using an example of adaptive mesh refinement data model, we demonstrate the proposed scheme can produce scalable performance results for both data and metadata creation and access.
Kui Gao, Alok N. Choudhary, Wei-keng Liao
HiPC4
2011 Improving the Average Response Time in Collective I/O
Saba Sehrish, Wei-keng Liao, Alok N. Choudhary, Karen Schuchardt
EuroMPI3
2011 Anatomy of a hash-based long read sequence mapping algorithm for next generation DNA sequencing
abstract
MOTIVATION: Recently, a number of programs have been proposed for mapping short reads to a reference genome. Many of them are heavily optimized for short-read mapping and hence are very efficient for shorter queries, but that makes them inefficient or not applicable for reads longer than 200 bp. However, many sequencers are already generating longer reads and more are expected to follow. For long read sequence mapping, there are limited options; BLAT, SSAHA2, FANGS and BWA-SW are among the popular ones. However, resequencing and personalized medicine need much faster software to map these long sequencing reads to a reference genome to identify SNPs or rare transcripts. RESULTS: We present AGILE (AliGnIng Long rEads), a hash table based high-throughput sequence mapping algorithm for longer 454 reads that uses diagonal multiple seed-match criteria, customized q-gram filtering and a dynamic incremental search approach among other heuristics to optimize every step of the mapping process. In our experiments, we observe that AGILE is more accurate than BLAT, and comparable to BWA-SW and SSAHA2. For practical error rates (< 5%) and read lengths (200-1000 bp), AGILE is significantly faster than BLAT, SSAHA2 and BWA-SW. Even for the other cases, AGILE is comparable to BWA-SW and several times faster than BLAT and SSAHA2. AVAILABILITY: http://www.ece.northwestern.edu/~smi539/agile.html.
Sanchit Misra, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary
Bioinform.3
2011 Design and Evaluation of MPI File Domain Partitioning Methods under Extent-Based File Locking Protocol
abstract
MPI collective I/O has been an effective method for parallel shared-file access and maintaining the canonical orders of structured data in files. Its implementation commonly uses a two-phase I/O strategy that partitions a file into disjoint file domains, assigns each domain to a unique process, redistributes the I/O data based on their locations in the domains, and has each process perform I/O for the assigned domain. The partitioning quality determines the maximal performance achievable by the underlying file system, as the shared-file I/O has long been impeded by the cost of file system's data consistency control, particularly due to the conflicted locks. This paper proposes a few file domain partitioning methods designed to reduce lock conflicts under the extent-based file locking protocol. Experiments from four I/O benchmarks on the IBM GPFS and Lustre parallel file systems show that the partitioning method producing minimum lock conflicts wins the highest performance. The benefit of removing conflicted locks can be so significant that more than thirty times of write bandwidth differences are observed between the best and worst methods.
Wei-keng Liao
IEEE Trans. Parallel Distributed Syst.1
2010 Enabling active storage on parallel I/O software stacks
abstract
As data sizes continue to increase, the concept of active storage is well fitted for many data analysis kernels. Nevertheless, while this concept has been investigated and deployed in a number of forms, enabling it from the parallel I/O software stack has been largely unexplored. In this paper, we propose and evaluate an active storage system that allows data analysis, mining, and statistical operations to be executed from within a parallel I/O interface. In our proposed scheme, common analysis kernels are embedded in parallel file systems. We expose the semantics of these kernels to parallel file systems through an enhanced runtime interface so that execution of embedded kernels is possible on the server. In order to allow complete server-side operations without file format or layout manipulation, our scheme adjusts the file I/O buffer to the computational unit boundary on the fly. Our scheme also uses server-side collective communication primitives for reduction and aggregation using interserver communication. We have implemented a prototype of our active storage system and demonstrate its benefits using four data analysis benchmarks. Our experimental results show that our proposed system improves the overall performance of all four benchmarks by 50.9% on average and that the compute-intensive portion of the k-means clustering kernel can be improved by 58.4% through GPU offloading when executed with a larger computational load. We also show that our scheme consistently outperforms the traditional storage model with a wide variety of input dataset sizes, number of nodes, and computational loads.
Seung Woo Son 0001, Samuel Lang, Philip H. Carns, Robert B. Ross, Rajeev Thakur, Berkin Özisikyilmaz, Prabhat Kumar 0002, Wei-keng Liao, Alok N. Choudhary
MSST8
2010 Automated Tracing of I/O Stack
Seong Jo Kim, Seung Woo Son 0001, Ramya Prabhakar, Mahmut T. Kandemir, Christina M. Patrick, Wei-keng Liao, Alok N. Choudhary
EuroMPI7
2009 Combining I/O operations for multiple array variables in parallel netCDF
abstract
Parallel netCDF (PnetCDF) is a popular library used in many scientific applications to store scientific datasets. It provides high-performance parallel I/O while maintaining file-format compatibility with Unidata's netCDF. Array variables comprise the bulk of the data in a netCDF dataset, and for accesses to large regions of single array variables, PnetCDF attains very high performance. However, the current PnetCDF interface only allows access to one array variable per call. If an application instead accesses a large number of small-sized array variables, this interface limitation can cause significant performance degradation, because high end network and storage systems deliver much higher performance with larger request sizes. Moreover, the record variables data is stored interleaved by record, and the contiguity information is lost, so the existing MPI-IO collective I/O optimization can not help. This paper presents a new mechanism for PnetCDF to combine multiple I/O operations for better I/O performance. This mechanism can be used in a new function that takes arguments for reading/writing multiple array variables, allowing application programmers to explicitly access multiple array variables in a single call. It can also be used in the implementation of asynchronous I/O functions, so that the combination is carried out implicitly, without changes to the application. Our performance results demonstrate significant improvement using well-known application benchmarks.
Kui Gao, Wei-keng Liao, Alok N. Choudhary, Robert B. Ross, Robert Latham
CLUSTER2
2009 Detailed analysis of I/O traces for large scale applications
abstract
In this paper, we present a tool to extract I/O traces from very large applications running at full scale during their production runs. We analyze these traces to gain information about the application. We analyze the traces of three applications. The analysis showed that the I/O traces reveal much information about the application even without access to the source code. In particular, these I/O traces provide multiple indications towards the algorithmic nature of the application by observing the changes of data amount and I/O request distribution at the checkpoints. Adaptive Mesh Refinement (AMR) is one of the kind of algorithms that can exhibit such I/O behavior. This is the first study of I/O characteristics of unbalanced AMR-supported applications at scale. The key observations that we made in the trace were (1) Variation in aggregate data sizes across checkpoints for AMR and non-AMR applications, (2) Variation in the number of I/O calls by a client depending on the nature of the application, (3) Use of temporary files by applications and possible erroneous calls to I/O functions, (4) Variation in average data transfer size according as whether the application has AMR support or not, (5) Aggregation of I/O for processes executing on a single physical node through MPI-IO calls, and (6) Updates to specific data structures in the checkpoint file.
Nithin Nakka, Alok N. Choudhary, Wei-keng Liao, Lee Ward, Ruth Klundt, Marlow I. Weston
HiPC3
2009 Using Subfiling to Improve Programming Flexibility and Performance of Parallel Shared-file I/O
abstract
There are two popular parallel I/O programming styles used by modern scientific computational applications: unique-file and shared-file. Unique-file I/O usually gives satisfactory performance, but its major drawback is that managing a large number of files can overwhelm the task of post-simulation data processing. Shared-file I/O produces fewer files and allows arrays partitioned among processes to be saved in the canonical order. As the number of processors on modern parallel machines increases into thousands and more, the problem size and in turn the global array size also increase proportionally. It is not practical to manage files of size each larger than a few hundreds of GB. Hence, to seek a middle ground between these two I/O styles, we propose a subfiling scheme that divides a large multi-dimensional global array into smaller subarrays, each saved in a smaller file, named subfile. Subfiling is implemented on top of MPI-IO. We also incorporate it into the parallel netCDF library in order to preserve the partitioning information in the netCDF file header, so that the global array can later be reconstructed. In addition, since the subfiling scheme decreases the number of processes sharing a file, it can reduce the overhead of file system's data consistency control. Our experimental results with several I/O benchmarks show that subfiling can provide improved I/O performance.
Kui Gao, Wei-keng Liao, Arifa Nisar, Alok N. Choudhary, Robert B. Ross, Robert Latham
ICPP2
2009 High Performance Parallel/Distributed Biclustering Using Barycenter Heuristic
abstract
Biclustering refers to simultaneous clustering of objects and their features. Use of biclustering is gaining momentum in areas such as text mining, gene expression analysis and collaborative filtering. Due to requirements for high performance in large scale data processing applications such as Collaborative filtering in E-commerce systems and large scale genome-wide gene expression analysis in microarray experiments, a high performance prallel/distributed solution for biclustering problem is highly desirable. Recently, Ahmad et al [1] showed that Bipartite Spectral Partitioning, which is a popular technique for biclustering, can be reformulated as a graph drawing problem where objective is to minimize Hall's energy of the bipartite graph representation of the input data. They showed that optimal solution to this problem is achieved when nodes are placed at the barycenter of their neighbors. In this paper, we provide a parallel algorithm for biclustering based on this formulation. We show that parallel energy minimization using barycenter heuristic is embarrassingly parallel. The challenge is to design a bi-cluster identification algorithm which is scalable as well as accurate. We show that our parallel implementation is not just extremely scalable, it is comparable in accuracy as well with serial implementation. We have evaluated proposed parallel biclustering algorithm with large synthetic data sets on upto 256 processors. Experimental evaluation shows large superlinear speedups, scalability and high level of accuracy.
Arifa Nisar, Waseem Ahmad, Wei-keng Liao, Alok N. Choudhary
SDM3
2008 AHPIOS: An MPI-Based Ad Hoc Parallel I/O System
abstract
This paper presents the design and implementation of a portable ad-hoc parallel I/O system (AHPIOS). AHPIOS virtualizes on-demand available distributed storage resources and allows the files to be striped over several storage devices. Additionally, the design unifies the configuration of the MPI-IO library and the AHPIOS data servers. By a strong integration of the application, MPI-IO library and file system, a significant performance improvement can be achieved. The experimental section shows that the full MPI-IO integrated AHPIOS implementation of file access operations outperforms the existing MPI-IO implementation by as much as 495% for file writes and 522% for file reads.
Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Wei-keng Liao, Alok N. Choudhary
ICPADS4
2008 Dynamically adapting file domain partitioning methods for collective I/O based on underlying parallel file system locking protocols
abstract
Collective I/O, such as that provided in MPI-IO, enables process collaboration among a group of processes for greater I/O parallelism. Its implementation involves file domain partitioning, and having the right partitioning is a key to achieving high-performance I/O. As modern parallel file systems maintain data consistency by adopting a distributed file locking mechanism to avoid centralized lock management, different locking protocols can have significant impact to the degree of parallelism of a given file domain partitioning method. In this paper, we propose dynamic file partitioning methods that adapt according to the underlying locking protocols in the parallel file systems and evaluate the performance of four partitioning methods under two locking protocols. By running multiple I/O benchmarks, our experiments demonstrate that no single partitioning guarantees the best performance. Using MPI-IO as an implementation platform, we provide guidelines to select the most appropriate partitioning methods for various I/O patterns and file systems.
Wei-keng Liao, Alok N. Choudhary
SC1
2008 Scaling parallel I/O performance through I/O delegate and caching system
abstract
Increasingly complex scientific applications require massive parallelism to achieve the goals of fidelity and high computational performance. Such applications periodically offload checkpointing data to file system for post-processing and program resumption. As a side effect of high degree of parallelism, I/O contention at servers doesn't allow overall performance to scale with increasing number of processors. To bridge the gap between parallel computational and I/O performance, we propose a portable MPI-IO layer where certain tasks, such as file caching, consistency control, and collective I/O optimization are delegated to a small set of compute nodes, collectively termed as I/O Delegate nodes. A collective cache design is incorporated to resolve cache coherence and hence alleviates the lock contention at I/O servers. By using popular parallel I/O benchmark and application I/O kernels, our experimental evaluation indicates considerable performance improvement with a small percentage of compute resources reserved for I/O.
Arifa Nisar, Wei-keng Liao, Alok N. Choudhary
SC2
2007 Improving MPI Independent Write Performance Using A Two-Stage Write-Behind Buffering Method
abstract
Many large-scale production applications often have very long executions times and require periodic data checkpoints in order to save the state of the computation for program restart and/or tracing application progress. These write-only operations often dominate the overall application runtime, which makes them a good optimization target. Existing approaches for write-behind data buffering at the MPI I/O level have been proposed, but challenges still exist for addressing system-level I/O issues. We propose a two-stage write-behind buffering scheme for handing checkpoint operations. The first-stage of buffering accumulates write data for better network utilization and the second-stage of buffering enables the alignment for the write requests to the file stripe boundaries. Aligned I/O requests avoid file lock contention that can seriously degrade I/O performance. We present our performance evaluation using BTIO benchmarks on both GPFS and Lustre file systems. With the two-stage buffering, the performance of BTIO through MPI independent I/O is significantly improved and even surpasses that of collective I/O.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Mahmut T. Kandemir
IPDPS1
2007 An Implementation and Evaluation of Client-Side File Caching for MPI-IO
abstract
Client-side file caching has long been recognized as a file system enhancement to reduce the amount of data transfer between application processes and I/O servers. However, caching also introduces cache coherence problems when a file is simultaneously accessed by multiple processes. Existing coherence controls tend to treat the client processes independently and ignore the aggregate I/O access pattern. This causes a serious performance degradation for parallel I/O applications. In this paper we discuss our new implementation and present an extended performance evaluation on GPFS and Lustre parallel file systems. In addition to comparing our methods to traditional approaches, we examine the performance of MPI-IO caching under direct I/O mode to bypass the underlying file system cache. We also investigate the performance impact of two file domain partitioning methods to MPI collective I/O operations: one which creates a balanced workload and the other which aligns accesses to the file system stripe size. In our experiments, alignment results in better performance by reducing file lock contention. When the cache page size is set to a multiple of the stripe size, MPI-IO caching inherits the same advantage and produces significantly improved I/O bandwidth.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Lee Ward
IPDPS1
2007 Noncontiguous locking techniques for parallel file systems
abstract
Many parallel scientific applications use high-level I/O APIs that offer atomic I/O capabilities. Atomic I/O in current parallel file systems is often slow when multiple processes simultaneously access interleaved, shared files. Current atomic I/O solutions are not optimized for handling noncontiguous access patterns because current locking systems have a fixed file system block-based granularity and do not leverage high-level access pattern information.
Avery Ching, Wei-keng Liao, Alok N. Choudhary, Robert B. Ross, Lee Ward
SC2
2007 Using MPI file caching to improve parallel write performance for large-scale scientific applications
abstract
Typical large-scale scientific applications periodically write checkpoint files to save the computational state throughout execution. Existing parallel file systems improve such write-only I/O patterns through the use of client-side file caching and write-behind strategies. In distributed environments where files are rarely accessed by more than one client concurrently, file caching has achieved significant success; however, in parallel applications where multiple clients manipulate a shared file, cache coherence control can serialize I/O. We have designed a thread based caching layer for the MPI I/O library, which adds a portable caching system closer to user applications so more information about the application’s I/O patterns is available for better coherence control. We demonstrate the impact of our caching solution on parallel write performance with a comprehensive evaluation that includes a set of widely used I/O benchmarks and production application I/O kernels. 1.
Wei-keng Liao, Avery Ching, Kenin Coloma, Arifa Nisar, Alok N. Choudhary, Jacqueline Chen, Ramanan Sankaran, Scott Klasky
SC1
2006 A New Flexible MPI Collective I/O Implementation
abstract
The MPI-IO standard creates a huge opportunity to break out of the traditional file system I/O methods. As a software layer between the user and the file system, an MPI-IO library can potentially optimize I/O on behalf of the user with little to no user intervention. This is all possible because of the rich data description and communication infrastructure MPI-2 offers. Powerful data descriptions and some of the other desirable features of MPI-2, however, make MPI-IO challenging to implement. By creating a new collective I/O implementation that allows developers to easily tinker and play with new optimizations or combinations of different techniques, research can proceed faster and be quickly and reliably deployed
Kenin Coloma, Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, Rajeev Thakur, Lee Ward
CLUSTER4
2006 Evaluating I/O characteristics and methods for storing structured scientific data
abstract
Many large-scale scientific simulations generate large, structured multi-dimensional datasets. Data is stored at various intervals on high performance I/O storage systems for checkpointing, post-processing, and visualization. Data storage is very I/O intensive and can dominate the overall running time of an application, depending on the characteristics of the I/O access pattern. Our NCIO benchmark determines how I/O characteristics greatly affect performance (up to 2 orders of magnitude) and provides scientific application developers with guidelines for improvement. In this paper, we examine the impact of various I/O parameters and methods when using the MPI-IO interface to store structured scientific data in an optimized parallel file system.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Lee Ward, Neil Pundit
IPDPS3
2006 Mining Frequent Patterns by Differential Refinement of Clustered Bitmaps
abstract
Existing algorithms for mining frequent patterns are facing challenges to handle databases (a) of increasingly large sizes, (b) consisting of variable-length, irregularly-spaced data, and (c) with mixed or even unknown properties. In this paper, we propose a novel self-adaptive algorithm D-CLUB that thoroughly addresses these issues by progressively clustering the database into condensed association bitmaps, applying a differential technique to digest and remove dense patterns, and then mining the remaining tiny bitmaps directly through fast aggregate bit operations. The bitmaps are well organized into rectangular two-dimensional matrices and adaptively refined in regions that necessitate further computation. We show that this approach not only drastically cuts down the original database size but also largely reduces and simplifies the mining computation for a wide variety of datasets and parameters. We compare D-CLUB with various state-of-the-art algorithms and show significant performance improvement in all cases.
Alok N. Choudhary, Wei-keng Liao
SDM4
2006 Distributed smart disks for I/O-intensive workloads on switched interconnects
Steve C. Chiu, Wei-keng Liao, Alok N. Choudhary
Future Gener. Comput. Syst.2
2006 Multicollective I/O: A technique for exploiting inter-file access patterns
abstract
The increasing gap between processor cycle times and access times to storage devices makes it necessary to use powerful optimizations. This is especially true for applications in the parallel computing domain that frequently perform large amounts of file I/O. Collective I/O strategy that coordinates the processes to perform I/O on each other's behalf has demonstrated a significant performance improvement. This article proposes a new concept called Multicollective I/O (MCIO) that expands the collective I/O to allow data from multiple files to be requested in a single I/O request, in contrast to allowing only multiple segments for a single file to be specified together. MCIO considers multiple arrays simultaneously by having a more global view of the overall I/O behavior exhibited by parallel applications. This article shows that determining the optimal MCIO access pattern is an NP-complete problem, and proposes two different heuristics for the access pattern detection problem, also called the assignment problem. Both heuristics have been implemented within a runtime library, and tested using a large-scale scientific application. Our results show that MCIO outperforms collective I/O by as much as 87%. Our runtime library-based implementation can be used by application users as well as by optimizing compilers. Based on our results, we recommend that future library designers for I/O-intensive applications include MCIO in their suite of optimizations.
Gokhan Memik, Mahmut T. Kandemir, Wei-keng Liao, Alok N. Choudhary
ACM Trans. Storage3
2006 Scalable Design and Implementations for MPI Parallel Overlapping I/O
abstract
We investigate the Message Passing Interface Input/Output (MPI I/O) implementation issues for two overlapping access patterns: the overlaps among processes within a single I/O operation and the overlaps across a sequence of I/O operations. The former case considers whether I/O atomicity can be obtained in the overlapping regions. The latter focuses on the file consistency problem on parallel machines with client-side file caching enabled. Traditional solutions for both overlapping I/O problems use whole file or byte-range file locking to ensure exclusive access to the overlapping regions and bypass the file system cache. Unfortunately, not only can file locking serialize I/O, but it can also increase the aggregate communication overhead between clients and I/O servers. For atomicity, we first differentiate MPI's requirements from the Portable Operating System Interface (POSIX) standard and propose two scalable approaches, graph coloring and process-rank ordering, which can resolve access conflicts and maintain I/O parallelism. For solving the file consistency problem across multiple I/O operations, we propose a method called Persistent File Domains, which tackles cache coherency with additional information and coordination to guarantee safe cache access without using file locks.
Wei-keng Liao, Kenin Coloma, Alok N. Choudhary, Lee Ward, Eric Russell, Neil Pundit
IEEE Trans. Parallel Distributed Syst.1
2005 Collective caching: application-aware client-side file caching
abstract
Parallel file subsystems in today's high-performance computers adopt many I/O optimization strategies that were designed for distributed systems. These strategies, for instance client-side file caching, treat each I/O request process independently, due to the consideration that clients are unlikely related with each other in a distributed environment. However, it is inadequate to apply such strategies directly in the high-performance computers where most of the I/O requests come from the processes that work on the same parallel applications. We believe that client-side could perform more effectively if the subsystem is aware of the process scope of an application and regards all the application processes as a single client. In this paper, we propose the idea of caching which coordinates the application processes to manage cache data and achieve cache coherence without involving the I/O servers. To demonstrate this idea, we implemented a collective subsystem at user space as a library, which can be incorporated into any message passing interface implementation to increase its portability. The performance evaluation is presented with three I/O benchmarks on an IBM SP using its native parallel file system, GPFS. Our results show significant performance enhancement obtained by collective over the traditional approaches.
Wei-keng Liao, Kenin Coloma, Alok N. Choudhary, Lee Ward, Eric Russell, Sonja Tideman
HPDC1
2005 Design and Evaluation of Database Layouts for MEMS-Based Storage Systems
abstract
MEMS-based storage systems have recently generated significant interest due to their potential to be faster and more efficient than disks, while providing the non-volatility property. Designing data layouts for these devices is a challenging, important and interesting problem. In this paper, we explore various ways of placing a database on a MEMS-based storage architecture. Three novel data layouts are proposed after considering the MEMS device characteristics and the access patterns arising from queries. We then design the access methodology for each layout and evaluate these layouts based on their respective I/O service times. Overall, our results were able to identify the intricacies of placing data on a MEMS-based storage and also ascertain the large potential of MEMS-based devices for databases.
Jayaprakash Pisharath, Wei-keng Liao, Alok N. Choudhary
IDEAS2
2005 A Two-Phase Algorithm for Fast Discovery of High Utility Itemsets
Ying Liu 0039, Wei-keng Liao, Alok N. Choudhary
PAKDD2
2005 Processor-embedded distributed smart disks for I/O-intensive workloads: architectures, performance models and evaluation
Steve C. Chiu, Wei-keng Liao, Alok N. Choudhary, Mahmut T. Kandemir
J. Parallel Distributed Comput.2
2005 Performance Evaluation of a Parallel Pipeline Computational Model for Space-Time Adaptive Processing
Wei-keng Liao, Alok N. Choudhary, Donald Weiner, Pramod K. Varshney
J. Supercomput.1
2004 Processor-Embedded Distributed MEMS-Based Storage Systems for High-Performance I/O
abstract
Summary form only given. Built upon new data organization and access characteristics, MEMS-based storage devices have come under consideration as an alternative to disks for large data-intensive applications. While not already in commercial production, MEMS-based storage devices have outperformed disks in device-level simulations. Processor-embedded distributed disks improved performance of workloads by offloading application-level processing to the storage. To exploit the potential benefits offered by these emerging storage technologies and offloading models, we propose a processor-embedded distributed MEMS-based storage architecture, and evaluate the proposed architecture with representative database and data mining workloads. Our results show that MEMS-based storage improved the overall performance of these workloads over disk-based systems, and transformed the characteristics of several workloads, impacting the design points for future storage architectures.
Steve C. Chiu, Wei-keng Liao, Alok N. Choudhary
IPDPS2
2004 Scalable High-level Caching for Parallel I/O
abstract
Summary form only given. In order for I/O systems to achieve high performance in a parallel environment, they must either sacrifice client-side file caching, or keep caching and deal with complex coherency issues. The most common technique for dealing with cache coherency in multiclient file caching environments uses file locks to bypass the client-side cache. Aside from effectively disabling cache usage, file locking is sometimes unavailable on larger systems. The high-level abstraction layer of MPI allows us to tackle cache coherency with additional information and coordination without using file locks. By approaching the cache coherency issue further up, the underlying I/O accesses can be modified in such a way as to ensure access to coherent data while satisfying the user's I/O request. We can effectively exploit the benefits of a file system's client-side cache while minimizing its management costs.
Kenin Coloma, Alok N. Choudhary, Wei-keng Liao, Lee Ward, Eric Russell, Neil Pundit
IPDPS3
2004 Processor-embedded distributed smart disks for I/O-intensive workloads: architectures, performance models and evaluation
Steve C. Chiu, Wei-keng Liao, Alok N. Choudhary, Mahmut T. Kandemir
J. Parallel Distributed Comput.2
2003 Noncontiguous I/O Accesses Through MPI-IO
abstract
I/O performance remains a weakness of parallel computing systems today. While this weakness is partly attributed to rapid advances in other system components, I/O interfaces available to programmers and the I/O methods supported by file systems have traditionally not matched efficiently with the types of I/O operations that scientific applications perform, particularly noncontiguous accesses. The MPI-IO interface allows for rich descriptions of the I/O patterns desired for scientific applications and implementations such as ROMIO have taken advantage of this ability while remaining limited by underlying file system methods. A method of noncontiguous data access, list I/O, was recently implemented in the Parallel Virtual File System (PVFS). We implement support for this interface in the ROMIO MPI-IO implementation. Through a suite of noncontiguous I/O tests we compared ROMIO list I/O to current methods of ROMIO noncontiguous access and found that the list I/O interface provides performance benefits in many noncontiguous cases.
Avery Ching, Alok N. Choudhary, Kenin Coloma, Wei-keng Liao, Robert B. Ross, William Gropp
CCGRID4
2003 Efficient Structured Data Access in Parallel File Systems
abstract
Parallel scientific applications store and retrieve very large, structured datasets. Directly supporting these structured accesses is an important step in providing high-performance I/O solutions for these applications. High-level interfaces such as HDF5 and Parallel netCDF provide convenient APIs for accessing structured datasets, and the MPI-IO interface also supports efficient access to structured data. However, parallel file systems do not traditionally support such access. In this work we present an implementation of structured data access support in the context of the parallel virtual file system (PVFS). We call this support "datatype I/O" because of its similarity to MPI datatypes. This support is built by using a reusable datatype-processing component from the MPICH2 MPI implementation. We describe how this component is leveraged to efficiently process structured data representations resulting from MPI-IO operations. We quantitatively assess the solution using three test applications. We also point to further optimizations in the processing path that could be leveraged for even more efficient operation.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, William Gropp
CLUSTER3
2003 Scalable Implementations of MPI Atomicity for Concurrent Overlapping I/O
abstract
For concurrent I/O operations, atomicity defines the results in the overlapping file regions simultaneously read/written by requesting processes. Atomicity has been well studied at the file system level, such as POSIX standard. We investigate the problems arising from the implementation of MPI atomicity for concurrent overlapping write access and provide two programming solutions. Since the MPI definition of atomicity differs from the POSIX one, an implementation that simply relies on the POSIX file systems does not guarantee correct MPI semantics. To have a correct implementation of atomic I/O in MPI, we examine the efficiency of three approaches: I) file locking, 2) graph-coloring, and 3) process-rank ordering. Performance complexity for these methods are analyzed and their experimental results are presented for file systems including NFS, SGI's XFS, and IBM's GPFS.
Wei-keng Liao, Alok N. Choudhary, Kenin Coloma, George K. Thiruvathukal, Lee Ward, Eric Russell, Neil Pundit
ICPP1
2003 Parallel netCDF: A High-Performance Scientific I/O Interface
abstract
Dataset storage, exchange, and access play a critical role in scientific applications. For such purposes netCDF serves as a portable, efficient file format and programming interface, which is popular in numerous scientific application domains. However, the original interface does not provide an efficient mechanism for parallel data storage and access. In this work, we present a new parallel interface for writing and reading netCDF datasets. This interface is derived with minimal changes from the serial netCDF interface but defines semantics for parallel access and is tailored for high performance. The underlying parallel I/O is achieved through MPI-IO, allowing for substantial performance gains through the use of collective I/O optimizations. We compare the implementation strategies and performance with HDF5. Our tests indicate programming convenience and significant I/O performance improvement with this parallel netCDF (PnetCDF) interface.
Wei-keng Liao, Alok N. Choudhary, Robert B. Ross, Rajeev Thakur, William Gropp, Robert Latham, Andrew R. Siegel, Brad Gallagher, Michael Zingale
SC2
2003 A high-performance application data environment for large-scale scientific computations
abstract
Effective high-level data management is becoming an important issue with more and more scientific applications manipulating huge amounts of secondary-storage and tertiary-storage data using parallel processors. A major problem facing the current solutions to this data management problem is that these solutions either require a deep understanding of specific data storage architectures and file layouts to obtain the best performance (as in high-performance storage management systems and parallel file systems), or they sacrifice significant performance in exchange for ease-of-use and portability (as in traditional database management systems). We discuss the design, implementation, and evaluation of a novel application development environment for scientific computations. This environment includes a number of components that make it easy for the programmers to code and run their applications without much programming effort and, at the same time, to harness the available computational and storage power on parallel architectures.
Xiaohui Shen, Wei-keng Liao, Alok N. Choudhary, Gokhan Memik, Mahmut T. Kandemir
IEEE Trans. Parallel Distributed Syst.2
2002 Noncontiguous I/O through PVFS
abstract
With the tremendous advances in processor and memory technology, I/O has risen to become the bottleneck in high-performance computing for many applications. The development of parallel file systems has helped to ease the performance gap, but I/O still remains an area needing significant performance improvement. Research has found that noncontiguous I/O access patterns in scientific applications combined with current file system methods, to perform these accesses lead to unacceptable performance for large data sets. To enhance performance of noncontiguous I/O, we have created list I/O, a native version of noncontiguous I/O. We have used the Parallel Virtual File System (PVFS) to implement our ideas. Our research and experimentation shows that list I/O outperforms current noncontiguous I/O access methods in most I/O situations and can substantially enhance the performance of real-world scientific applications.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, William Gropp
CLUSTER3
2002 I/O Analysis and Optimization for an AMR Cosmology Application
abstract
In this paper we investigate the data access patterns and file I/O behaviors of a production cosmology application that uses the adaptive mesh refinement (AMR) technique for its domain decomposition. This application was originally developed using Hierarchical Data Format (HDF version 4) I/O library and since HDF4 does not provide parallel I/O facilities, the global file I/O operations were carried out by one of the allocated processors. When the number of processors becomes large, the I/O performance of this design degrades significantly due to the high communication cost and sequential file access. In this work, we present two additional I/O implementations, using MPI-IO and parallel HDF version 5, and analyze their impacts to the I/O performance for this typical AMR application. Based on the I/O patterns discovered in this application, we also discuss the interaction between user level parallel I/O operations and different parallel file systems and point out the advantages and disadvantages. The performance results presented in this work are obtained from an SGI Origin2000 using XFS, an IBM SP using GPFS, and a Linux cluster using PVFS.
Wei-keng Liao, Alok N. Choudhary, Valerie Taylor 0001
CLUSTER2
2001 An Integrated Graphical User Interface for High Performance Distributed Computing
abstract
It is very common that modern large-scale scientific applications employ multiple compute and storage resources in a heterogeneously distributed environment. Working effectively and efficiently in such an environment is one of the major concerns for designing meta-data management systems. The authors present an integrated graphical user interface (GUI) that makes the entire environment virtually an easy-to-use control platform for managing complex programs and their large datasets. To hide the I/O latency when the the user carries out interactive visualization, aggressive prefetching and caching techniques are employed in our GUI. The performance numbers show that the design of our Java GUI has achieved the goals of both high performance and ease-of-use.
Xiaohui Shen, Wei-keng Liao, Alok N. Choudhary
IDEAS2
2000 Meta-data Management System for High-Performance Large-Scale Scientific Data Access
Wei-keng Liao, Xiaohui Shen, Alok N. Choudhary
HiPC1
2000 A novel application development environment for large-scale scientific computations
abstract
Our results demonstrate that our novel application development environment provides both ease-of-use and high performance for large-scale, I/O-intensive scientific applications.
Xiaohui Shen, Wei-keng Liao, Alok N. Choudhary, Gokhan Memik, Mahmut T. Kandemir, Sachin More, George K. Thiruvathukal, Arti Singh
ICS2
2000 Design and Evaluation of I/O Strategies for Parallel Pipelined STAP Applications
abstract
This paper presents experimental results for a parallel pipeline STAP system with I/O task implementation using the parallel file systems on the Intel Paragon and the IBM SP. In our previous work, a parallel pipeline model was designed for radar signal processing applications on parallel computers. Based on this model, we implemented a real STAP application which demonstrated the performance scalability of this model in terms of throughput and latency. In this paper we study the effect on system performance when the I/O task is incorporated in the parallel pipeline model. There are two alternative for I/O implementation: embedding I/O in the pipeline or having a separate I/O task. From these two I/O implementations, we discovered that the latency may be improved when the structure of the pipeline is reorganized by merging multiple tasks into a single task. All the performance results shown in this paper demonstrated the scalability of parallel I/O implementation on the parallel pipeline STAP system.
Wei-keng Liao, Alok N. Choudhary, Donald Weiner, Pramod K. Varshney
IPDPS1
1999 I/O Implementation and Evaluation of Parallel Pipelined STAP on High Performance Computers
Wei-keng Liao, Alok N. Choudhary, Donald Weiner, Pramod K. Varshney
HiPC1