VLDB 2026 Research / reviewers in the wild / expert
Jong Choi 0001
dblp:54/4861 · also Jong Youl Choi
· DBLP profile ↗
51ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-6459-6152ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 1 since 2021Security and privacy · 5 · 1 first-authorArtificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Concurrent GHZ State Distribution in Quantum Networks
Qing Cao 0001, Weisheng Si, Jong Choi 0001, Sajal K. Das 0001 |
CCGrid | 3 |
| 2026 | Scalable Hybrid Learning Techniques for Scientific Data CompressionabstractData compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). This paper presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios. Tania Banerjee, Jong Choi 0001, Jaemoon Lee, Qian Gong, Jieyang Chen, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Scaling Laws of Graph Neural Networks for Atomistic Materials ModelingabstractAtomistic materials modeling is a critical task with wide-ranging applications, from drug discovery to materials science, where accurate predictions of the target material property can lead to significant advancements in scientific discovery. Graph Neural Networks (GNNs) represent the state-of-the-art approach for modeling atomistic material data thanks to their capacity to capture complex relational structures. While machine learning performance has historically improved with larger models and datasets, GNNs for atomistic materials modeling remain relatively small compared to large language models (LLMs), which leverage billions of parameters and terabyte-scale datasets to achieve remarkable performance in their respective domains. To address this gap, we explore the scaling limits of GNNs for atomistic materials modeling by developing a foundational model with billions of parameters, trained on extensive datasets in terabytescale. Our approach incorporates techniques from LLM libraries to efficiently manage large-scale data and models, enabling both effective training and deployment of these large-scale GNN models. This work addresses three fundamental questions in scaling GNNs: the potential for scaling GNN model architectures, the effect of dataset size on model accuracy, and the applicability of LLM-inspired techniques to GNN architectures. Specifically, the outcomes of this study include (1) insights into the scaling laws for GNNs, highlighting the relationship between model size, dataset volume, and accuracy, (2) a foundational GNN model optimized for atomistic materials modeling, and (3) a GNN codebase enhanced with advanced LLM-based training techniques. Our findings lay the groundwork for large-scale GNNs with billions of parameters and terabyte-scale datasets, establishing a scalable pathway for future advancements in atomistic materials modeling. Chaojian Li, Zhifan Ye, Massimiliano Lupo Pasini, Jong Choi 0001, Cheng Wan 0005, Prasanna Balaprakash |
DAC | 4 |
| 2025 | Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNNabstractWe present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier. Massimiliano Lupo Pasini, Jong Choi 0001, Kshitij Mehta, David M. Rogers 0001, Jonghyun Bae, Khaled Z. Ibrahim, Ashwin M. Aji, Karl W. Schulz, Jorda Polo, Prasanna Balaprakash |
J. Supercomput. | 2 |
| 2025 | Lustre Unveiled: Evolution, Design, Advancements, and Current TrendsabstractThe Lustre filesystem serves as a vital element in high-performance parallel storage, meeting the rising demands of scientific, research, and enterprise environments. Widely deployed across HPC environments, ranging from small-scale applications in AI/ML, to domains like oil and gas, drug discovery, and meteorology, and manufacturing, Lustre addresses the universal challenge of efficiently accessing vast and ever-increasing volumes of data. Lustre is the filesystem of choice on six out of the top 10 fastest supercomputers in the world today, over 65% of the top 100, and also for over 60% of the top 500. Despite its widespread popularity, there is a lack of a complete and up-to-date reference, covering Lustre’s evolution, design, and various advancements made over the years. In this journal, we aim to fill this gap by providing a comprehensive journey of Lustre, including its history with significant contributions to HPC, detailed architecture and design elements, exploration of advancements added through its evolution, and future directions. Additionally, we present a comparison of Lustre with other prominent storage technologies of the era. To illustrate the current state of Lustre, we analyze several filesystem trends, including utilization, performance, and usage patterns on Orion, the Lustre filesystem on the first exascale supercomputer Frontier. We hope that this journal serves as a comprehensive educational reference for the current and future generations interested in HPC filesystem storage aspects. Anjus George, Andreas Dilger, Michael J. Brim, Rick Mohr, Amir Shehata, Jong Choi 0001, Ahmad Maroof Karimi, Jesse Hanley, James Simmons, Dominic Manno, Verónica G. Vergara Larrea, Sarp Oral, Christopher Zimmer 0001 |
ACM Trans. Storage | 6 |
| 2023 | Online and Scalable Data Compression Pipeline with Guarantees on Quantities of InterestabstractData compression is becoming critical for data-intensive scientific applications. Scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Prior work has shown that a pipeline can be built to guarantee error on the primary data (PD) within user-defined bounds and achieve near-floating point QoI errors. In this paper, we present novel computational approaches for accelerating the pipeline and demonstrate results that enable concurrent execution of compression in parallel with the simulation nodes. This allows compression, including the writing of the required compression data, for the previous time step to be completed while the simulation proceeds with the current time step. Overall, the approach presented in this paper results in a 6–8 times improvement in computational overhead compared to previous work. These results were obtained using data generated by a large-scale fusion code called XGC, which produces hundreds of terabytes of data in a single day. Tania Banerjee, Jaemoon Lee, Jong Choi 0001, Qian Gong, Jieyang Chen, Choong-Seock Chang, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
e-Science | 3 |
| 2023 | Fast Algorithms for Scientific Data CompressionabstractMany scientific simulations and experiments generate terabytes to petabytes of data daily, necessitating data compression techniques. Unlike video and image compression, scientists require methods that accurately preserve primary data (PD) and derived quantities of interest (QoIs). In our previous work, we demonstrated the effectiveness of hybrid compression techniques that combine machine learning with traditional approaches. This paper presents innovative computational techniques aimed at expediting the compression pipeline. Our experiments, conducted on two distinct platforms with a large-scale XGC-based fusion simulation, demonstrate that the overhead incurred by these new approaches is less than one percent of the computational resources needed for the simulation. Tania Banerjee, Jaemoon Lee, Jong Choi 0001, Qian Gong, Jieyang Chen, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
HiPC | 3 |
| 2023 | Analyzing File Access Patterns on Large-Scale HPC Systems: Opportunities for File PrefetchingabstractThis paper explores the potential opportunities for implementing file prefetching techniques on large-scale high-performance computing (HPC) systems. Specifically, we investigate the file access patterns of various applications across multiple scientific domains using two years' worth of Darshan I/O traces obtained from the Summit supercomputer. We identify recurring trends and patterns which indicate that prefetching can be effectively leveraged to improve data access performance on HPC systems. This study serves as a valuable reference for system architects and developers in the HPC community, providing insights into the opportunities and challenges associated with enabling file prefetching on large-scale HPC systems. Ahmad Maroof Karimi, Arnab Kumar Paul, Jong Choi 0001, Lipeng Wan 0001, Feiyi Wang |
MASCOTS | 3 |
| 2022 | Hybrid Analysis of Fusion Data for Online Understanding of Complex Science on Extreme Scale ComputersabstractThe current practice for fusion scientists running first principle simulations on high performance computing plat-forms is to either run their simulations and output their data for post-hoc analysis, or to place in situ analytics into their code. In this paper we examine a complex workflow using XGC fusions simulation run on the Oak Ridge Leadership Computing Facility's supercomputer Summit, which also involve three anal-yses as part of the results necessary for scientific discovery. We discuss the challenges faced when implementing these algorithms and present an original hybrid staging technique to help enable the physicists to make discoveries during the execution of the simulation. By creating this infrastructure, we can examine complicated physics results, which may not have been possible without the infrastructure. For example, our work enables the online visualization of turbulent homoclinic tangle around the magnetic X-point, breaking the last confinement surface. This visualization could help fusion scientists to better understand and improve the turbulence spread of plasma exhaust heat, which is crucial toward realizing plasmas beyond the currently accessible physics regimes of present-day tokamak reactors. The physics of turbulent homoclinic tangle will be reported in a future physics publication, by utilizing the original online analysis/visualization framework presented in this paper. Eric Suchyta, Jong Choi 0001, Seung-Hoe Ku, David Pugmire, Ana Gainaru, Kevin A. Huck, Ralph Kube, Aaron Scheinberg, Frédéric Suter, Choong-Seock Chang, Todd S. Munson, Norbert Podhorszki, Scott Klasky |
CLUSTER | 2 |
| 2022 | An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error GuaranteesabstractEfficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists. Tania Banerjee, Jong Choi 0001, Jaemoon Lee, Qian Gong, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka |
HIPC | 2 |
| 2022 | Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage SystemsabstractMonitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue. Arnab Kumar Paul, Jong Choi 0001, Ahmad Maroof Karimi, Feiyi Wang |
HPDC | 2 |
| 2022 | A codesign framework for online data analysis and reductionabstractAbstract Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation‐analysis‐reduction pipelines coupled in‐memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application‐ and system‐specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. This toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next‐generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers. Kshitij Mehta, Bryce Allen, Matthew Wolf, Jeremy Logan, Eric Suchyta, Swati Singhal, Jong Choi 0001, Keichi Takahashi, Kevin A. Huck, Igor Yakushin, Alan Sussman, Todd S. Munson, Ian T. Foster, Scott Klasky |
Concurr. Comput. Pract. Exp. | 7 |
| 2021 | Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUsabstractRapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput-83% of theoretical peak-on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software. Jieyang Chen, Lipeng Wan 0001, Xin Liang 0001, Ben Whitney, Qing Liu 0002, David Pugmire, Nicholas Thompson, Jong Choi 0001, Matthew Wolf, Todd S. Munson, Ian T. Foster, Scott Klasky |
IPDPS | 8 |
| 2020 | Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and ImplicationsabstractThis article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries. We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications. In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable. Sarp Oral, Christopher Zimmer 0001, Jong Choi 0001, David Dillow, Scott Klasky, Jay F. Lofstead, Norbert Podhorszki, Jeffrey S. Chase |
ACM Trans. Storage | 4 |
| 2019 | Scalable Performance Awareness for In Situ Scientific ApplicationsabstractPart of the promise of exascale computing and the next generation of scientific simulation codes is the ability to bring together time and spatial scales that have traditionally been treated separately. This enables creating complex coupled simulations and in situ analysis pipelines, encompassing such things as "whole device" fusion models or the simulation of cities from sewers to rooftops. Unfortunately, the HPC analysis tools that have been built up over the preceding decades are ill suited to the debugging and performance analysis of such computational ensembles. In this paper, we present a new vision for performance measurement and understanding of HPC codes, MonitoringAnalytics (MONA). MONA is designed to be a flexible, high performance monitoring infrastructure that can perform monitoring analysis in place or in transit by embedding analytics and characterization directly into the data stream, without relying upon delivering all monitoring information to a central database for post-processing. It addresses the trade-offs between the prohibitively expensive capture of all performance characteristics and not capturing enough to detect the features of interest. We demonstrate several uses of MONA; capturing and indexing multi-executable performance profiles to enable later processing, extraction of performance primitives to enable the generation of customizable benchmarks and performance skeletons, and extracting communication and application behaviors to enable better control and placement for the current and future runs of the science ensemble. Relevant performance information based on a system for MONA built from ADIOS and SOSflow technologies is provided for DOE science applications and leadership machines. Matthew Wolf, Julien Dominski, Gabriele Merlo, Jong Choi 0001, Greg Eisenhauer, Stéphane Ethier, Kevin A. Huck, Scott Klasky, Jeremy Logan, Allen D. Malony, Chad Wood |
eScience | 4 |
| 2019 | Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems?abstractFor a production high-performance computing (HPC) system, where storage devices are shared between multiple applications and managed in a best effort manner, I/O contention is often a major problem. In this paper, we propose a balanced messaging-based re-routing in conjunction with throttling at the middleware level. This work tackles two key challenges that have not been fully resolved in the past: whether I/O variability can be reduced on a QoS-less HPC storage system, and how to design a runtime scheduling system that can scale up to a large amount of cores. The proposed scheme uses a two-level messaging system to re-route I/O requests to a less congested storage location so that write performance is improved, while limiting the impact on read by throttling re-routing. An analytical model is derived to guide the setup of optimal throttling factor. We thoroughly analyze the virtual messaging layer overhead and explore whether the in-transit buffering is effective in managing I/O variability. Contrary to the intuition, in-transit buffer cannot completely solve the problem. It can reduce the absolute variability but not the relative variability. The proposed scheme is verified against a synthetic benchmark as well as being used by production applications. Dan Huang 0001, Qing Liu 0002, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Jeremy Logan, George Ostrouchov, Xubin He, Matthew Wolf |
IEEE Trans. Computers | 3 |
| 2019 | Harnessing Data Movement in Virtual Clusters for In-Situ ExecutionabstractAs a result of increasing data volume and velocity, Big Data science at exascale has shifted towards the in-situ paradigm, where large scale simulations run concurrently alongside data analytics. With in-situ, data generated from simulations can be processed while still in memory, thereby avoiding the slow storage bottleneck. However, running simulations and analytics together on shared resources will likely result in substantial contention if left unmanaged, as demonstrated in this work, leading to much reduced efficiency of simulations and analytics. Recently, virtualization technologies such as Linux containers have been widely applied to data centers and physical clusters to provide highly efficient and elastic resource provisioning for consolidated workloads including scientific simulations and data analytics. In this paper, we investigate to facilitate network traffic manipulation and reduce mutual interference on the network for in-situ applications in virtual clusters. In order to dynamically allocate the network bandwidth when it is needed, we adopt SARIMA-based techniques to analyze and predict MPI traffic issued from simulations. Although this can be an effective technique, the naïve usage of network virtualization can lead to performance degradation for bursty asynchronous transmissions within an MPI job. We analyze and resolve this performance degradation in virtual clusters. Dan Huang 0001, Qing Liu 0002, Scott Klasky, Jun Wang 0001, Jong Choi 0001, Jeremy Logan, Norbert Podhorszki |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 1 |
| 2018 | A View from ORNL: Scientific Data Research Opportunities in the Big Data AgeabstractOne of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1]. Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001 |
ICDCS | 5 |
| 2018 | Understanding and Modeling Lossy Compression Schemes on HPC Scientific DataabstractScientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions. Tao Lu 0014, Qing Liu 0002, Xubin He, Huizhang Luo, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Matthew Wolf, Tong Liu 0030, Zhenbo Qiao |
IPDPS | 6 |
| 2017 | TGE: Machine Learning Based Task Graph Embedding for Large-Scale Topology MappingabstractTask mapping is an important problem in parallel and distributed computing. The goal in task mapping is to find an optimal layout of the processes of an application (or a task) onto a given network topology. We target this problem in the context of staging applications. A staging application consists of two or more parallel applications (also referred to as staging tasks) which run concurrently and exchange data over the course of computation. Task mapping becomes a more challenging problem in staging applications, because not only data is exchanged between the staging tasks, but also the processes of a staging task may exchange data with each other. We propose a novel method, called Task Graph Embedding (TGE), that harnesses the observable graph structures of parallel applications and network topologies. TGE employs a machine learning based algorithm to find the best representation of a graph, called an embedding, onto a space in which the task-to-processor mapping problem can be solved. We evaluate and demonstrate the effectiveness of TGE experimentally with the communication patterns extracted from runs of XGC, a large-scale fusion simulation code, on Titan. Jong Choi 0001, Jeremy Logan, Matthew Wolf, George Ostrouchov, Tahsin M. Kurç, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Melissa Romanus, Manish Parashar, Michael Churchill, Choong-Seock Chang |
CLUSTER | 1 |
| 2017 | Extending Skel to Support the Development and Optimization of Next Generation I/O SystemsabstractAs the memory and storage hierarchy get deeper and more complex, it is important to have new benchmarks and evaluation tools that allow us to explore the emerging middleware solutions to use this hierarchy. Skel is a tool aimed at automating and refining this process of studying HPC I/O performance. It works by generating application I/O kernel/benchmarks as determined by a domain-specific model. This paper provides some techniques for extending Skel to address new situations and to answer new research questions. For example, we document use cases as diverse as using Skel to troubleshoot I/O performance issues for remote users, refining an I/O system model, and facilitating the development and testing of a mechanism for runtime monitoring and performance analytics. We also discuss data oriented extensions to Skel to support the study of compression techniques for Exascale scientific data management. Jeremy Logan, Jong Choi 0001, Matthew Wolf, George Ostrouchov, Lipeng Wan 0001, Norbert Podhorszki, William F. Godoy, Scott Klasky, Erich Lohrmann, Greg Eisenhauer, Chad Wood, Kevin A. Huck |
CLUSTER | 2 |
| 2017 | Canopus: A Paradigm Shift Towards Elastic Extreme-Scale Data Analytics on HPC StorageabstractScientific simulations on high performance computing (HPC) platforms generate large quantities of data. To bridge the widening gap between compute and I/O, and enable data to be more efficiently stored and analyzed, simulation outputs need to be refactored, reduced, and appropriately mapped to storage tiers. However, a systematic solution to support these steps has been lacking on the current HPC software ecosystem. To that end, this paper develops Canopus, a progressive JPEGlike data management scheme for storing and analyzing big scientific data. It co-designs the data decimation, compression and data storage, taking the hardware characteristics of each storage tier into considerations. With reasonably low overhead, our approach refactors simulation data into a much smaller, reduced-accuracy base dataset, and a series of deltas that is used to augment the accuracy if needed. The base dataset and deltas are compressed and written to multiple storage tiers. Data saved on different tiers can then be selectively retrieved to restore the level of accuracy that satisfies data analytics. Thus, Canopus provides a paradigm shift towards elastic data analytics and enables end users to make trade-offs between analysis speed and accuracy on-the-fly. We evaluate the impact of Canopus on unstructured triangular meshes, a pervasive data model used by scientific modeling and simulations. In particular, we demonstrate the progressive data exploration of Canopus using the “blob detection” use case on the fusion simulation data. Tao Lu 0014, Eric Suchyta, David Pugmire, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Norbert Podhorszki, Mark Ainsworth, Matthew Wolf |
CLUSTER | 4 |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo |
Euro-Par | 6 |
| 2017 | Canopus: Enabling Extreme-Scale Data Analytics on Big HPC Storage via Progressive Refactoring
Tao Lu 0014, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Qing Liu 0002, David Pugmire, Matthew Wolf, Mark Ainsworth |
HotStorage | 3 |
| 2017 | Predicting Output Performance of a Petascale SupercomputerabstractIn this paper, we develop a predictive model useful for output performance prediction of supercomputer file systems under production load. Our target environment is Titan---the 3rd fastest supercomputer in the world---and its Lustre-based multi-stage write path. We observe from Titan that although output performance is highly variable at small time scales, the mean performance is stable and consistent over typical application run times. Moreover, we find that output performance is non-linearly related to its correlated parameters due to interference and saturation on individual stages on the path. These observations enable us to build a predictive model of expected write times of output patterns and I/O configurations, using feature transformations to capture non-linear relationships. We identify the candidate features based on the structure of the Lustre/Titan write path, and use feature transformation functions to produce a model space with 135,000 candidate models. By searching for the minimal mean square error in this space we identify a good model and show that it is effective. Yezhou Huang, Jeffrey S. Chase, Jong Choi 0001, Scott Klasky, Jay F. Lofstead, Sarp Oral |
HPDC | 4 |
| 2017 | Exacution: Enhancing Scientific Data Management for ExascaleabstractAs we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed. Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001 |
ICDCS | 7 |
| 2017 | Comprehensive Measurement and Analysis of the User-Perceived I/O Performance in a Production Leadership-Class Storage SystemabstractWith the increase of the scale and intensity of the parallel I/O workloads generated by those scientific applications running on high performance computing facilities, understanding the I/O dynamics, especially the root cause of the I/O performance variability and degradation in HPC environment, have become extremely critical to the HPC community. In this paper, we run extensive I/O measuring tests on a production leadership-class storage system to capture the performance variabilities of large-scale parallel I/O. Analyzing these results and its statistic correlation revealed some valuable insights into the characteristics of the storage system and the root cause of I/O performance variability. Further, we leverage these findings and propose an I/O middleware design refactoring which can improve the performance of the parallel I/O by optimizing the data striping and placement. Our preliminary evaluation results demonstrate the proposed approach can reduce the average per-process write latency by at least 80% and the maximum per-process write latency by at least 20%. Lipeng Wan 0001, Matthew Wolf, Feiyi Wang, Jong Choi 0001, George Ostrouchov, Scott Klasky |
ICDCS | 4 |
| 2017 | Personalized Search Inspired Fast Interactive Estimation of Distribution Algorithm and Its ApplicationabstractInteractive evolutionary algorithms have been applied to personalized search, in which less user fatigue and efficient search are pursued. Motivated by this, we present a fast interactive estimation of distribution algorithm (IEDA) by using the domain knowledge of personalized search. We first induce a Bayesian model to describe the distribution of the new user's preference on the variables from the social knowledge of personalized search. Then we employ the model to enhance the performance of IEDA in two aspects, that is: 1) dramatically reducing the initial huge space to a preferred subspace and 2) generating the individuals of estimation of distribution algorithm(EDA) by using it as a probabilistic model. The Bayesian model is updated along with the implementation of the EDA. To effectively evaluate individuals, we further present a method to quantitatively express the preference of the user based on the human-computer interactions and train a radial basis function neural network as the fitness surrogate. The proposed algorithm is applied to a laptop search, and its superiorities in alleviating user fatigue and speeding up the search procedure are empirically demonstrated. Yang Chen 0007, Xiaoyan Sun 0002, Dun-Wei Gong, Yong Zhang 0016, Jong Choi 0001, Scott Klasky |
IEEE Trans. Evol. Comput. | 5 |
| 2016 | A synthesized ranking-assisted NSGA-II for interval multi-objective optimizationabstractMulti-objective optimization problems with interval (MOPs-I) uncertainties parameters are common in practice. Evolutionary multi-objective (EMO) algorithms are popularly employed to solve these problems due to their powerful explorations. The comparison strategies among interval objectives of MOPs-I are crucially important for obtaining a superior Pareto front when applying EMOs. By effectively combining two different intervals ranking methods together, i.e., μ and P metrics, we present an improved NSGA-II with a synthesized intervals ranking strategy for optimizing MOPs-I. The characteristics of μ and P in ranking intervals are first analyzed, and then the synthesized ranking method termed as μ ⊕ P is developed to compare and select individuals within the NSGA-II framework. The proposed algorithm is experimentally validated by four MOPs-I functions and a practical problem, and the results empirically demonstrate its merits in obtaining Pareto front with outstanding convergence and spread. Ruidong Xu, Xiaoyan Sun 0002, Dun-Wei Gong, Yong Zhang 0016, Jong Choi 0001 |
CEC | 6 |
| 2016 | Towards Real-Time Detection and Tracking of Spatio-Temporal Features: Blob-Filaments in Fusion PlasmaabstractA novel algorithm and implementation of real-time identification and tracking of blob-filaments in fusion reactor data is presented. Similar spatio-temporal features are important in many other applications, for example, ignition kernels in combustion and tumor cells in a medical image. This work presents an approach for extracting these features by dividing the overall task into three steps: local identification of feature cells, grouping feature cells into extended feature, and tracking movement of feature through overlapping in space. Through our extensive work in parallelization, we demonstrate that this approach can effectively make use of a large number of compute nodes to detect and track blob-filaments in real time in fusion plasma. On a set of 30 GB fusion simulation data, we observed linear speedup on 1,024 processes and completed blob detection in less than three milliseconds using Edison, a Cray XC30 system at NERSC. Lingfei Wu 0001, Kesheng Wu, Alex Sim, Michael Churchill, Jong Choi 0001, Andreas Stathopoulos, Choong-Seock Chang, Scott Klasky |
IEEE Trans. Big Data | 5 |
| 2015 | Combining phase identification and statistic modeling for automated parallel benchmark generationabstractParallel application benchmarks are indispensable for evaluating/optimizing HPC software and hardware. However, it is very challenging and costly to obtain high-fidelity benchmarks reflecting the scale and complexity of state-of-the-art parallel applications. Hand-extracted synthetic benchmarks are time- and labor-intensive to create. Real applications themselves, while offering most accurate performance evaluation, are expensive to compile, port, recon- figure, and often plainly inaccessible due to security or ownership concerns. This work contributes APPRIME, a novel tool for trace-based automatic parallel benchmark generation. Taking as input standard communication-I/O traces of an application’s execution, it couples accurate automatic phase identification with statistical regeneration of event parameters to create compact, portable, and to some degree reconfigurable parallel application benchmarks. Experiments with four NAS Parallel Benchmarks (NPB) and three real scientific simulation codes confirm the fidelity of APPRIME benchmarks. They retain the original applications’ performance characteristics, in particular the relative performance across platforms. Xiaosong Ma, Qing Liu 0002, Jeremy Logan, Norbert Podhorszki, Jong Choi 0001, Scott Klasky |
PPoPP | 7 |
| 2015 | Combining Phase Identification and Statistic Modeling for Automated Parallel Benchmark GenerationabstractParallel application benchmarks are indispensable for evaluating/optimizing HPC software and hardware. However, it is very challenging and costly to obtain high-fidelity benchmarks reflecting the scale and complexity of state-of-the-art parallel applications. Hand-extracted synthetic benchmarks are time- and labor-intensive to create. Real applications themselves, while offering most accurate performance evaluation, are expensive to compile, port, reconfigure, and often plainly inaccessible due to security or ownership concerns. This work contributes APPrime, a novel tool for trace-based automatic parallel benchmark generation. Taking as input standard communication-I/O traces of an application's execution, it couples accurate automatic phase identification with statistical regeneration of event parameters to create compact, portable, and to some degree reconfigurable parallel application benchmarks. Experiments with four NAS Parallel Benchmarks (NPB) and three real scientific simulation codes confirm the fidelity of APPrime benchmarks. They retain the original applications' performance characteristics, in particular their relative performance across platforms. Also, the result benchmarks, already released online, are much more compact and easy-to-port compared to the original applications. Xiaosong Ma, Qing Liu 0002, Jeremy Logan, Norbert Podhorszki, Jong Choi 0001, Scott Klasky |
SIGMETRICS | 7 |
| 2014 | Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworksabstractSUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd. Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu |
Concurr. Comput. Pract. Exp. | 6 |
| 2013 | ADIOS Visualization Schema: A First Step Towards Improving Interdisciplinary Collaboration in High Performance ComputingabstractScientific communities have benefitted from a significant increase of available computing and storage resources in the last few decades. For science projects that have access to leadership scale computing resources, the capacity to produce data has been growing exponentially. Teams working on such projects must now include, in addition to the traditional application scientists, experts in various disciplines including applied mathematicians for development of algorithms, visualization specialists for large data, and I/O specialists. Sharing of knowledge and data is becoming a requirement for scientific discovery, providing useful mechanisms to facilitate this sharing is a key challenge for e-Science. Our hypothesis is that in order to decrease the time to solution for application scientists we need to lower the barrier of entry into related computing fields. We aim at improving users' experience when interacting with a vast software ecosystem and/or huge amount of data, while maintaining focus on their primary research field. In this context we present our approach to bridge the gap between the application scientists and the visualization experts through a visualization schema as a first step and proof of concept for a new way to look at interdisciplinary collaboration among scientists dealing with big data. The key to our approach is recognizing that our users are scientists who mostly work as islands. They tend to work in very specialized environment but occasionally have to collaborate with other researchers in order to take full advantage of computing innovations and get insight from big data. We present an example of identifying the connecting elements between one of such relationships and offer a liaison schema to facilitate their collaboration. Roselyne Tchoua, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Jeremy Logan, Kenneth Moreland, Jingqing Mu, Manish Parashar, Norbert Podhorszki, David Pugmire, Matthew Wolf |
e-Science | 2 |
| 2012 | Mining hidden mixture context with ADIOS-P to improve predictive pre-fetcher accuracyabstractPredictive pre-fetcher, which predicts future data access events and loads the data before users requests, has been widely studied, especially in file systems or web contents servers, to reduce data load latency. Especially in scientific data visualization, pre-fetching can reduce the IO waiting time. In order to increase the accuracy, we apply a data mining technique to extract hidden information. More specifically, we apply a data mining technique for discovering the hidden contexts in data access patterns and make prediction based on the inferred context to boost the accuracy. In particular, we performed Probabilistic Latent Semantic Analysis (PLSA), a mixture model based algorithm popular in the text mining area, to mine hidden contexts from the collected user access patterns and, then, we run a predictor within the discovered context. We further improve PLSA by applying the Deterministic Annealing (DA) method to overcome the local optimum problem. In this paper we demonstrate how we can apply PLSA and DA optimization to mine hidden contexts from users data access patterns and improve predictive pre-fetcher performance. Jong Choi 0001, Hasan Abbasi, David Pugmire, Norbert Podhorszki, Scott Klasky, Cristian Capdevila, Manish Parashar, Matthew Wolf, Judy Qiu, Geoffrey C. Fox |
eScience | 1 |
| 2011 | Browsing large-scale cheminformatics data with dimension reductionabstractSUMMARY Visualization of large‐scale high dimensional data is highly valuable for data analysis facilitating scientific discovery in many fields. We present PubChemBrowse, a customized visualization tool for cheminformatics research. It provides a novel 3D data point browser that displays complex properties of massive data on commodity clients. As in Geographic Information System browsers for Earth and Environment data, chemical compounds with similar properties are nearby in the browser. PubChemBrowse is built around in‐house high performance parallel Multi‐dimensional scaling and Generative topographic mapping services and supports fast interaction with an external property database. These properties can be overlaid on 3D mapped compound space or queried for individual points. We prototype the integration with Chem2Bio2RDF system using SPARQL endpoint to access over 20 publicly accessible bioinformatics databases. We describe our design and implementation of the integrated PubChemBrowse application and outline its use in drug discovery. The same core technologies are generally applicable to develop high performance scientific data browsing systems for other applications. Copyright © 2011 John Wiley & Sons, Ltd. Jong Choi 0001, Seung-Hee Bae, Judy Qiu, Bin Chen 0002, David J. Wild 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | Cloud computing paradigms for pleasingly parallel biomedical applicationsabstractSUMMARY Cloud computing offers exciting new approaches for scientific computing that leverage major commercial players’ hardware and software investments in large‐scale data centers. Loosely coupled problems are very important in many scientific fields, and with the ongoing move towards data‐intensive computing, they are on the rise. There exist several different approaches to leveraging clouds and cloud‐oriented data processing frameworks to perform pleasingly parallel (also called embarrassingly parallel) computations. In this paper, we present three pleasingly parallel biomedical applications: (i) assembly of genome fragments; (ii) sequence alignment and similarity search; and (iii) dimension reduction in the analysis of chemical structures, which are implemented utilizing a cloud infrastructure service‐based utility computing models of Amazon Web Services ( http://Amazon.com Inc., Seattle, WA, USA) and Microsoft Windows Azure (Microsoft Corp., Redmond, WA, USA) as well as utilizing MapReduce‐based data processing frameworks Apache Hadoop (Apache Software Foundation, Los Angeles, CA, USA) and Microsoft DryadLINQ. We review and compare each of these frameworks, performing a comparative study among them based on performance, cost, and usability. High latency, eventually consistent cloud infrastructure service‐based frameworks that rely on off‐the‐node cloud storage were able to exhibit performance efficiencies and scalability comparable to the MapReduce‐based frameworks with local disk‐based storage for the applications considered. In this paper, we also analyze variations in cost among the different platform choices (e.g., Elastic Compute Cloud instance types), highlighting the importance of selecting an appropriate platform based on the nature of the computation. Copyright © 2011 John Wiley & Sons, Ltd. Thilina Gunarathne, Tak-Lon Wu, Jong Choi 0001, Seung-Hee Bae, Judy Qiu |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | High Performance Dimension Reduction and Visualization for Large High-Dimensional Data AnalysisabstractLarge high dimension datasets are of growing importance in many fields and it is important to be able to visualize them for understanding the results of data mining approaches or just for browsing them in a way that distance between points in visualization (2D or 3D) space tracks that in original high dimensional space. Dimension reduction is a well understood approach but can be very time and memory intensive for large problems. Here we report on parallel algorithms for Scaling by MAjorizing a Complicated Function (SMACOF) to solve Multidimensional Scaling problem and Generative Topographic Mapping (GTM). The former is particularly time consuming with complexity that grows as square of data set size but has advantage that it does not require explicit vectors for dataset points but just measurement of inter-point dissimilarities. We compare SMACOF and GTM on a subset of the NIH PubChem database which has binary vectors of length 166 bits. We find good parallel performance for both GTM and SMACOF and strong correlation between the dimension-reduced PubChem data from these two methods. Jong Choi 0001, Seung-Hee Bae, Xiaohong Qiu, Geoffrey C. Fox |
CCGRID | 1 |
| 2010 | Dimension reduction and visualization of large high-dimensional data via interpolationabstractThe recent explosion of publicly available biology gene sequences and chemical compounds offers an unprecedented opportunity for data mining. To make data analysis feasible for such vast volume and high-dimensional scientific data, we apply high performance dimension reduction algorithms. It facilitates the investigation of unknown structures in a three dimensional visualization. Among the known dimension reduction algorithms, we utilize the multidimensional scaling and generative topographic mapping algorithms to configure the given high-dimensional data into the target dimension. However, both algorithms require large physical memory as well as computational resources. Thus, the authors propose an interpolated approach to utilizing the mapping of only a subset of the given data. This approach effectively reduces computational complexity. With minor trade-off of approximation, interpolation method makes it possible to process millions of data points with modest amounts of computation and memory requirement. Since huge amount of data are dealt, we represent how to parallelize proposed interpolation algorithms, as well. For the evaluation of the interpolated MDS by STRESS criteria, it is necessary to compute symmetric all pairwise computation with only subset of required data per process, so we also propose a simple but efficient parallel mechanism for the symmetric all pairwise computation when only a subset of data is available to each process. Our experimental results illustrate that the quality of interpolated mapping results are comparable to the mapping results of original algorithm only. In parallel performance aspect, those interpolation methods are well parallelized with high efficiency. With the proposed interpolation method, we construct a configuration of two-million out-of-sample data into the target dimension, and the number of out-of-sample data can be increased further. Seung-Hee Bae, Jong Choi 0001, Judy Qiu, Geoffrey C. Fox |
HPDC | 2 |
| 2010 | Browsing large scale cheminformatics data with dimension reductionabstractVisualization of large-scale high dimensional data tool is highly valuable for scientific discovery in many fields. We presentPubChemBrowse, acustomizedvisualizationtoolfor cheminformatics research. It provides a novel 3D data point browser that displays complex properties of massive data on commodity clients. As in GIS browsers for Earth and Environment data, chemical compounds with similar properties are nearby in the browser. PubChemBrowse is built around in-househighperformanceparallel MDS(Multi-Dimensional Scaling) and GTM (Generative Topographic Mapping) services andsupports fast interaction with anexternalproperty database. These properties can be overlaid on 3D mapped compound space or queried for individual points. We prototype use with Chem2Bio2RDF system using SPARQLquery language to access over 20 publicly accessible bioinformatics databases. We describe our design and implementation of the integrated PubChemBrowse application and outline its use in drug discovery. The same core technologies can be used to develop similar high dimensional browsers in other scientific areas. Jong Choi 0001, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox, Bin Chen 0002, David J. Wild 0001 |
HPDC | 1 |
| 2010 | Hybrid cloud and cluster computing paradigms for life science applicationsabstractBACKGROUND: Clouds and MapReduce have shown themselves to be a broadly useful approach to scientific computing especially for parallel data intensive applications. However they have limited applicability to some areas such as data mining because MapReduce has poor performance on problems with an iterative structure present in the linear algebra that underlies much data analysis. Such problems can be run efficiently on clusters using MPI leading to a hybrid cloud and cluster environment. This motivates the design and implementation of an open source Iterative MapReduce system Twister. RESULTS: Comparisons of Amazon, Azure, and traditional Linux and Windows environments on common applications have shown encouraging performance and usability comparisons in several important non iterative cases. These are linked to MPI applications for final stages of the data analysis. Further we have released the open source Twister Iterative MapReduce and benchmarked it against basic MapReduce (Hadoop) and MPI in information retrieval and life sciences applications. CONCLUSIONS: The hybrid cloud (MapReduce) and cluster (MPI) approach offers an attractive production environment while Twister promises a uniform programming environment for many Life Sciences applications. METHODS: We used commercial clouds Amazon and Azure and the NSF resource FutureGrid to perform detailed comparisons and evaluations of different approaches to data intensive computing. Several applications were developed in MPI, MapReduce and Twister in these different environments. Judy Qiu, Jaliya Ekanayake, Thilina Gunarathne, Jong Choi 0001, Seung-Hee Bae, Bingjing Zhang, Tak-Lon Wu, Yang Ruan 0001, Saliya Ekanayake, Adam Hughes, Geoffrey C. Fox |
BMC Bioinform. | 4 |
| 2009 | Biomedical Case Studies in Data Intensive Computing
Geoffrey C. Fox, Xiaohong Qiu, Scott Beason, Jong Choi 0001, Jaliya Ekanayake, Thilina Gunarathne, Mina Rho, Haixu Tang, Neil Devadasan, Gilbert C. Liu |
CloudCom | 4 |
| 2009 | Using Web 2.0 for scientific applications and scientific communitiesabstractAbstract Web 2.0 approaches are revolutionizing the Internet, blurring lines between developers and users and enabling collaboration and social networks that scale into the millions of users. As discussed in our previous work, the core technologies of Web 2.0 effectively define a comprehensive distributed computing environment that parallels many of the more complicated service‐oriented systems such as Web service and Grid service architectures. In this paper we build upon this previous work to discuss the applications of Web 2.0 approaches to four different scenarios: client‐side JavaScript libraries for building and composing Grid services; integrating server‐side portlets with ‘rich client’ AJAX tools and Web services for analyzing Global Positioning System data; building and analyzing folksonomies of scientific user communities through social bookmarking; and applying microformats and GeoRSS to problems in scientific metadata description and delivery. Copyright © 2009 John Wiley & Sons, Ltd. Marlon E. Pierce, Geoffrey C. Fox, Jong Choi 0001, Zhenhua Guo 0004, Xiaoming Gao |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | SALSA Project: Parallel Data Mining of GIS, Web, Medical, Physics, Chemical, and Biology DataabstractThe multicore revolution promises potentially hundreds of cores in desktop computers. The ever increasing number of cores per chip will be accompanied by a pervasive data deluge whose size will probably increase even faster than CPU core count over the next few years. This suggests the importance of parallel data analysis and data mining applications with good multicore, cluster and grid performance. The SALSA project at Community Grid Lab of Indiana University is looking to revolutionize the way software is written in parallel for real applications that advance scientific discovery and improve the quality of people's life. Xiaohong Qiu, Geoffrey C. Fox, Seung-Hee Bae, Jong Choi 0001, Jaliya Ekanayake, Yang Ruan 0001 |
eScience | 4 |
| 2008 | BioVLAB-Microarray: Microarray Data Analysis in Virtual EnvironmentabstractMicroarray technology is a high-throughput experimental technique that can measure expression levels of hundreds of thousands of genes simultaneously. To interpret massive data from gene-expression microarray experiments, biologists encounter computational and analytical challenges. This is especially challenging for small research labs that lack local computing and bioinformatics expertise. Here, we introduce a virtual analysis system for microarray gene expression data in computing clouds with flexible and configurable GUI workflow engine so that biologists are able to analyze the data in many angles without worrying about computational and bioinformatics issues. Youngik Yang, Jong Choi 0001, Kwangmin Choi, Marlon E. Pierce, Dennis Gannon, Sun Kim |
eScience | 2 |
| 2008 | PRECIP: Towards Practical and Retrofittable Confidential Information Protection
XiaoFeng Wang 0001, Zhuowei Li 0001, Ninghui Li 0001, Jong Choi 0001 |
NDSS | 4 |
| 2008 | Fast and Black-box Exploit Detection and Signature Generation for Commodity SoftwareabstractIn biology, a vaccine is a weakened strain of a virus or bacterium that is intentionally injected into the body for the purpose of stimulating antibody production. Inspired by this idea, we propose a packet vaccine mechanism that randomizes address-like strings in packet payloads to carry out fast exploit detection and signature generation. An exploit with a randomized jump address behaves like a vaccine: it will likely cause an exception in a vulnerable program’s process when attempting to hijack the control flow, and thereby expose itself. Taking that exploit as a template, our signature generator creates a set of new vaccines to probe the program in an attempt to uncover the necessary conditions for the exploit to happen. A signature is built upon these conditions to shield the underlying vulnerability from further attacks. In this way, packet vaccine detects exploits and generates signatures in a black-box fashion, that is, not relying on the knowledge of a vulnerable program’s source and binary code. Therefore, it even works on the commodity software obfuscated for the purpose of copyright protection. In addition, since our approach avoids the expense of tracking the program’s execution flow, it performs almost as fast as a normal run of the program and is capable of generating a signature of high quality within seconds or even subseconds. We present the design of the packet vaccine mechanism and an example of its application. We also describe our proof-of-concept implementation and the evaluation of our technique using real exploits. XiaoFeng Wang 0001, Zhuowei Li 0001, Jong Choi 0001, Jun Xu 0003, Michael K. Reiter, Chongkyung Kil |
ACM Trans. Inf. Syst. Secur. | 3 |
| 2007 | SpyShield: Preserving Privacy from Spy Add-Ons
Zhuowei Li 0001, XiaoFeng Wang 0001, Jong Choi 0001 |
RAID | 3 |
| 2006 | Packet vaccine: black-box exploit detection and signature generationabstractIn biology,a vaccine is a weakened strain of a virus or bacterium that is intentionally injected into the body for the purpose of stimulating antibody production.Inspired by this idea, we propose a packet vaccine mechanism that randomizes address-like strings in packet payloads to carry out fast exploit detection, vulnerability diagnosis and signature generation. An exploit with a randomized jump address behaves like a vaccine: it will likely cause an exception in a vulnerable program's process when attempting to hijack the control flow,and thereby expose itself. Taking that exploit as a template, our signature generator creates a set of new vaccines to probe the program, in an attempt to uncover the necessary conditions for the exploit to happen. A signature is built upon these conditions to shield the underlying vulnerability from further attacks. In this way, packet vaccine detects and fllters exploits in a black-box fashion,i.e., avoiding the expense of tracking the program's execution flow. We present the design of the packet vaccine mechanism and an example of its application. We also describe our proof-of-concept implementation and the evaluation of our technique using real exploits. XiaoFeng Wang 0001, Zhuowei Li 0001, Jun Xu 0003, Michael K. Reiter, Chongkyung Kil, Jong Choi 0001 |
CCS | 6 |
| 2006 | Tamper-Evident Digital Signature Protecting Certification Authorities Against MalwareabstractWe introduce the notion of tamper-evidence for digital signature generation in order to defend against attacks aimed at covertly leaking secret information held by corrupted signing nodes. This is achieved by letting observers (which need not be trusted) verify the absence of covert channels by means of techniques we introduce herein. We call our signature schemes tamper-evident since any deviation from the protocol is immediately detectable. We demonstrate our technique for the RSA-PSS (known as RSA's probabilistic signature scheme) and DSA signature schemes and show how the same technique can be applied to the Schnorr and Feige-Fiat-Shamir (FFS) signature schemes. Our technique does not modify the distribution of the generated signature transcripts, and has only a minimal overhead in terms of computation, communication, and storage Jong Choi 0001, Philippe Golle, Markus Jakobsson |
DASC | 1 |