Houjun Tang

dblp:19/8438 · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-7038-8360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 AURORA-Q: Asynchronous Unified Resource Optimizer for Quantum Simulation on HPC System
Changjong Kim, Alex Sim, Kesheng Wu, Houjun Tang, Yongseok Son, Jisung Park 0001, Sunggon Kim
ICDCS4
2025 Swiftn: Accelerating Quantum Circuit Simulation Through Tensor Optimization
abstract
Quantum computers are evolving at a rapid pace and are considered next-generation computers with high computational capabilities. However, due to the unique characteristics of qubits, state-of-the-art quantum computers are vulnerable to noise caused by qubit instability. To overcome this, highperformance computing (HPC) systems are utilized for quantum circuit simulations to evaluate complex quantum algorithms with great accuracy. However, quantum circuit simulations have high computational demands, and the data volume increases exponentially as the number of qubits increases. In this paper, we propose SWIFTN, a quantum circuit simulation optimization framework for HPC systems with scalability. To achieve this, it enhances parallelism by dividing the tensor networks and distributing them across multiple GPUs and nodes. Additionally, it reduces computational costs by bypassing tasks through intermittent tensor contraction. Finally, to mitigate the degradation in accuracy due to intermittent tensor contraction,SWIFTNperforms amplitude adjustments. We implement and evaluateSWIFTNusing a Perlmutter supercomputer. Our evaluation results using popular quantum algorithm benchmark (i.e., QAOA) shows thatSWIFTNcan improve the performance by$7.85 \times$with 99.997 % accuracy.
Changjong Kim, Alex Sim, Kesheng Wu, Houjun Tang, Sunggon Kim
CCGrid5
2025 Data Management in the Continuum: Cross-facility Object-based Data Transfers
abstract
Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.
Jean Luca Bez, Houjun Tang, Chen Wang 0004, Surendra Byna
SBAC-PAD2
2025 Regen: An object layout regenerator on large-scale production HPC systems
Dong Kyu Sung, Sunggon Kim, Sangjin Lee 0003, Houjun Tang, Alex Sim, Kesheng Wu, Surendra Byna, Yongseok Son
Future Gener. Comput. Syst.4
2024 IDIOMS: Index-powered Distributed Object-centric Metadata Search for Scientific Data Management
abstract
Affix-oriented metadata search is one of the essential fuzzy search capabilities that allow users to find data of interest in their voluminous data set with incomplete query conditions. With the recent transition towards object-centric data management systems in the science community, there is a paramount need for the support of such features in distributed settings. However, existing metadata search solutions either do not support efficient affix-oriented metadata search or do not suit well in a distributed setting of object-centric data management systems. To bridge this gap, we introduce IDIOMS, a metadata search solution underpinned by a distributed metadata index, meticulously designed to enable high-performance affix-oriented metadata search for parallel object-centric storage. One of the standout features of IDIOMS is its efficiency in supporting four distinct types of highly demanded metadata queries. Furthermore, IDIOMS is flexibly catering to both independent and collective metadata search operations. Our experimental comparisons with SoMeta, a state-of-the-art metadata query method, demonstrate more than 400× performance boost for independent queries and up to 300× performance improvements for collective queries, while keeping a small index management overhead.
Wei Zhang 0097, Houjun Tang, Surendra Byna
CCGrid2
2024 A2FL: Autonomous and Adaptive File Layout in HPC through Real-time Access Pattern Analysis
abstract
Various scientific applications with different I/O characteristics are executed in HPC systems. However, underlying parallel file systems are unaware of these characteristics of applications, and using a single fixed file layout for all applications can degrade the performance of HPC systems. In this paper, we propose A2FL, an autonomous and adaptive file layout adjustment scheme that optimizes parallel file system configurations by analyzing the access pattern of the applications. The key steps of A2FL are as follows: (1) A2FL initially intercepts the I/O operations of the application, recording their access patterns in real-time. (2) The access patterns are then transformed into a graphical representation used for predicting I/O performance and providing adjustment recommendations. (3) A2FL autonomously adjusts the file layout based on the prediction results, delivering an optimal file layout within the parallel file system. Moreover, we propose A2FL-Compound which analyzes an access pattern by dividing it into smaller components to optimize the file layout in a fine-grained manner. Our evaluations demonstrate that A2FL significantly enhances I/O performance, with improvements of up to 65.9× compared to the default file layout.
Dong Kyu Sung, Yongseok Son, Alex Sim, Kesheng Wu, Surendra Byna, Houjun Tang, Hyeonsang Eom, Changjong Kim, Sunggon Kim
IPDPS6
2024 h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre-exascale platforms
abstract
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.
Jean Luca Bez, Houjun Tang, M. Scot Breitenfeld, Huihuo Zheng, Wei-keng Liao, Kaiyuan Hou, Zanhua Huang, Surendra Byna
Concurr. Comput. Pract. Exp.2
2024 PROV-IO$^+$+: A Cross-Platform Provenance Framework for Scientific Data on HPC Systems
abstract
Data provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. In this paper, we analyze four representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs. Based on the first-hand analysis, we propose a provenance framework called PROV-IO$^+$, which includes an I/O-centric provenance model for describing scientific data and the associated I/O operations and environments precisely. Moreover, we build a prototype of PROV-IO$^+$to enable end-to-end provenance support on real HPC systems with little manual effort. The PROV-IO$^+$framework can support both containerized and non-containerized workflows on different HPC platforms with flexibility in selecting various classes of provenance. Our experiments with realistic workflows show that PROV-IO$^+$can address the provenance needs of the domain scientists effectively with reasonable performance (e.g., less than 3.5% tracking overhead for most experiments). Moreover, PROV-IO$^+$outperforms a state-of-the-art system (i.e., ProvLake) in our experiments.
Runzhou Han, Mai Zheng, Surendra Byna, Houjun Tang, Bin Dong 0002, Dong Dai 0001, Yong Chen 0001, Dongkyun Kim, Joseph Hassoun, David Thorsley
IEEE Trans. Parallel Distributed Syst.4
2023 Evaluating Asynchronous Parallel I/O on HPC Systems
abstract
Parallel I/O is an effective method to optimize data movement between memory and storage for many scientific applications. Poor performance of traditional disk-based file systems has led to the design of I/O libraries which take advantage of faster memory layers, such as on-node memory, present in high-performance computing (HPC) systems. By allowing caching and prefetching of data for applications alternating computation and I/O phases, a faster memory layer also provides opportunities for hiding the latency of I/O phases by overlapping them with computation phases, a technique called asynchronous I/O. Since asynchronous parallel I/O in HPC systems is still in the initial stages of development, there hasn't been a systematic study of the factors affecting its performance.In this paper, we perform a systematic study of various factors affecting the performance and efficacy of asynchronous I/O, we develop a performance model to estimate the aggregate I/O bandwidth achievable by iterative applications using synchronous and asynchronous I/O based on past observations, and we evaluate the performance of the recently developed asynchronous I/O feature of a parallel I/O library (HDF5) using benchmarks and real-world science applications. Our study covers parallel file systems on two large-scale HPC systems: Summit and Cori, the former with a GPFS storage and the latter with a Lustre parallel file system.
John Ravi, Surendra Byna, Quincey Koziol, Houjun Tang, Michela Becchi
IPDPS4
2023 AMRIC: A Novel In Situ Lossy Compression Framework for Efficient I/O in Adaptive Mesh Refinement Applications
abstract
As supercomputers advance towards exascale capabilities, computational intensity increases significantly, and the volume of data requiring storage and transmission experiences exponential growth. Adaptive Mesh Refinement (AMR) has emerged as an effective solution to address these two challenges. Concurrently, error-bounded lossy compression is recognized as one of the most efficient approaches to tackle the latter issue. Despite their respective advantages, few attempts have been made to investigate how AMR and error-bounded lossy compression can function together. To this end, this study presents a novel in-situ lossy compression framework that employs the HDF5 filter to improve both I/O costs and boost compression quality for AMR applications. We implement our solution into the AMReX framework and evaluate on two real-world AMR applications, Nyx and WarpX, on the Summit supercomputer. Experiments with 4096 CPU cores demonstrate that AMRIC improves the compression ratio by up to 81× and the I/O performance by up to 39× over AMReX's original compression solution.
Daoce Wang, Jesus Pulido, Pascal Grosset, Jiannan Tian, Sian Jin, Houjun Tang, Jean M. Sexton, Sheng Di, Kai Zhao 0008, Bo Fang 0002, Zarija Lukic, Franck Cappello, James P. Ahrens, Dingwen Tao
SC6
2022 HDF5 Cache VOL: Efficient and Scalable Parallel I/O through Caching Data on Node-local Storage
abstract
Modern-era high performance computing (HPC) systems are providing multiple levels of memory and storage layers to bridge the performance gap between fast memory and slow disk-based storage system managed by Lustre or GPFS. Several of the recent HPC systems are equipped with SSD and NVMe-based storage that is attached locally to compute nodes. A few systems are providing an SSD-based “burst buffer” intermediate storage layer that is accessible by all compute nodes as a single file system. Although these hardware layers are intended to reduce the latency gap between memory and disk-based long-term storage, how to utilize them has been left to the users. High-level I/O libraries, such as HDF5 and netCDF, can potentially take advantage of the node-local storage as a cache for reducing I/O latency from capacity storage. However, it is challenging to use node-local storage in parallel I/O especially for a single shared file. In this paper, we present an approach to integrate node-local storage as transparent caching or staging layers in a high-level parallel I/O library without placing the burden of managing these layers on users. We designed this to move data asynchronously between the caching storage layer and a parallel file system to overlap the data movement overhead in performing I/O with compute phases. We implement this approach as an external HDF5 Virtual Object Layer (VOL) connector, named Cache VOL. HDF5 VOL is a layer of abstraction in HDF5 that allows intercepting the public HDF5 application programming interface (API) and performing various optimizations to data movement after the interception. Existing HDF5 applications can use Cache VOL with minimal code modifications. We evaluated the performance of Cache VOL in HPC applications such as VPIC-10, and deep learning applications such as ImageNet and CosmoFlow. We show that using Cache VOL, one can achieve higher observed I/O performance, more scalable and stable I/O compared to direct I/O to the parallel file system, thus achieving faster time-to-solution in scientific simulations. While the caching approach is implemented in HDF5, the methods are applicable in other high-level I/O libraries.
Huihuo Zheng, Venkatram Vishwanath, Quincey Koziol, Houjun Tang, John Ravi, John Mainzer, Surendra Byna
CCGRID4
2022 PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC Systems
abstract
cData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations.
Runzhou Han, Surendra Byna, Houjun Tang, Bin Dong 0002, Mai Zheng
HPDC3
2022 Accelerating Parallel Write via Deeply Integrating Predictive Lossy Compression with HDF5
abstract
Lossy compression is one of the most efficient solutions to reduce storage overhead and improve I/O performance for HPC applications. However, existing parallel I/O libraries cannot fully utilize lossy compression to accelerate parallel write due to the lack of deep understanding on compression-write performance. To this end, we propose to deeply integrate predictive lossy compression with HDF5 to significantly improve the parallel-write performance. Specifically, we propose analytical models to predict the time of compression and parallel write before the actual compression to enable compression-write overlapping. We also introduce an extra space in the process to handle possible data overflows resulting from prediction uncertainty in compression ratios. Moreover, we propose an optimization to reorder the compression tasks to increase the overlapping efficiency. Experiments with up to 4,096 cores from Summit show that our solution improves the write performance by up to$4.5\times$and$2.9\times$over the non-compression and lossy compression solutions, respectively, with only 1.5% storage overhead (compared to original data) on two real-world HPC applications.
Sian Jin, Dingwen Tao, Houjun Tang, Sheng Di, Surendra Byna, Zarija Lukic, Franck Cappello
SC3
2022 Improving nonnegative matrix factorization with advanced graph regularization
Degang Chen 0002, Hong Yu 0007, Guoyin Wang 0001, Houjun Tang, Kesheng Wu
Inf. Sci.5
2022 Transparent Asynchronous Parallel I/O Using Background Threads
abstract
Moving toward exascale computing, the size of data stored and accessed by applications is ever increasing. However, traditional disk-based storage has not seen improvements that keep up with the explosion of data volume or the speed of processors. Multiple levels of non-volatile storage devices are being added to handle bursty I/O, however, moving data across the storage hierarchy can take longer than the data generation or analysis. Asynchronous I/O can reduce the impact of I/O latency as it allows applications to schedule I/O early and to check their status later. I/O is thus overlapped with application communication or computation or both, effectively hiding some or all of the I/O latency. POSIX and MPI-I/O provide asynchronous read and write operations, but lack the support for non-data operations such as file open and close. Users also have to manually manage data dependencies and use low-level byte offsets, which requires significant effort and expertise to adopt. In this article, we present an asynchronous I/O framework that supports all types of I/O operations, manages data dependencies transparently and automatically, provides implicit and explicit modes for application flexibility, and error information retrieval. We implemented these techniques in HDF5. Our evaluation of several benchmarks and application workloads demonstrates it effectiveness on hiding the I/O cost from the application.
Houjun Tang, Quincey Koziol, John Ravi, Surendra Byna
IEEE Trans. Parallel Distributed Syst.1
2021 Tuning Parallel Data Compression and I/O for Large-scale Earthquake Simulation
abstract
Scientific applications, such as those simulating earthquakes, the origins of universe, etc., often produce massive amounts of data as high-performance computing (HPC) systems are moving toward exascale. The ever-increasing volumes of data are posing challenges for scientists to store, share, analyze, and visualize. Compression algorithms have become a crucial component for data management in scientific workflows. Data reduction enables simulations to output more data without worrying about exceeding storage quotas, and could capture more insights in the simulation. However, due to the complexity and poor performance of I/O and compression libraries as well as parallel file systems, the overall compression and I/O performance varies significantly. In this paper, we explore tuning parallel compression of data produced by a large-scale earthquake simulation. We show that our strategies achieve up to 13X performance improvement and a compression ratio of up to 251.
Houjun Tang, Surendra Byna, N. Anders Petersson, David McCallen
IEEE BigData1
2021 Battle of the Defaults: Extracting Performance Characteristics of HDF5 under Production Load
abstract
Popular parallel I/O libraries, such as HDF5, provide tuning parameters to obtain superior performance. However, the selection of effective parameters on production systems is complex due to the interdependence of I/O software and file system layers. Hence, application developers typically use the default parameters and often experience poor I/O performance. This work conducts a benchmarking-based analysis on the HDF5 behaviors with a wide variety of I/O patterns to extract performance characteristics under the production workload. To make the analysis well controlled, we exercise I/O benchmarks on POSIX-IO, MPI-IO, and HDF5 using the same I/O patterns and in the same jobs. To address high performance variability in production environments, we repeat the benchmarks across I/O patterns, storage devices, and time intervals. Based on the results, we identified consistent HDF5 behaviors that appropriate configurations and operations on dataset layout and file-metadata placement can improve performance significantly. We apply our findings and evaluate the tuned I/O library on two supercomputers: Summit and Cori. The results show that our tuned parameters can achieve more than 10× I/O performance speedup than that with default parameters on both systems, suggesting the effectiveness, stability, and generality of our solution.
Houjun Tang, Surendra Byna, Jesse Hanley, Quincey Koziol, Tonglin Li, Sarp Oral
CCGRID2
2020 Interfacing HDF5 with a scalable object-centric storage system on hierarchical storage
abstract
Summary Object storage technologies that take advantage of multitier storage on HPC systems are emerging. However, to use these technologies at present, applications have to be modified significantly from current I/O libraries. HDF5, a widely used I/O middleware on HPC systems, provides a virtual object layer (VOL) that allows applications to connect to different storage mechanisms transparently without requiring significant code modifications. We recently designed the proactive data containers (PDC) object‐centric storage system that provides the capabilities of transparent, asynchronous, and autonomous data movement taking advantage of multiple storage tiers—a decision that has so far been left upon the user on most current systems. To enable PDC's features through HDF5 without modifying application codes, we have developed an HDF5 VOL connector that interfaces with PDC. We present in this article the connector interface and evaluate its performance on Cori, a Cray XC40 supercomputer located at the National Energy Research Scientific Computing Center (NERSC). Our evaluation demonstrates up to an 8× improvement compared with HDF5 that has the most recent optimizations.
Jingqing Mu, Jérome Soumagne, Surendra Byna, Quincey Koziol, Houjun Tang, Richard Warren
Concurr. Comput. Pract. Exp.5
2020 ExaHDF5: Delivering Efficient Parallel I/O on Exascale Computing Systems
Surendra Byna, M. Scot Breitenfeld, Bin Dong 0002, Quincey Koziol, Elena Pourmal, Dana Robinson, Jérome Soumagne, Houjun Tang, Venkatram Vishwanath, Richard Warren
J. Comput. Sci. Technol.8
2019 Tuning Object-Centric Data Management Systems for Large Scale Scientific Applications
abstract
Efficient management of scientific data on high-performance computing (HPC) systems has been a challenge, as it often requires knowledge of various hardware and software components of the system, as well as tedious manual effort in optimizing parallel I/O for each application. This situation is exacerbated by the fact that storage systems on upcoming exascale supercomputers are equipped with an unprecedented level of complexity due to a deep storage and memory hierarchy with heterogeneous hardware and their management software. Simple and effective data management methods are critical for numerous scientific applications that are storing and analyzing massive amounts of data on HPC systems. Object-centric data management systems (ODMS) provide an easy-to-use interface, allow for massive scalability with relaxed consistency, and have been gaining popularity in the HPC community. However, tuning an ODMS to achieve its full potential on existing HPC systems with large-scale science use cases still remains a challenging task. In this paper, we explore and evaluate various well-known I/O tuning techniques on a new ODMS called Proactive Data Containers (PDC). Our experiments using real science applications and I/O kernels demonstrate that the benefits of these tuning methods with up to 9X I/O performance speedup over the previous version of PDC, and 47X over a highly optimized HDF5 implementation.
Houjun Tang, Surendra Byna, Zarija Lukic, Jialin Liu 0002, Quincey Koziol, Bin Dong 0002
HiPC1
2019 Analysis in the Data Path of an Object-Centric Data Management System
abstract
Emerging high performance computing (HPC) systems are expected to be deployed with an unprecedented level of complexity due to a deep system memory and storage hierarchy. Efficient and scalable methods of data management and movement through the multi-level storage hierarchy of upcoming HPC systems will be critical for scientific applications at exascale. In this paper, we propose in locus analysis that allows registering user-defined functions (UDFs) and running those functions automatically while the data is moving between levels of a storage hierarchy. We implement this analysis in the data path approach in our object-centric data management system, called Proactive Data Containers (PDC). The transparent invocation of analysis functions as part of PDC object mapping is an optimized approach to minimize latency to access data as it moves within the storage hierarchy. Because a user defined analysis or transform function will be invoked automatically by the PDC runtime, the user simply registers their functions for PDC to identify the function name as well as the required list of actual parameters. To demonstrate the validity and flexibility of this analysis approach, we have implemented several scientific analysis kernels to compare against other HPC analysis-oriented approaches.
Richard Warren, Jérome Soumagne, Jingqing Mu, Houjun Tang, Surendra Byna, Bin Dong 0002, Quincey Koziol
HiPC4
2019 MIQS: metadata indexing and querying service for self-describing file formats
abstract
Scientific applications often store datasets in self-describing data file formats, such as HDF5 and netCDF. Regrettably, to efficiently search the metadata within these files remains challenging due to the sheer size of the datasets. Existing solutions extract the metadata and store it in external database management systems (DBMS) to locate desired data. However, this practice introduces significant overhead and complexity in extraction and querying. In this research, we propose a novel Metadata Indexing and Querying Service (MIQS), which removes the external DBMS and utilizes in-memory index to achieve efficient metadata searching. MIQS follows the self-contained data management paradigm and provides portable and schema-free metadata indexing and querying functionalities for self-describing file formats. We have evaluated MIQS with the state-of-the-art MongoDB-based metadata indexing solution. MIQS achieved up to 99% time reduction in index construction and up to 172kx search performance improvement with up to 75% reduction in memory footprint.
Wei Zhang 0097, Surendra Byna, Houjun Tang, Brody Williams, Yong Chen 0001
SC3
2018 DART: distributed adaptive radix tree for efficient affix-based keyword search on HPC systems
abstract
Affix-based search is a fundamental functionality for storage systems. It allows users to find desired datasets, where attributes of a dataset match an affix. While building inverted index to facilitate efficient affix-based keyword search is a common practice for standalone databases and for desktop file systems, building local indexes or adopting indexing techniques used in a standalone data store is insufficient for high-performance computing (HPC) systems due to the massive amount of data and distributed nature of the storage devices within a system. In this paper, we propose Distributed Adaptive Radix Tree (DART), to address the challenge of distributed affix-based keyword search on HPC systems. This trie-based approach is scalable in achieving efficient affix-based search and alleviating imbalanced keyword distribution and excessive requests on keywords at scale. Our evaluation at different scales shows that, comparing with the "full string hashing" use case of the most popular distributed indexing technique - Distributed Hash Table (DHT), DART achieves up to 55× better throughput with prefix search and with suffix search, while achieving comparable throughput with exact and infix searches. Also, comparing to the "initial hashing" use case of DHT, DART maintains a balanced keyword distribution on distributed nodes and alleviates excessive query workload against popular keywords.
Wei Zhang 0097, Houjun Tang, Surendra Byna, Yong Chen 0001
PACT2
2018 ARCHIE: Data Analysis Acceleration with Array Caching in Hierarchical Storage
abstract
Scientific data analysis typically involves reading massive amounts of data that was generated by simulations, experiments, and observations. Performance of reading such large volumes of data from disk-based file systems is often poor because of the slow and mechanical components in the disks. Recent supercomputing systems are adding non-volatile storage layers in a hierarchy to handle the performance gap between fast main memory and slow disk-based storage. Software libraries for managing this hierarchy not only need efficient reading of data but also reduce user-involvement for cross-layer data movement. Furthermore, these libraries need to support array data access patterns into hierarchical storage management as scientific data is often organized in array-based data structures. Existing software typically manage individual storage layers requiring significant manual process in moving data among them. In this paper, we introduce a new array caching in hierarchical storage (ARCHIE) to accelerate array data analysis in a seamless fashion. ARCHIE evaluates array access patterns and prefetches data with array semantics between storage layers. Our evaluation shows that ARCHIE outperforms state-of-the-art file systems, i.e., Lustre and DataWarp, on a production supercomputing system by up to 5.8× in accessing data by scientific analysis applications.
Bin Dong 0002, Houjun Tang, Quincey Koziol, Kesheng Wu, Surendra Byna
IEEE BigData3
2018 Toward Scalable and Asynchronous Object-Centric Data Management for HPC
abstract
Emerging high performance computing (HPC) systems are expected to be deployed with an unprecedented level of complexity due to a deep system memory and storage hierarchy. Efficient and scalable methods of data management and movement through this hierarchy is critical for scientific applications using exascale systems. Moving toward new paradigms for scalable I/O in the extreme-scale era, we introduce novel object-centric data abstractions and storage mechanisms that take advantage of the deep storage hierarchy, named Proactive Data Containers (PDC). In this paper, we formulate object-centric PDCs and their mappings in different levels of the storage hierarchy. PDC adopts a client-server architecture with a set of servers managing data movement across storage layers. To demonstrate the effectiveness of the proposed PDC system, we have measured performance of benchmarks and I/O kernels from scientific simulation and analysis applications using PDC programming interface, and compared the results with existing highly tuned I/O libraries. Using asynchronous I/O along with data and metadata optimizations, PDC demonstrates up to 23× speedup over HDF5 and PLFS in writing and reading data from a plasma physics simulation. PDC achieves comparable performance with HDF5 and PLFS in reading and writing data of a single timestep at small scale, and outperforms them at a scale of larger than 10K cores. In contrast to existing storage systems, PDC offers user-space data management with the flexibility to allocate the number of PDC servers depending on the workload.
Houjun Tang, Surendra Byna, Francois Tessier, Bin Dong 0002, Jingqing Mu, Quincey Koziol, Jérome Soumagne, Venkatram Vishwanath, Jialin Liu 0002, Richard Warren
CCGrid1
2018 A Transparent Server-Managed Object Storage System for HPC
abstract
On the road to exascale, the high-performance computing (HPC) community is seeing the emergence of multi-tier storage systems. However, existing data management solutions for HPC applications are no longer suitable for handling the increased level of storage complexity and currently delegate that task back to the user. We describe a novel object-based data abstraction that takes advantage of deep memory hierarchies by providing a simplified programming interface that enables autonomous, asynchronous, and transparent data movement with a server-driven architecture. Users can define a mapping between the application memory and abstract storage objects, creating a linkage between either all or part of an object's content without data copy or transfer, avoiding explicit management of complex data movement across multiple storage hierarchies. We evaluate our system by storing plasma physics simulation data with different storage layouts.
Jingqing Mu, Jérome Soumagne, Houjun Tang, Surendra Byna, Quincey Koziol, Richard Warren
CLUSTER3
2018 UniviStor: Integrated Hierarchical and Distributed Storage for HPC
abstract
High performance computing (HPC) architectures have been adding new layers of storage, such as burst buffers, to tolerate latency between memory and disk-based file systems. However, existing file system and burst buffer management software typically manage each storage layer separately. As a result, the burden of moving data across multiple layers falls upon HPC system users. To hide the complexity of managing the scattered storage devices from applications, we introduce UniviStor, a data management service offering a unified view of storage layers. By considering each layer's distinct characteristics, UniviStor provides performance optimizations and data structures tailored for distributed and hierarchical data placement, interference aware data movement scheduling, adaptive data striping, and lightweight workflow management. UniviStor supports parallel I/O library APIs, such as MPI-IO and HDF5. Our evaluations on a large-scale supercomputer demonstrated that UniviStor outperforms Data Elevator, a state-of-the-art transparent caching solution for burst buffers by up to 17x, and Lustre by up to 46x.
Surendra Byna, Bin Dong 0002, Houjun Tang
CLUSTER4
2017 SoMeta: Scalable Object-Centric Metadata Management for High Performance Computing
abstract
Scientific data sets, which grow rapidly in volume, are often attached with plentiful metadata, such as their associated experiment or simulation information. Thus, it becomes difficult for them to be utilized and their value is lost over time. Ideally, metadata should be managed along with its corresponding data by a single storage system, and can be accessed and updated directly. However, existing storage systems in high-performance computing (HPC) environments, such as Lustre parallel file system, still use a static metadata structure composed of non-extensible and fixed amount of information. The burden of metadata management falls upon the end-users and require ad-hoc metadata management software to be developed.With the advent of "object-centric" storage systems, there is an opportunity to solve this issue. In this paper, we present SoMeta, a scalable and decentralized metadata management approach for object-centric storage in HPC systems. It provides a flat namespace that is dynamically partitioned, a tagging approach to manage metadata that can be efficiently searched and updated, and a light-weight and fault tolerant management strategy. In our experiments, SoMeta achieves up to 3.7X speedup over Lustre in performing common metadata operations, and up to 16X faster than SciDB and MongoDB for advanced metadata operations, such as adding and searching tags. Additionally, in contrast to existing storage systems, SoMeta offers scalable user-space metadata management by allowing users with the capability to specify the number of metadata servers depending on their workload.
Houjun Tang, Surendra Byna, Bin Dong 0002, Jialin Liu 0002, Quincey Koziol
CLUSTER1
2016 Exploring memory hierarchy and network topology for runtime AMR data sharing across scientific applications
abstract
Runtime data sharing across applications is of great importance for avoiding high I/O overhead for scientific data analytics. Sharing data on a staging space running on a set of dedicated compute nodes is faster than writing data to a slow disk-based parallel file system (PFS) and then reading it back for post-processing. Originally, the staging space has been purely based on main memory (DRAM), and thus was several orders of magnitude faster than the PFS approach. However, storing all the data produced by large-scale simulations on DRAM is impractical. Moving data from memory to SSD-based burst buffers is a potential approach to address this issue. However, SSDs are about one order of magnitude slower than DRAM. To optimize data access performance over the staging space, methods such as prefetching data from SSDs according to detected spatial access patterns and distributing data across the network topology have been explored. Although these methods work well for uniform mesh data, which they were designed for, they are not well suited for adaptive mesh refinement (AMR) data. Two major issues must be addressed before constructing such a memory hierarchy and topology-aware runtime AMR data sharing framework: (1) spatial access pattern detection and prefetching for AMR data; (2) AMR data distribution across the network topology at runtime. We propose a framework that addresses these challenges and demonstrate its effectiveness with extensive experiments on AMR data. Our results show the framework's spatial access pattern detection and prefetching methods demonstrate about 26% performance improvement for client analytical processes. Moreover, the framework's topology-aware data placement can improve overall data access performance by up to 18%.
Wenzhao Zhang, Houjun Tang, Stephen Ranshous, Surendra Byna, Daniel F. Martin, Kesheng Wu, Bin Dong 0002, Scott Klasky, Nagiza F. Samatova
IEEE BigData2
2016 Usage Pattern-Driven Dynamic Data Layout Reorganization
abstract
As scientific simulations and experiments move toward extremely large scales and generate massive amounts of data, the data access performance of analytic applications becomes crucial. A mismatch often happens between write and read patterns of data accesses, typically resulting in poor read performance. Data layout reorganization has been used to improve the locality of data accesses. However, current data reorganizations are static and focus on generating a single (or set of) optimized layouts that rely on prior knowledge of exact future access patterns. We propose a framework that dynamically recognizes the data usage patterns, replicates the data of interest in multiple reorganized layouts that would benefit common read patterns, and makes runtime decisions on selecting a favorable layout for a given read pattern. This framework supports reading individual elements and chunks of a multi-dimensional array of variables. Our pattern-driven layout selection strategy achieves multi-fold speedups compared to reading from the original dataset.
Houjun Tang, Surendra Byna, Steve Harenberg, Xiaocheng Zou, Wenzhao Zhang, Kesheng Wu, Bin Dong 0002, Oliver Rübel, Kristofer E. Bouchard, Scott Klasky, Nagiza F. Samatova
CCGrid1
2016 AMRZone: A Runtime AMR Data Sharing Framework for Scientific Applications
abstract
Frameworks that facilitate runtime data sharingacross multiple applications are of great importance for scientificdata analytics. Although existing frameworks work well overuniform mesh data, they can not effectively handle adaptive meshrefinement (AMR) data. Among the challenges to construct anAMR-capable framework include: (1) designing an architecturethat facilitates online AMR data management, (2) achievinga load-balanced AMR data distribution for the data stagingspace at runtime, and (3) building an effective online indexto support the unique spatial data retrieval requirements forAMR data. Towards addressing these challenges to supportruntime AMR data sharing across scientific applications, wepresent the AMRZone framework. Experiments over real-worldAMR datasets demonstrate AMRZone's effectiveness at achievinga balanced workload distribution, reading/writing large-scaledatasets with thousands of parallel processes, and satisfyingqueries with spatial constraints. Moreover, AMRZone's performance and scalability are even comparable with existing state-of-the-art work when tested over uniform mesh data with up to16384 cores, in the best case, our framework achieves a 46% performance improvement.
Wenzhao Zhang, Houjun Tang, Steve Harenberg, Surendra Byna, Xiaocheng Zou, Dharshi Devendran, Daniel F. Martin, Kesheng Wu, Bin Dong 0002, Scott Klasky, Nagiza F. Samatova
CCGrid2
2016 In Situ Storage Layout Optimization for AMR Spatio-temporal Read Accesses
abstract
Analyses of large simulation data often concentrate on regions in space and in time that contain important information. As simulations adopt Adaptive Mesh Refinement (AMR), the data records from a region of interest could be widely scattered on storage devices and accessing interesting regions results in significantly reduced I/O performance. In this work, we study the organization of block-structured AMR data on storage to improve performance of spatio-temporal data accesses. AMR has a complex hierarchical multi-resolution data structure that does not fit easily with the existing approaches that focus on uniform mesh data. To enable efficient AMR read accesses, we develop an in situ data layout optimization framework. Our framework automatically selects from a set of candidate layouts based on a performance model, and reorganizes the data before writing to storage. We evaluate this framework with three AMR datasets and access patterns derived from scientific applications. Our performance model is able to identify the best layout scheme and yields up to a 3X read performance improvement compared to the original layout. Though it is not possible to turn all read accesses into contiguous reads, we are able to achieve 90% of contiguous read throughput with the optimized layouts on average.
Houjun Tang, Surendra Byna, Steve Harenberg, Wenzhao Zhang, Xiaocheng Zou, Daniel F. Martin, Bin Dong 0002, Dharshi Devendran, Kesheng Wu, David Trebotich, Scott Klasky, Nagiza F. Samatova
ICPP1
2015 Parallel In Situ Detection of Connected Components in Adaptive Mesh Refinement Data
abstract
Adaptive Mesh Refinement (AMR) represents a significant advance for scientific simulation codes, greatly reducing memory and compute requirements by dynamically varying simulation resolution over space and time. As simulation codes transition to AMR, existing analysis algorithms must also make this transition. One such algorithm, connected component detection, is of vital importance in many simulation and analysis contexts, with some simulation codes even relying on parallel, in situ connected component detection for correctness. Yet, current detection algorithms designed for uniform meshes are not applicable to hierarchical, non-uniform AMR, and to the best of our knowledge, AMR connected component detection has not been explored in the literature. Therefore, in this paper, we formally define the general problem of connected component detection for AMR, and present a general solution. Beyond solving the general detection problem, achieving viable in situ detection performance is even more challenging. The core issue is the conflict between the communication-intensive nature of connected component detection (in general, and especially for AMR data) and the requirement that in situ processes incur minimal performance impact on the co-located simulation. We address this challenge by presenting the first connected component detection methodology for structured AMR that is applicable in a parallel, in situ context. Our key strategy is the incorporation of an multi-phase AMR-aware communication pattern that synchronizes connectivity information across the AMR hierarchy. In addition, we distil our methodology to a generic framework within the Combo AMR infrastructure, making connected component detection services available for many existing applications. We demonstrate our method's efficacy by showing its ability to detect ice calving events in real time within the real-world BISICLES ice sheet modelling code. Results show up to a 6.8x speedup of our algorithm over the existing specialized BISICLES algorithm. We also show scalability results for our method up to 4,096 cores using a parallel Combo-based benchmark.
Xiaocheng Zou, Kesheng Wu, David A. Boyuka II, Daniel F. Martin, Surendra Byna, Houjun Tang, Kushal Bansal, Terry J. Ligocki, Hans Johansen, Nagiza F. Samatova
CCGRID6
2015 Exploring Memory Hierarchy to Improve Scientific Data Read Performance
abstract
Improving read performance is one of the major challenges with speeding up scientific data analytic applications. Utilizing the memory hierarchy is one major line of researches to address the read performance bottleneck. Related methods usually combine solide-state-drives(SSDs) with dynamic random-access memory(DRAM) and/or parallel file system(PFS) to mitigate the speed and space gap between DRAM and PFS. However, these methods are unable to handle key performance issues plaguing SSDs, namely read contention that may cause up to 50% performance reduction. In this paper, we propose a framework that exploits the memory hierarchy resource to address the read contention issues involved with SSDs. The framework employs a general purpose online read algorithm that able to detect and utilize memory hierarchy resource to relieve the problem. To maintain a near optimal operating environment for SSDs, the framework is able to orchastrate data chunks across different memory layers to facilitate the read algorithm. Compared to existing tools, our framework achieves up to 50% read performance improvement when tested on datasets from real-world scientific simulations.
Wenzhao Zhang, Houjun Tang, Xiaocheng Zou, Steve Harenberg, Qing Liu 0002, Scott Klasky, Nagiza F. Samatova
CLUSTER2
2015 The hyperdyadic index and generalized indexing and query with PIQUE
abstract
Many scientists rely on indexing and query to identify trends and anomalies within extreme-scale scientific data. Compressed bitmap indexing (e.g., FastBit) is the go-to indexing method for many scientific datasets and query workloads. Recently, the ALACRITY compressed inverted index was shown as a viable alternative approach. Notably, though FastBit and ALACRITY employ very different data structures (inverted list vs. bitmap) and binning methods (bit-wise vs. decimal-precision), close examination reveals marked similarities in index structure.
David A. Boyuka II, Houjun Tang, Kushal Bansal, Xiaocheng Zou, Scott Klasky, Nagiza F. Samatova
SSDBM2
2014 Improving Read Performance with Online Access Pattern Analysis and Prefetching
Houjun Tang, Xiaocheng Zou, John Jenkins, David A. Boyuka II, Stephen Ranshous, Dries Kimpe, Scott Klasky, Nagiza F. Samatova
Euro-Par1
2014 Fast Set Intersection through Run-Time Bitmap Construction over PForDelta-Compressed Indexes
Xiaocheng Zou, Sriram Lakshminarasimhan, David A. Boyuka II, Stephen Ranshous, Houjun Tang, Scott Klasky, Nagiza F. Samatova
Euro-Par5
2013 A Generic High-Performance Method for Deinterleaving Scientific Data
Eric R. Schendel, Steve Harenberg, Houjun Tang, Venkatram Vishwanath, Michael E. Papka, Nagiza F. Samatova
Euro-Par3
1997 A unified surface smoothing scheme for automobile body shape modeling design
abstract
A unified surface smoothing scheme is proposed to fit grid points, scatter points and boundary points, which are digitized from physical models with the Coordinate Measurement Machine, to smoothness parameter surfaces. A variety of curves and surfaces with different smoothness can be obtained with the aid of parameters such as approximation and smoothness weights, segment or patch numbers and the distribution of approximation and smoothness weights. A pickup car body is finished with the proposed method.
Bingyan Zhao, Houjun Tang, Shigenori Okubo
Shape Modeling International2