EDBT 2026 Demo / reviewers in the wild / expert
Hyogi Sim
dblp:169/1762
· DBLP profile ↗
11ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0003-2485-2171ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Storage systems · 84% High-performance computing · 10% Cloud and datacenter computing · 5% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 13 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems › distributed file system
parallel file system |
0.8 | 2 | 2021 | Exploiting user activeness for data retention in HPC systems · SC 2021 Scientific user behavior and data-sharing trends in a petascale file system · SC 2017 |
Storage systems
key-value storage |
0.8 | 2 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 bespoKV: application tailored scale-out key-value stores · SC 2018 |
Storage systems › file systems
distributed file system |
0.7 | 2 | 2020 | An Integrated Indexing and Search Service for Distributed File Systems · IEEE Trans. Parallel Distributed Syst. 2020 Tagit: an integrated indexing and search service for file systems · SC 2017 |
Storage systems › data management › scientific data storage
analysis workflow-aware storage |
0.6 | 2 | 2019 | An Analysis Workflow-Aware Storage System for Multi-Core Active Flash Arrays · IEEE Trans. Parallel Distributed Syst. 2019 AnalyzeThis: an analysis workflow-aware storage system · SC 2015 |
Storage systems › computational storage
in-storage computing |
0.6 | 2 | 2019 | An Analysis Workflow-Aware Storage System for Multi-Core Active Flash Arrays · IEEE Trans. Parallel Distributed Syst. 2019 AnalyzeThis: an analysis workflow-aware storage system · SC 2015 |
Storage systems › storage reliability › durability
data retention |
0.5 | 1 | 2021 | Exploiting user activeness for data retention in HPC systems · SC 2021 |
Storage systems › key-value storage
distributed key-value store |
0.4 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems › key-value storage
scalable key-value store |
0.4 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems › flash and SSD › flash memory
flash storage |
0.4 | 1 | 2019 | An Analysis Workflow-Aware Storage System for Multi-Core Active Flash Arrays · IEEE Trans. Parallel Distributed Syst. 2019 |
High-performance computing
supercomputing |
0.1 | 1 | 2021 | Exploiting user activeness for data retention in HPC systems · SC 2021 |
Information retrieval › distributed information retrieval
distributed search |
0.1 | 1 | 2020 | An Integrated Indexing and Search Service for Distributed File Systems · IEEE Trans. Parallel Distributed Syst. 2020 |
High-performance computing › scientific computing
HPC applications |
0.1 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2017 | Scientific user behavior and data-sharing trends in a petascale file system · SC 2017 |
Methods — techniques the papers use, named apart from their topics
distributed indexing · 0.9active operator offloading · 0.9integrated indexing · 0.6trace analysis · 0.5activeness-based prioritization · 0.5event-driven simulation · 0.4emulation · 0.4quantitative file system metrics · 0.3metadata snapshot analysis · 0.3active flash emulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | An Analysis of System Balance and Architectural Trends Based on Top500 SupercomputersabstractSupercomputer design is a complex, multi-dimensional optimization process, wherein several subsystems need to be reconciled to meet a desired figure of merit performance for a portfolio of applications and a budget constraint. However, overall, the HPC community has been gravitating towards ever more Flops, at the expense of many other subsystems. To draw attention to overall system balance, in this paper, we analyze balance ratios and architectural trends in the world’s most powerful supercomputers. Specifically, we have collected the performance characteristics of systems between 1993 and 2019 based on the Top500 lists and then analyzed their architectures from diverse system design perspectives. Notably, our analysis studies the performance balance of the machines, across a variety of subsystems such as compute, memory, I/O, interconnect, intra-node connectivity and power. Our analysis reveals that balance ratios of the various subsystems need to be considered carefully alongside the application workload portfolio to provision the subsystem capacity and bandwidth specifications, which can help achieve optimal performance. Awais Khan 0002, Hyogi Sim, Sudharshan S. Vazhkudai, Ali Raza Butt, Youngjae Kim 0001 |
HPC Asia | 2 |
| 2021 | Exploiting user activeness for data retention in HPC systemsabstractHPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology. Wei Zhang 0097, Surendra Byna, Hyogi Sim, Sankeun Lee 0001, Sudharshan S. Vazhkudai, Yong Chen 0001 |
SC | 3 |
| 2020 | Customizable Scale-Out Key-Value StoresabstractEnterprise KV stores are often not well suited for HPC applications, and thus cumbersome end-to-end KV design customization is required to meet the needs of modern HPC applications. To this end, in this article we present bespoKV, an adaptive, extensible, and scale-out KV store framework. bespoKV decouples the KV store design into the control plane for distributed management and the data plane for local data store. For the control plane, bespoKVprovides pre-built modules, called controlets, supporting common distributed functionalities (e.g., replication, consistency, and topology) and their various combinations. This decoupling allows bespoKV to take a user-provided single-server KV store, called a datalet, and transparently enables a scalable and fault-tolerant distributed KV store service. The resulting distributed stores are also adaptive to consistency or topology requirement changes and can be easily extended for new types of services. Such specializations enable innovative uses of KV stores in HPC applications, especially for emerging applications that utilize KV-friendly workloads. We evaluate bespoKV in a local testbed as well as in a public cloud settings. Experiments show that bespoKV-enabled distributed KV stores scale horizontally to a large number of nodes, and performs comparably and sometimes 1.2× to 2.6× better than the state-of-the-art systems. Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | An Integrated Indexing and Search Service for Distributed File SystemsabstractData services such as search, discovery, and management in scalable distributed environments have traditionally been decoupled from the underlying file systems, and are often deployed using external databases and indexing services. However, modern data production rates, looming data movement costs, and the lack of metadata, entail revisiting the decoupled file system-data services design philosophy. In this article, we present TagIt, a scalable data management service framework aimed at scientific datasets, which can be integrated into prevalent distributed file system architectures. A key feature of TagIt is a scalable, distributed metadata indexing framework, which facilitates a flexible tagging capability to support data discovery. Furthermore, the tags can also be associated with an active operator, for pre-processing, filtering, or automatic metadata extraction, which we seamlessly offload to file servers in a load-aware fashion. We have integrated TagIt into two popular distributed file systems, i.e., GlusterFS and CephFS. Our evaluation demonstrates that TagIt can expedite data search operation by up to 10× over the extant decoupled approach. Hyogi Sim, Awais Khan 0002, Sudharshan S. Vazhkudai, Seung-Hwan Lim, Ali Raza Butt, Youngjae Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Profiling the Usage of an Extreme-Scale Archival Storage SystemabstractProfiling the archival storage system in scientific computing environments has received much less attention compared to the parallel file system, but is equally important since it stores the final data products safely, for a long duration. In this paper, we analyze eight years worth of data transfer logs for accessing the archival file system (HPSS) in the Oak Ridge Leadership Computing Facility (OLCF), which has been hosting the world's largest supercomputers and file systems. Our analysis encompasses about 135 million data transfer activities to the 80 PB High Performance Storage System (HPSS), between 2010 and 2017. We analyze the logs from several dimensions, including studying the workload characteristics (e.g., access patterns, frequency of accesses and temporal behavior), file system characteristics (e.g., directory depth, file system scaling trends, file types), and scientific user behavior (e.g., domain-specific usage and organization). Based on the analysis, we derive insights into the future evolution of the archive in terms of provisioning, desired features and functionality from the archive software, role and right sizing of the archive tiers, quota management, and the importance of smart and efficient metadata and storage management. We believe our study will prove useful for both operating current archival storage and better provisioning future systems. Hyogi Sim, Sudharshan S. Vazhkudai |
MASCOTS | 1 |
| 2019 | An Analysis Workflow-Aware Storage System for Multi-Core Active Flash ArraysabstractThe need for novel data analysis is urgent in the face of a data deluge from modern applications. Traditional approaches to data analysis incur significant data movement costs, moving data back and forth between the storage system and the processor. Emerging Active Flash devices enable processing on the flash, where the data already resides. An array of such Active Flash devices allows us to revisit how analysis workflows interact with storage systems. By seamlessly blending together the flash storage and data analysis, we create an analysis workflow-aware storage system, AnalyzeThis. Our guiding principle is that analysis-awareness be deeply ingrained in each and every layer of the storage, elevating data analyses as first-class citizens, and transforming AnalyzeThis into a potent analytics-aware appliance. To evaluate the AnalyzeThis system, we have adopted both emulation and simulation approaches. In particular, we have evaluated AnalyzeThis by implementing the AnalyzeThis storage system on top of the Active Flash Array's emulation platform. We have also implemented an event-driven AnalyzeThis simulator, called AnalyzeThisSim, which allows us to address the limitations of the emulation platform, e.g., performance impact of using multi-core SSDs. The results from our emulation and simulation platforms indicate that AnalyzeThis is a viable approach for expediting workflow execution and minimizing data movement. Hyogi Sim, Geoffroy Vallée, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Devesh Tiwari, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | bespoKV: application tailored scale-out key-value stores
Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
SC | 5 |
| 2017 | AnalyzeThat: A Programmable Shared-Memory System for an Array of Processing-In-Memory DevicesabstractProcessing In Memory (PIM), the concept of integrating processing directly with memory, has been attracting a lot of attention since PIM can assist in overcoming the throughput limitation caused by data movement between CPU and memory. The challenge, however, is that it requires the programmers to have a deep understanding of the PIM architecture to maximize the benefits such as data locality and parallel thread execution on multiple PIM devices. In this study, we present AnalyzeThat, a programmable shared-memory system for parallel data processing with PIM devices. Thematic to AnalyzeThat is a rich PIM-Aware Data Structure (PADS), which is an encapsulation that integrally ties together the data, the analysis tasks and the runtime needed to interface with the PIM device array. The PADS abstraction provides (i) a key-value data container that allows programmers to easily store data on multiple PIMs, (ii) a suite of parallel operations with which users can easily implement data analysis applications, and (iii) a runtime, hidden to programmers, which provides the mechanisms needed to overlay both the data and the tasks on the PIM device array in an intelligent fashion, based on PIM-specific information collected from the hardware. We have developed a PIM emulation framework called AnalyzeThat. Our experimental evaluation with representative data analytics applications suggests that the proposed system can significantly reduce the PIM programming effort without losing its technology benefits. Sangkeun Matt Lee, Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai |
CCGrid | 2 |
| 2017 | Scientific user behavior and data-sharing trends in a petascale file systemabstractThe Oak Ridge Leadership Computing Facility (OLCF) runs the No. 4 supercomputer in the world, supported by a petascale file system, to facilitate scientific discovery. In this paper, using the daily file system metadata snapshots collected over 500 days, we have studied the behavioral trends of 1, 362 active users and 380 projects across 35 science domains. In particular, we have analyzed both individual and collective behavior of users and projects, highlighting needs from individual communities and the overall requirements to operate the file system. We have analyzed the metadata across three dimensions, namely (i) the projects' file generation and usage trends, using quantitative file system-centric metrics, (ii) scientific user behavior on the file system, and (iii) the data sharing trends of users and projects. To the best of our knowledge, our work is the first of its kind to provide comprehensive insights on user behavior from multiple science domains through metadata analysis of a large-scale shared file system. We envision that this OLCF case study will provide valuable insights for the design, operation, and management of storage systems at scale, and also encourage other HPC centers to undertake similar such efforts. Seung-Hwan Lim, Hyogi Sim, Raghul Gunasekaran, Sudharshan S. Vazhkudai |
SC | 2 |
| 2017 | Tagit: an integrated indexing and search service for file systemsabstractData services such as search, discovery, and management in scalable distributed environments have traditionally been decoupled from the underlying file systems, and are often deployed using external databases and indexing services. However, modern data production rates, looming data movement costs, and the lack of metadata, entail revisiting the decoupled file system-data services design philosophy. Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Geoffroy Vallée, Seung-Hwan Lim, Ali Raza Butt |
SC | 1 |
| 2015 | AnalyzeThis: an analysis workflow-aware storage systemabstractThe need for novel data analysis is urgent in the face of a data deluge from modern applications. Traditional approaches to data analysis incur significant data movement costs, moving data back and forth between the storage system and the processor. Emerging Active Flash devices enable processing on the flash, where the data already resides. An array of such Active Flash devices allows us to revisit how analysis workflows interact with storage systems. By seamlessly blending together the flash storage and data analysis, we create an analysis workflow-aware storage system, AnalyzeThis. Our guiding principle is that analysis-awareness be deeply ingrained in each and every layer of the storage, elevating data analyses as first-class citizens, and transforming AnalyzeThis into a potent analytics-aware appliance. We implement the AnalyzeThis storage system atop an emulation platform of the Active Flash array. Our results indicate that AnalyzeThis is viable, expediting workflow execution and minimizing data movement. Hyogi Sim, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Devesh Tiwari, Ali Anwar 0001, Ali Raza Butt, Lavanya Ramakrishnan |
SC | 1 |