EDBT 2026 Demo / reviewers in the wild / expert
Sungyong Park
dblp:38/1991 · also Sung-Yong Park
· DBLP profile ↗
47ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quicktopia: Iteration-Level GPU Frequency Control for Energy-Latency Co-Optimization in LLM Inference
Soyang Baek, Bodon Jeong, Hongsu Byun, Sungyong Park |
CCGrid | 4 |
| 2026 | Dual-Blade: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
Bodon Jeong, Hongsu Byun, Youngjae Kim 0001, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park |
ICDCS | 7 |
| 2026 | Resystance: Unleashing Hidden Performance of Compaction in LSM-Trees Via eBPFabstractThe development of high-speed storage devices such as NVMe SSDs has shifted the primary I/O bottleneck from hardware to software. Modern database systems also rely on kernel-based I/O paths, where frequent system call invocations and kernel-user space transitions lead to relatively large overheads and performance degradation. This issue is particularly pronounced in Log-Structured Merge-tree (LSM-tree)-based NoSQL databases. We identified that, in particular, the background compaction process generates a large number of read system calls, causing significant overhead. To address this problem, we propose RESYSTANCE, which leverages eBPF and io_uring to free compaction from system calls and unlock hidden performance potential. RESYSTANCE improves disk I/O efficiency during read operations via io uring and significantly reduces software stack overhead by handling compaction directly inside the kernel through eBPF. Moreover, RESYSTANCE minimizes user-kernel transitions by offloading key I/O routines into the kernel without modifying the LSM-tree structure or compaction algorithm. RESYSTANCE was extensively evaluated using db_bench, YCSB, and OLTP workloads. Compared to baseline RocksDB, it reduced the average number of system call invocations during compaction by 99% and shortened compaction time by 50%. Consequently, in write-intensive workloads, RESYSTANCE improved throughput by up to 75% and reduced the p99 latency by 40%. Hongsu Byun, Honghyeon Yoo, Myoungjoon Kim, Sungyong Park |
ICDE | 5 |
| 2025 | SIDL: A Real-World Dataset for Restoring Smartphone Images with Dirty LensesabstractSmartphone cameras are ubiquitous in daily life, yet their performance can be severely impacted by dirty lenses, leading to degraded image quality. This issue is often overlooked in image restoration research, which assumes ideal or controlled lens conditions. To address this gap, we introduced SIDL (Smartphone Images with Dirty Lenses), a novel dataset designed to restore images captured through contaminated smartphone lenses. SIDL contains diverse real-world images taken under various lighting conditions and environments. These images feature a wide range of lens contaminants, including water drops, fingerprints, and dust. Each contaminated image is paired with a clean reference image, enabling supervised learning approaches for restoration tasks. To evaluate the challenge posed by SIDL, various state-of-the-art restoration models were trained and compared on this dataset. Their performances achieved some level of restoration but did not adequately address the diverse and realistic nature of the lens contaminants in SIDL. This challenge highlights the need for more robust and adaptable image restoration techniques for restoring images with dirty lenses. Sooyoung Choi, Sungyong Park |
AAAI | 2 |
| 2025 | Streamrag: a Lock-Aware and Traffic-Aware Query Coordinator in Stream-Based Rag SystemsabstractStream-based retrieval augmented generation (RAG) systems integrate stream processing engines (SPEs) with real-time document retrieval to support dynamic indexing and search over unstructured datasets. However, efficiently executing queries in these systems is challenging, as seamless coordination between SPEs and vector databases is essential for maintaining low latency and high throughput. The lack of mutual awareness between these components results in two major performance bottlenecks. First, SPEs are unaware of ongoing indexing operations in vector databases, leading to metadata lock contention when indexing and search operations overlap, which increases query latency. Second, vector databases lack visibility into query traffic from SPEs and rely solely on internal metrics for scaling. As a result, they respond reactively to traffic spikes, often leading to instance overload and delayed query processing. To address these issues, we propose StreamRAG, a lock-aware and traffic-aware query coordination mechanism that facilitates real-time exchange of metadata lock statuses and query traffic metrics between SPEs and vector databases. By optimizing query routing and enabling proactive instance scaling, StreamRAG enhances the performance and scalability of real-time RAG systems. Experimental results demonstrate that Streamrag reduces tail latency by up to$4 \times$at the 99 th percentile and significantly improves overall system performance under varying traffic conditions. Yeonwoo Jeong, Kyuli Park, Sungyong Park |
CCGrid | 3 |
| 2025 | ECO-KVS: Energy-Aware Compaction Offloading Mechanism for LSM-Tree Based Key-Value Stores in Edge FederationabstractIn recent years, the rise in energy consumption across infrastructure has highlighted the need for more energy-efficient technologies. This is particularly critical in edge computing environments, where resources and power are limited. Consequently, there is increasing interest in improving the energy efficiency of resource-intensive tasks on edge servers. Edge servers commonly use Log-Structured Merge-tree-based Key-Value Store (LSM-KVS), to manage continuous data streams from edge devices. A key operation in LSM-KVS, known as compaction, merges key-value pairs in a CPU-intensive and energy-demanding process. Additionally, delays during compaction can cause write stalls, blocking I/O operations and degrading performance. This creates a significant challenge in balancing energy consumption and system performance. To address these challenges, we propose ECO-KVS, a solution that improves both energy efficiency and performance in LSM-KVS by offloading compaction tasks across edge servers in an edge federation. ECO-KVS leverages a real-time learning model to predict compaction time and energy consumption, reducing write stalls and enhancing overall energy efficiency. Implemented on RocksDB, ECO-KVS achieves up to 21% higher throughput compared to the baseline RocksDB and improves the performance-to-energy efficiency ratio by up to 18 % compared to EdgePilot, a state-of-the-art solution for edge environments. Jeeseob Kim, Hongsu Byun, Myoungjoon Kim, Youngjae Kim 0001, Zaipeng Xie, Sungyong Park |
CCGrid | 7 |
| 2025 | Equilibria: Co-Optimizing Energy and Latency in Online ML-Based Stream Processing SystemsabstractAn Online machine learning (ML)-based streaming processing system (SPS) combines real-time stream processing with continuous, incremental learning through simultaneous model training and inference. This system processes large, dynamic, high-velocity data streams while adapting its models to improve performance over time. However, balancing the tradeoff between latency and energy efficiency remains a critical challenge, which has not been adequately addressed in prior research. This paper introduces EQUILIBRIA, a novel framework designed to co-optimize power consumption and latency in Online ML-based SPS. EQUILIBRIA integrates dynamic voltage and frequency scaling (DVFS) with two innovative energy optimization strategies. First, a Pareto-based clock frequency adjustment mechanism dynamically tunes both core and memory clock frequencies to reduce latency while minimizing energy consumption. Second, a two-tier threshold training management technique optimizes energy use by periodically pausing and resuming model training once accuracy requirements are met, all while preserving latency. Experimental evaluations across various queries and traffic scenarios demonstrate that EQUILIBRIA achieves up to 58% energy savings without compromising latency, making a significant step forwards in energy-efficient, highperformance streaming analytics for modern, rapidly evolving data environments. Sejeong Oh, Soyang Baek, Gordon Euhyun Moon, Sungyong Park |
CCGrid | 4 |
| 2025 | DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AIabstractRobotic manipulation in embodied AI critically depends on large-scale, high-quality datasets that reflect realistic object interactions and physical dynamics. However, existing data collection pipelines are often slow, expensive, and heavily reliant on manual efforts. We present DynScene, a diffusion-based framework for generating dynamic robotic manipulation scenes directly from textual instructions. Unlike prior methods that focus solely on static environments or isolated robot actions, DynScene decomposes the generation into two phases static scene synthesis and action trajectory generation allowing fine-grained control and diversity. Our model enhances realism and physical feasibility through scene refinement (layout sampling, quaternion quantization) and leverages residual action representation to enable action augmentation, generating multiple diverse trajectories from a single static configuration. Experiments show DynScene achieves 26.8× faster generation, 1.84× higher accuracy, and 28% greater action diversity than human-crafted data. Furthermore, agents trained with DynScene exhibit up to 19.4% higher success rates across complex manipulation tasks. Our approach paves the way for scalable, automated dataset generation in robot learning. Sungyong Park |
CVPR | 2 |
| 2025 | CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector DatabasesabstractEmbedding models capture both semantic and syntactic structures of queries, often mapping different queries to similar regions in vector space. This results in nonuniform cluster access patterns in modern disk-based vector databases. While existing approaches optimize individual queries, they overlook the impact of cluster access patterns, failing to account for the locality effects of queries that access similar clusters. This oversight increases cache miss penalty. To minimize the cache miss penalty, we propose CALL, a context-aware query grouping mechanism that organizes queries based on shared cluster access patterns. Additionally, CALL incorporates a group-aware prefetching method to minimize cache misses during transitions between query groups and latency-aware cluster loading. Experimental results show that CALL reduces the 99th percentile tail latency by up to 33 % while consistently maintaining a higher cache hit ratio, substantially reducing search latency. Yeonwoo Jeong, Hyunji Cho, Kyuli Park, Youngjae Kim 0001, Sungyong Park |
HiPC | 5 |
| 2025 | Revisiting Multi-threaded Compaction in LSM-trees: Enabling Compaction PipeliningabstractWe reveal that modern LSM-tree multi-threaded compaction suffers from limited cross-level parallelism, which prevents concurrent compactions across multiple levels. This limitation leads to an imbalance in thread assignment and causes throughput to saturate even when more threads are added. To address this limitation, we propose a compaction strategy called DownForce. DownForce enables multiple compactions to be executed across levels by introducing non-blocking pipelined compaction, allowing level-wise compactions to proceed simultaneously. This resolves thread imbalance and achieves fully multi-threaded compaction. DownForce is implemented in RocksDB, a representative LSM-tree-based key-value store, and supports both leveled and tiered compaction. In our evaluation, leveled compaction enhanced with DownForce achieves an average of 1.44 × higher thread-level parallelism and delivers up to 1.81 × higher throughput under write-intensive workloads, compared to the conventional multi-threaded leveled compaction. Hongsu Byun, Honghyeon Yoo, Sungyong Park |
ICPP | 3 |
| 2025 | PACMAN: Power Usage-Aware Capping and Management for All-Flash Storage ServersabstractRecently, power management has become increasingly important in storage-centric infrastructures, where SSD arrays are emerging as major power consumers. In such environments, static storage power capping (SSPC) has been used to reduce the peak power consumption of SSD arrays by fixing the power states of individual devices, aiming to stay within constrained power budgets. However, SSPC does not account for workload I/O characteristics or internal SSD behaviors such as garbage collection. This often leads to over-capping, which degrades performance, or underutilization of available power. To address these limitations, we propose PACMAN, a power usageaware capping and management framework designed for the storage layer in all-flash storage servers. PACMAN dynamically reallocates SSD power caps in real time through predictive adjustments based on recent per-device power usage. Without relying on performance counters, PACMAN efficiently adapts to workload variability and optimizes performance and energy efficiency within a given power budget. PACMAN has been extensively evaluated using both the FIO benchmark and a realworld application, RocksDB. Compared with SSPC, PACMAN improved throughput by up to 41% and energy efficiency by up to 24%. Bodon Jeong, Hongsu Byun, Kyungkeun Lee, Bumjun Kim, Jeong-Uk Kang, Sungyong Park |
MASCOTS | 6 |
| 2024 | Coordinating Compaction Between LSM-Tree Based Key-Value Stores for Edge FederationabstractEdge computing environments increasingly demand real-time data processing, leading to the adoption of log-structured merge-tree based key-value stores (LSM-KVS) for efficient data handling. LSM-KVS periodically runs compaction operations in the background to manage the database. However compaction delays cause write stalls, which lead to degraded throughput of LSM-KVS and system performance on resource-limited edge servers. An edge federation environment, which shares resources and tasks between edge servers, can alleviate the resource limitations. Such environments can leverage com-paction offloading where another server performs CPU-intensive compaction operations instead. But coordinating compaction offloading is an important challenge, as the performance of the server performing the compaction can be degraded. In this paper, we propose Edgepilot. Edgepilot is scheduling mechanism of compaction offloading that is designed for LSM-KVS within edge federation. Edgepilot schedules where to reallocate compactions among the edge servers. This is achieved by considering the resource and computing power of each server. As a result, the overall resource efficiency and compaction throughput are increased. In addition, Edgepi-lotprovides Edgecode to determine the effectiveness of compaction offloading. Edgecode is a mathematical modeling based on compaction processing data to approximate inter-server compaction processing times. Edgepilot is implemented on the prominent LSM-KVS, RocksDB v8.3.2, and demonstrates notable improvements compared to the conventional RocksDB. The overall write stall duration of the system is reduced by up to 71 %, and throughput is increased by 17%. Jeeseob Kim, Honghyeon Yoo, Hongsu Byun, Sungyong Park |
CLOUD | 5 |
| 2024 | ML-Based Dynamic Operator-Level Query Mapping for Stream Processing Systems in Heterogeneous Computing EnvironmentsabstractMapping queries to optimal computing devices at the operator-level presents a significant challenge in stream processing systems (SPS) with heterogeneous computing resources. Inefficient query mapping can degrade the performance of the SPS. To address this issue, existing approaches employ static methods, such as mapping all queries to either CPUs or GPUs, or maintaining static mapping tables for queries or operators based on their predetermined device preferences. However, the static mapping scheme fails to provide an optimal solution, as the device preference for different query operators changes dynamically at runtime. In this paper, we propose DynO, a high performance SPS that dynamically maps queries to devices at the operator-level using a tree-based machine learning algorithm. To effectively determine an optimized device mapping plan for query operators, DynO employs a tree-based gradient boosting model to accurately predict the execution time for all potential mapping plan combinations. DynO also introduces a novel turn-based updating scheme to maximize performance in stream processing while training a tree-based gradient boosting model. Additionally, we devise an efficient device mapping scheme to expedite the process of determining the optimal device mapping plan by leveraging a direct acyclic graph (DAG) shortest path algorithm. DynO completely hides any overhead caused by the extra computation needed to find the optimal plan by utilizing prefetching and GPU idle periods. Experimental results using a variety of queries and traffic patterns show that DynO outperforms existing state-of-the-art approaches by ensuring high throughput, low latency, and high efficiency. Sejeong Oh, Gordon Euhyun Moon, Sungyong Park |
CLUSTER | 3 |
| 2024 | An Analytical Model-based Capacity Planning Approach for Building CSD-based Storage SystemsabstractThe data movement in large-scale computing facilities (from compute nodes to data nodes) is categorized as one of the major contributors to high cost and energy utilization. To tackle it, in-storage processing (ISP) within storage devices, such as Solid-State Drives (SSDs), has been explored actively. The introduction of computational storage drives (CSDs) enabled ISP within the same form factor as regular SSDs and made it easy to replace SSDs within traditional compute nodes. With CSDs, host systems can offload various operations such as search, filter, and count. However, commercialized CSDs have different hardware resources and performance characteristics. Thus, it requires careful consideration of hardware, performance, and workload characteristics for building a CSD-based storage system within a compute node. Therefore, storage architects are hesitant to build a storage system based on CSDs as there are no tools to determine the benefits of CSD-based compute nodes to meet the performance requirements compared to traditional nodes based on SSDs. In this work, we proposed an analytical model-based storage capacity planner called CsdPlan for system architects to build performance-effective CSD-based compute nodes. Our model takes into account the performance characteristics of the host system, targeted workloads, and hardware and performance characteristics of CSDs to be deployed and provides optimal configuration based on the number of CSDs for a compute node. Furthermore, CsdPlan estimates and reduces the total cost of ownership (TCO) for building a CSD-based compute node. To evaluate the efficacy of CsdPlan , we selected two commercially available CSDs and four representative big data analysis workloads. Hongsu Byun, Safdar Jamil, Jungwook Han, Sungyong Park, Myungcheol Lee, Changsoo Kim, Beongjun Choi, Youngjae Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Chronica: A Data-Imbalance-Aware Scheduler for Distributed Deep LearningabstractOne of the major challenges in distributed deep learning is attenuating straggler problem. The straggler increases synchronization latency and significantly inhibits the convergence of deep learning model. We empirically observe that the imbal-anced data samples worsen the straggler problem and make the convergence of the deep learning model slower. However, existing approaches such as BOA and EP4DDL have not addressed data imbalance issues while solving the straggler problem. To overcome the straggler and data imbalance problems, we propose Chronica,a new data-imbalance-aware scheduler. Based on the size of the data samples and the configuration of each worker, Chronicaelaborately predicts the training time required for each worker. Chronicathen provides equivalent training time to each of the workers, alleviating both step- and epoch-level straggler problems. Furthermore, Chronicasuggests a new parameter synchronization scheme to achieve fast convergence based on the weighted average of the training workload on each worker. Our extensive evaluation using four deep learning models on 32 Amazon EC2 GPU instances showed that the new Chronicaachieves up to 3.19 times speedup over the state-of-the-art systems. Sanha Maeng, Gordon Euhyun Moon, Sungyong Park |
CCGrid | 3 |
| 2022 | Learning and Generalizing Cooperative Manipulation Skills Using Parametric Dynamic Movement PrimitivesabstractThis paper presents an approach that generates the overall trajectory of mobile manipulators for a complex mission consisting of several sub-tasks. Parametric dynamic movement primitives (PDMPs) can quickly generalize the online motion of robot manipulation by learning multiple demonstrations in offline. However, regarding complex missions consisting of multiple sub-tasks, a large number of demonstrations are required for full generalization, which is impractical. In this paper, we propose a framework that reduces the number of demonstrations for a complex mission. In the proposed method, complex demonstrations are segmented into multiple unit motions representing sub-tasks, and one PDMP is formed per each segment, resulting in multiple PDMPs. The phase decision process determines which sub-task and associated PDMPs to be executed online, allowing multiple PDMPs to be autonomously configured within an integrated framework. In order to generalize the execution time and regional goal in each phase, the Gaussian process regression (GPR) is applied. Simulation results from two different scenarios confirm that the proposed framework not only effectively reduces the number of demonstrations but also improves generalization performance. The actual experiments also demonstrate that the mobile manipulators effectively perform complex missions through the proposed framework. Note to Practitioners—This paper presents an approach of learning from demonstration (LfD) to generalize complex movements of robots. Parametric dynamic movement primitives (PDMPs) compute styles of movements from multiple demonstrations. However, the complexity of the PDMP increases as the mission involves more sub-tasks. In this paper, we resolve this issue by segmenting the complex mission into multiple sub-tasks and configuring multiple PDMPs. This work effectively reduces the number of required demonstrations for PDMPs, moderates the complexity of the algorithm. Also, the proposed approach allows flexible sub-task sequencing. It enables the mission in an unlearned sequence or a new combination of sub-tasks. The proposed approach is validated in both simulation and experimental results. Our approach is applicable for complex missions whose sub-tasks are clearly identified Hyoin Kim, Changsuk Oh, Inkyu Jang, Sungyong Park, Hoseong Seo, H. Jin Kim |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Model-based Domain Randomization of Dynamics System with Deep Bayesian Locally Linear EmbeddingabstractDomain randomization (DR) is a powerful tool to make a policy robust to the uncertainty of dynamics caused by unobservable environmental parameters. Conventional DR has adopted model-free reinforcement learning as a policy optimizer. However, the model-free methods in DR demand high time-complexity due to the randomization process where the environment is extremely changed. In this paper, we introduce model-based dynamics and policy learning for efficient DR. A Bayesian model of locally linear embedding is designed to fit the stochastic dynamics in DR. By virtue of locally linear dynamics, model-based optimal control is substituted for the policy optimization. Unlike previous works, our proposed Bayesian model with a MNIW prior allows the locally linear embedding to capture the dynamics in DR as a stochastic model. We show that a training method that combines variational and adversarial approaches is adequate for Bayesian embedding. Finally, a model-based controller is designed on our Bayesian locally linear embedding, and it shows better performance in DR environments compared with the non-Bayesian model of locally linear embedding. Jae Hyeon Park, Sungyong Park, H. Jin Kim |
ICRA | 2 |
| 2021 | A Probabilistic Machine Learning Approach to Scheduling Parallel Loops With Bayesian OptimizationabstractThis article proposes Bayesian optimization augmented factoring self-scheduling (BO FSS), a new parallel loop scheduling strategy. BO FSS is an automatic tuning variant of the factoring self-scheduling (FSS) algorithm and is based on Bayesian optimization (BO), a black-box optimization algorithm. Its core idea is to automatically tune the internal parameter of FSS by solving an optimization problem using BO. The tuning procedure only requires online execution time measurement of the target loop. In order to apply BO, we model the execution time using two Gaussian process (GP) probabilistic machine learning models. Notably, we propose a locality-aware GP model, which assumes that the temporal locality effect resembles an exponentially decreasing function. By accurately modeling the temporal locality effect, our locality-aware GP model accelerates the convergence of BO. We implemented BO FSS on the GCC implementation of the OpenMP standard and evaluated its performance against other scheduling algorithms. Also, to quantify our method's performance variation on different workloads, or workload-robustness in our terms, we measure the minimax regret. According to the minimax regret, BO FSS shows more consistent performance than other algorithms. Within the considered workloads, BO FSS improves the execution time of FSS by as much as 22% and 5% on average. Kyurai Kim, Youngjae Kim 0001, Sungyong Park |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | DISKSHIELD: A Data Tamper-Resistant Storage for Intel SGXabstractWith the increasing importance of data, the threat of malware which destroys data has been increasing. If malware acquires the highest software privilege, any attempt to detect and remove malware can be disabled. In this paper, we propose DISKSHIELD, a secure storage framework. DISKSHIELD uses Intel SGX to provide Trusted Execution Environment (TEE) to the host, implements the file system into SSD firmware that provides a Trusted Computing Base (TCB), and uses a two-way authentication mechanism to securely transfer data from the host TEE to the SSD TCB against data tampering attacks. This design frees DISKSHIELD from attacks to the kernel. To show the efficacy of DISKSHIELD, we prototyped a DISKSHIELD system by modifying Intel IPFS and developing a device file system on the Jasmine OpenSSD Platform in a Linux environment. Our results show that DISKSHIELD provides strong data tamper resistance the throughput of read and write is on average to 28%, 19% lower than IPFS. Jinwoo Ahn, Junghee Lee 0004, Yungwoo Ko, Donghyun Min, Jiyun Park, Sungyong Park, Youngjae Kim 0001 |
AsiaCCS | 6 |
| 2020 | Position: GPUKV: Towards a GPU-Driven Computing on Key-Value SSD
Min-Gyo Jeong, Chang-Gyu Lee, DongGyu Park, Sungyong Park, Youngjae Kim 0001, Jungki Noh, Woosuk Chung, Kyoung Park |
HotStorage | 4 |
| 2020 | A NUMA-aware NVM File System Design for Manycore Server ApplicationsabstractNOVA, a state-of-the-art NVM-based file system, is known to have scalability bottlenecks when multiple I/O threads read/write data simultaneously. Recent studies have identified the cause as the coarse-grained lock adopted by NOVA to provide consistency, and proposed fine-grained range-based locks to improve the scalability of NOVA. However, these variants of NOVA only scale on Uniform Memory Access (UMA) architecture and do not scale on Non-Uniform Memory Access (NUMA) architecture. This is because NOVA has no NUMA-aware memory allocation policy and still uses non-scalable file data structures. In this paper, we propose a NUMA-aware NOVA file system which virtualizes the NVM devices located across NUMA nodes so that they can be used as a single address space. The proposed file system adopts a local-first placement policy where file data and metadata are placed preferentially on the local NVM device to reduce the remote access problem. In addition, the lock-free per-core data structures proposed in this file system allow data to be updated concurrently while mitigating the remote memory access. Extensive evaluations show that our NUMA-aware NOVA for parallel writing is scalable with respect to the increased core count and outperforms vanilla NOVA by 2.56-19.18 times. June-Hyung Kim, Youngjae Kim 0001, Safdar Jamil, Sungyong Park |
MASCOTS | 4 |
| 2020 | Crocus: Enabling Computing Resource Orchestration for Inline Cluster-Wide Deduplication on Scalable Storage SystemsabstractInline deduplication dramatically improves storage space utilization. However, it degrades I/O throughput due to computeintensive deduplication operations such as chunking, fingerprinting or hashing of chunk content, and redundant lookup I/Os over the network in the I/O path. In particular, the fingerprint or hash generation of content contributes largely to the degraded I/O throughput and is computationally expensive. In this article, we propose CROCUS, a framework that enables compute resource orchestration to enhance cluster-wide deduplication performance. In particular, CROCUS takes into account all compute resources such as local and remote {CPU, GPU} by managing decentralized compute pools. An opportunistic Load-Aware Fingerprint Scheduler (LAFS), distributes and offloads compute-intensive deduplication operations in a load-aware fashion to compute pools. CROCUS is highly generic and can be adopted in both inline and offline deduplication with different storage tier configurations. We implemented CROCUS in Ceph scale-out storage system. Our extensive evaluation shows that CROCUS reduces the fingerprinting overhead by 86 percent with 4KB chunk size compared to Ceph with baseline deduplication while maintaining high disk-space savings. Our proposed LAFS scheduler, when tested in different internal and external contention scenarios also showed 54 percent improvement over a fixed or static scheduling approach. Prince Hamandawana, Awais Khan 0002, Chang-Gyu Lee, Sungyong Park, Youngjae Kim 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Towards Robust Data-Driven Parallel Loop Scheduling Using Bayesian OptimizationabstractEfficient parallelization of loops is critical to improving the performance of high-performance computing applications. Many classical parallel loop scheduling algorithms have been developed to increase parallelization efficiency. Recently, workload-aware methods were developed to exploit the structure of workloads. However, both classical and workload-aware scheduling methods lack what we call robustness. That is, most of these scheduling algorithms tend to be unpredictable in terms of performance or have specific workload patterns they favor. This causes application developers to spend additional efforts in finding the best suited algorithm or tune scheduling parameters. This paper proposes Bayesian Optimization augmented Factoring Self-Scheduling (BO FSS), a robust data-driven parallel loop scheduling algorithm. BO FSS is powered by Bayesian Optimization (BO), a machine learning based optimization algorithm. We augment a classical scheduling algorithm, Factoring Self-Scheduling (FSS), into a robust adaptive method that will automatically adapt to a wide range of workloads. To compare the performance and robustness of our method, we have implemented BO FSS and other loop scheduling methods on the OpenMP framework. A regret-based metric called performance regret is also used to quantify robustness. Extensive benchmarking results show that BO FSS performs fairly well in most workload patterns and is also very robust relative to other scheduling methods. BO FSS achieves an average of 4% performance regret. This means that even when BO FSS is not the best performing algorithm on a specific workload, it stays within a 4 percentage points margin of the best performing algorithm. Kyurai Kim, Youngjae Kim 0001, Sungyong Park |
MASCOTS | 3 |
| 2019 | iLSM-SSD: An Intelligent LSM-Tree Based Key-Value SSD for Data AnalyticsabstractSeveral key-value stores such as RocksDB and MongoDB are implemented on the file system using the Log-Structured Merge-Tree (LSM-tree). The LSM-tree involves high compaction overhead. To minimize this overhead, WiscKey, the state-of-the-art LSM-tree, separates key and value, appends the value to the Value Log file, and LSM-tree manages only the key and Value Log offset. This minimizes the compaction overhead by reducing the number of SSTables managed by the LSM-tree. However, WiscKey still has a high I/O stack overhead that must go through the OS file system and block-layer. Therefore, this paper proposes iLSM-SSD that implements WiscKey in SSD and supports near-data processing. iLSM-SSD has the following features: (i) iLSM-SSD implements a key-value separation based LSM-tree in a limited memory space inside the SSD. (ii) The Value Log offset update management overhead incurred during the Value Log cleaning has a significant performance impact on CPU and memory-constrained SSD environments. To minimize this overhead, iLSM-SSD implements Scattered Logging, which reuses invalidated Value Log pages on the Value Log. (iii) iLSM-SSD manages the data layout internally. This enables iLSM-SSD to eliminate the need for file system interactions to obtain the data layout for in-storage processing on traditional block-interface-based SSDs. We prototyped the iLSM-SSD on the Cosmos+ OpenSSD platform in a Linux environment. Extensive evaluations with synthetic benchmarks have shown that the PUT performance of iLSM-SSD is 1.6-4 times higher than that of WiscKey implemented in RocksDB. Chang-Gyu Lee, Hyeongu Kang, DongGyu Park, Sungyong Park, Youngjae Kim 0001, Jungki Noh, Woosuk Chung, Kyoung Park |
MASCOTS | 4 |
| 2018 | A Robust Fault-Tolerant and Scalable Cluster-Wide Deduplication for Shared-Nothing Storage SystemsabstractDeduplication has been largely employed in distributed storage systems to improve space efficiency. Traditional deduplication research ignores the design specifications of shared-nothing distributed storage systems such as no central metadata bottleneck, scalability, and storage rebalancing. Further, deduplication introduces transactional changes, which are prone to errors in the event of a system failure, resulting in inconsistencies in data and deduplication metadata. In this paper, we propose a robust, fault-tolerant and scalable cluster-wide deduplication that can eliminate duplicate copies across the cluster. We design a distributed deduplication metadata shard which guarantees performance scalability while preserving the design constraints of shared-nothing storage systems. The placement of chunks and deduplication metadata is made cluster-wide based on the content fingerprint of chunks. To ensure transactional consistency and garbage identification, we employ a flag-based asynchronous consistency mechanism. We implement the proposed deduplication on Ceph. The evaluation shows high disk-space savings with minimal performance degradation as well as high robustness in the event of sudden server failure. Awais Khan 0002, Chang-Gyu Lee, Prince Hamandawana, Sungyong Park, Youngjae Kim 0001 |
MASCOTS | 4 |
| 2017 | LAWC: Optimizing Write Cache Using Layout-Aware I/O Scheduling for All Flash StorageabstractFlash memory-based SSD-RAIDs are swiftly replacing conventional hard disk drives by exhibiting improved performance and stability, especially in I/O-intensive environments. However, the variations in latency and throughput occurring due to uncoordinated internal garbage collection cripples further boosting of performance. In addition, the unwanted variations in each SSD can influence the overall performance of the entire flash storage adversely. This performance bottleneck can be essentially reduced by an internal write cache in the RAID controller designed prudently by considering the crucial device characteristics. The state-of-the-art cache write for the RAID controller fails to incorporate device characteristics of flash memory-based SSDs and mitigates the performance gain. In this paper, we propose a novel cache design namely Layout-Aware Write Cache (LAWC) to overcome the performance barrier inculcated by independent garbage collections. LAWC implements (i) improved I/O scheduling for logically partitioned write caches, (ii) a destage write synchronization mechanism to allow individual write caches to flush write blocks into the SSD array in a coordinated manner, and (iii) a two-level hybrid cache algorithm utilizing small front level cache for the improved write cache efficiency. LAWC shows significant reduction in response time by 82.39 percent on RAID-0 and 68.51 percent on RAID-5 types of SSDs when compared with state-of-theart write cache algorithms. Kalidas Ganesh, Youngjae Kim 0001, Monobrata Debnath, Sungyong Park, Junghee Lee 0004 |
IEEE Trans. Computers | 4 |
| 2016 | Design and Analysis of Fault Tolerance Mechanisms for Big Data TransfersabstractIncreased growth of the data and the need to move the data between data centers, demands, an efficient data transfer tool which can, not only transfer the data at higher rates but also handle the faults occurred during the transfer. Absence of fault tolerance mechanisms, would need to retransmit the whole data, in case of any error during the transfer. Hence, fault tolerance is an important aspect of big data transfer tools. In this paper, we have considered LADS data transfer tool, which proved to be superior to the existing data transfer tools with respect to the speed of the transfer. However, absence of fault tolerance implementation in LADS might result in data retransmission and congestion issues upon errors. In this paper, we design and analyze fault tolerance mechanisms which can be used with LADS data transfer tool. We have proposed three different fault tolerance mechanisms, File logging, Transaction logging and Universal logging. Also, we have analyzed the space and performance overhead of these fault tolerance mechanisms on LADS data transfer tool. Preethika Kasu, Youngjae Kim 0001, Sungyong Park, Scott Atchley, Geoffroy Vallée |
CLUSTER | 3 |
| 2016 | Implementation of a large-scale language model adaptation in a cloud environment
Kwang-Ho Kim, Dae-Young Jung, Donghyun Lee 0001, Hyuk-Jun Lee, Sungyong Park, Myoung-Wan Koo, Jeong-Sik Park, Hyung-Bae Jeon |
Multim. Tools Appl. | 5 |
| 2016 | Android RMI: a user-level remote method invocation mechanism between Android devices
Hee-Eun Kang, Kihyun Jeong, Kwonyong Lee, Sungyong Park, Youngjae Kim 0001 |
J. Supercomput. | 4 |
| 2015 | A QoS Assured Network Service Chaining Algorithm in Network Function Virtualization ArchitectureabstractIn the Network Function Virtualization (NFV) architecture, Network Service Chaining (NSC) is consisted in a certain order of network elements so that it can provide flexible network services to users. Due to the complexity of network infrastructure, creating a service chain requires high operation cost especially in carrier-grade network service providers and supporting stringent QoS requirements is also a challenge. Although several vendors provide various solutions for the NSC, there is only few information and the detailed algorithm or implementation logic is hidden. This paper presents an NSC algorithm in NFV that assures QoS from the perspective of service providers. In order to formulate NSC path selection problem, we apply the NP complete genetic algorithm. The evaluation results show that the proposed algorithm minimizes the operation cost of service providers by approximately 10.6% while the requested QoS targets is not violated. Taekhee Kim, Siri Kim, Kwonyong Lee, Sungyong Park |
CCGRID | 4 |
| 2013 | A Service Path Selection and Adaptation Algorithm in Service-Oriented Network Virtualization ArchitectureabstractWith the advances of networking and cloud computing technologies, many web services are often implemented as cloud services by different service providers, and some SPs compose those services to create an end-to-end network service. In this service-oriented network virtualization environment, two important but somehow contradictory technical challenges are 1) how to select an optimal service path in order to meet the QoS requirements by users, and 2) how to balance the loads to fully utilize the whole systems. In this paper, we present an adaptive service path selection algorithm in service-oriented network virtualization environment, which assures QoS and balances the load at the same time. The proposed algorithm selects appropriate component services from multiple candidate service instances running on virtual machines and creates an optimal service path satisfying the users' QoS requirements. Our algorithm also dynamically adapts to the optimal service path based on various changes. Since the proposed algorithm handles a variation of the multi-constrained path selection problem known as NP-complete, we formulate the problem using ant colony optimization algorithm. The experimental results show that our algorithm guarantees the maximal QoS to the users while balancing the load at the same time, and adapts to the optimal service path based on the situational changes. Kwonyong Lee, Haerim Yoon, Sungyong Park |
ICPADS | 3 |
| 2009 | A QoS Based Migration Scheme for Virtual Machines in Data Center Environments
Saeyoung Han, Jinseok Kim 0003, Sungyong Park |
APNOMS | 3 |
| 2006 | A New Polynomial Time Algorithm for Bayesian Network Structure Learning
Sanghack Lee, Jihoon Yang, Sungyong Park |
ADMA | 3 |
| 2005 | Unit Volume Based Distributed Clustering Using Probabilistic Mixture Model
Keunjoon Lee, Jinu Joo, Jihoon Yang, Sungyong Park |
Discovery Science | 4 |
| 2005 | A Content-Based Load Balancing Algorithm for Metadata Servers in Cluster File Systems
Junho Jang, Saeyoung Han, Sungyong Park, Jihoon Yang |
ISPA | 3 |
| 2005 | A Video Streaming System for Mobile Phones: Practice and Experience
Hojung Cha, Jongmin Lee 0001, Jongho Nang, Sungyong Park, Jin-Hwan Jeong, Chuck Yoo |
Wirel. Networks | 4 |
| 2004 | Discovery of Hidden Similarity on Collaborative Filtering to Overcome Sparsity Problem
Sanghack Lee, Jihoon Yang, Sungyong Park |
Discovery Science | 3 |
| 2004 | Modeling and Dynamic Control of Compliant Framed wheeled Modular Mobile RobotsabstractDynamic models and controllers for compliant framed wheeled modular mobile robots are studied in this paper. This is a new type of wheeled mobile robot using axle and frame modules that can provide both full suspension and enhanced steering capability without additional hardware. In this research modular kinematic and dynamic models of the modules are developed and assembled into a scalable dynamic system model such that a wide variety of configurations could be described. Dynamic control is then achieved by coordinated control of the individual wheel torques as specified by a backstepping controller scaled to the dimension of the system configuration. These results are applied to a two-axle scout case study in order to demonstrate their implementation and performance. Simulation and experimental results illustrate dynamic control of trajectory tracking while path following. Sungyong Park, Mark A. Minor |
ICRA | 1 |
| 2004 | MPICH-GP: A Private-IP-Enabled MPI Over Grid Environments
Kumrye Park, Sungyong Park, Oh-Young Kwon, Hyoung-Woo Park |
ISPA | 2 |
| 2004 | An Adaptive Proximity Route Selection Scheme in DHT-Based Peer to Peer Systems
Jiyoung Song, Sungyong Park, Jihoon Yang |
PDCAT | 2 |
| 2002 | Email Categorization Using Fast Machine Learning Algorithms
Jihoon Yang, Sungyong Park |
Discovery Science | 2 |
| 1998 | A Multithreaded Message-Passing System for High-Performance Distributed Computing ApplicationsabstractNYNET (ATM wide area network testbed in New York state) Communication System (NCS) is a multithreaded message passing system developed at Syracase University that provides high performance and flexible communication services over asynchronous transfer mode (ATM) based high performance distributed computing (HPDC) environments. NCS capitalizes on thread based programming model to overlap computations and communications, and develop a dynamic message passing environment with separate data and control paths. This leads to a flexible and adaptive message passing environment that can support multiple flow control, error control, and multicasting algorithms. We provide an overview of the NCS architecture and present how NCS point to point communication services are implemented. We also analyze the overhead incurred by using multithreading and compare the performance of NCS point to point communication primitives with those of other message passing systems such as p4, PVM, and MPI. Benchmarking results indicate that NCS shows comparable performance to other systems for small message sizes but outperforms other systems for large message sizes. Sungyong Park, Joohan Lee, Salim Hariri |
ICDCS | 1 |
| 1997 | I/O and memory-efficient matrix multiplication with user-controllable parallel I/OabstractThe UPIO (user-controllable parallel I/O) proposed by the authors in 1996 allows users to determine a file's structure by considering the access patterns of particular applications and the distribution of data for parallel access, and them do I/O collectively. This enables users to produce high-performance external computation codes by planning I/O, computations, communication, and the reuse of data effectively in the codes. They show how well UPIO produces high performance external computation codes by designing I/O and memory-efficient external matrix multiplication algorithms and exploring the effects of UPIO with the codes. Jang Sun Lee, Sungyong Park, P. Bruce Berra, Sanjay Ranka |
ICPADS | 2 |
| 1997 | A High Performance Message-Passing System for Network of Workstations
Sungyong Park, Salim Hariri |
J. Supercomput. | 1 |
| 1996 | NYNET Communication System (NCS): A Multithreaded Message Passing Tool over ATM NetworkabstractCurrent advances in processor technology, and the rapid development of high speed networking technology, such as ATM, have made high performance network computing an attractive computing environment for large-scale high performance distributed computing (HPDC) applications. However, due to the communications overhead at the host-network interface, most of the HPDC applications are not getting the full benefit of high speed communication networks. This overhead can be attributed to the high cost of operating system calls, context switching, the use of inefficient communication protocols, and the coupling of data and control paths. We present an architecture and implementation for a low-latency, high-throughput message passing tool, that we refer to as the NYNET (ATM wide area network testbed in New York state) Communication System (NCS), which can support a variety of HPDC applications with different Quality of Services (QOS) requirements. NCS uses multithreading to provide efficient techniques that overlap computation and communication. NCS uses read/write trap routines to bypass traditional operating system calls. This reduces latency and avoids using inefficient communication protocols. By separating data and control paths, NCS eliminates unnecessary control transfers. This optimizes the data path and improves performance. Benchmarking results show that the performance of NCS is at least a factor of two better than the performance of corresponding p4 and PVM primitives. Sungyong Park, Salim Hariri, Yoonhee Kim, J. Stuart Harris, Rajesh Yadav |
HPDC | 1 |
| 1995 | Software Tool Evaluation MethodologyabstractThe recent development of parallel and distributed computing software has introduced a variety of software tools that support several programming paradigms and languages. This variety of tools makes the selection of the best tool to run a given class of applications on a parallel or distributed system a non-trivial task that requires some investigation. We expect tool evaluation to receive more attention as the deployment and usage of distributed systems increases. In this paper, we present a multi-level evaluation methodology for parallel/distributed tools in which tools are evaluated from different perspectives. We apply our evaluation methodology to three message passing tools viz Express, p4, and PVM. The approach covers several important distributed systems platforms consisting of different computers (e.g., IBM-SP1, Alpha cluster, SUN workstations) interconnected by different types of networks (e.g., Ethernet, FDDI, ATM). Salim Hariri, Sungyong Park, Rajashekar Reddy, Mahesh Subramanyan, Rajesh Yadav, Geoffrey C. Fox, Manish Parashar |
ICDCS | 2 |
| 1994 | A Concurrent Multi Target Tracker: Benchmarking and PortabilityabstractWith the current advances in computing and network technology and software, the gap between parallel and distributed computing environment is gradually becoming narrower. Consequently, parallel programs run on parallel as well as distributed systems. However, programming and porting complex applications to such environment is challenging task and not well understood. In this paper, we use a concurrent multi target tracker as a running example to analyze and evaluate performance of two different parallel implementations on parallel and distributed systems. We have benchmarked both these implementations on different architectures that vary from a network of worksta-tions{SUN, IBM RS6000) to parallel computers (CM5, iPSC 860) using different parallel/distributed message passing tools{PVM, p4, EXPRESS). Salim Hariri, Rajesh Yadav, Balaji Thiagarajan, Sungyong Park, Mahesh Subramanyan, Rajashekar Reddy, Geoffrey C. Fox |
ICPP (3) | 4 |