Suhyun Lee 0002

dblp:321/4116-2 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-0501-794XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 DMO-DB: Mitigating the Data Movement Bottlenecks of GPU-Accelerated Relational OLAP
abstract
Graphics Processing Units (GPUs) offer high computational throughput and memory bandwidth, making them promising accelerators for relational OnLine Analytical Processing (OLAP). GPU-accelerated relational OLAP executes the relational operations of a Structured Query Language (SQL) query on a GPU instead of the host Central Processing Unit (CPU). Depending on where input columns and their values reside in, a GPU-accelerated SQL query execution can be classified into two scenarios: 1) in-GPU, in which all input columns fit in the GPU memory, or 2) in-host, in which the input columns reside in the host memory and get transferred to the GPU memory when needed. However, both scenarios incur significant intraGPU and host-to-GPU data movement overheads, respectively. In-GPU executions incur excessive GPU cache misses and thus frequent off-chip GPU memory accesses. In-host executions suffer from the limited host-to-GPU data transfer bandwidth. This paper presents DMO-DB, a Data Movement-Optimized GPU-accelerated relational OLAP engine. Since modern GPUaccelerated relational OLAP decomposes SQL queries into multiple pipelines-sequences of relational operations that can be executed on input columns from the same table, DMO-DB leverages inter-pipeline dependencies to overcome the two data movement bottlenecks. DMO-DB introduces two key ideas: cache-fit bloom filtering and Ahead-of-Time value Discarding (AoTD), which preemptively eliminate unnecessary input values before their movement across the memory hierarchies. For in-GPU execution, GPU L1 data cache-fit filters discard non-contributing values before triggering costly off-chip DRAM accesses. For in-host execution, host CPU last level cache-fit filters strategically prune unnecessary input values, minimizing PCIe transfer overhead. After that, AoTD exploits multiple inter-pipeline dependencies by collecting these cache-fit bloom filters to earlier pipeline execution stages. Our evaluation using NVIDIA RTX A4000 and TITAN RTX GPUs shows that DMO-DB achieves speedups of $\mathbf{1. 5 3 x}$ over in-GPU Crystal-Opt and 6.10x over in-host HeavyDB.
Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Joonsung Kim 0001, Jinho Lee 0001, Youngsok Kim
PACT2
2024 SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs
abstract
Recent advances in Dual In-line Memory Modules (DIMMs) allow DIMMs to support Processing-In-DIMM (PID) by placing In-DIMM Processors (IDPs) near their memory banks. Prior studies have shown that in-memory joins can benefit from PID by offloading their operations onto the IDPs and exploiting the high internal memory bandwidth of DIMMs. Aimed at evenly balancing the computational loads between the IDPs, the existing algorithms perform IDP-wise global partitioning on input tables and then make each IDP process a partition of the input tables. Unfortunately, we find that the existing PID join algorithms achieve low performance and scalability with skewed input tables. With skewed input tables, the IDP-wise global partitioning incurs imbalanced loads between the IDPs, making the IDPs remain idle until the heaviest-load IDP completes processing its partition. To fully exploit the IDPs for accelerating in-memory joins involving skewed input tables, therefore, we need a new PID join algorithm which achieves high skew resistance by mitigating the imbalanced inter-IDP loads. In this paper, we present SPID-Join, a skew-resistant PID join algorithm which exploits two parallelisms inherent in DIMM architectures, namely bank- and rank-level parallelisms. By replicating join keys across the banks within a rank and across ranks, SPID-Join significantly increases the internal memory bandwidth and computational throughput allocated to each join key, improving the load balance between the IDPs and accelerating join executions. SPID-Join exploits the bank- and the rank-level parallelisms to minimize join key replication overheads and support a wider range of join key replication ratios. Despite achieving high skew resistance, SPID-Join exhibits a trade-off between the join key replication ratio and the join execution latency, making the best-performing join key replication ratio depend on join and PID system configurations. We, therefore, augment SPID-Join with a cost model which identifies the best-performing join key replication ratio for given join and PID system configurations. By accurately modeling and scaling the IDPs' throughput and the inter-IDP communication bandwidth, the cost model accurately captures the impact of the join key replication ratio on SPID-Join. Our experimental results using eight UPMEM DIMMs, which collectively provide a total of 1,024 IDPs, show that SPID-Join achieves up to 10.38x faster join executions over PID-Join, the state-of-the-art PID join algorithm, with highly skewed input tables.
Suhyun Lee 0002, Chaemin Lim, Jinwoo Choi 0003, Heelim Choi, Yongjun Park 0001, Kwanghyun Park 0001, Hanjun Kim 0001, Youngsok Kim
Proc. ACM Manag. Data1
2023 Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMs
abstract
Modern dual in-line memory modules (DIMMs) support processing-in-memory (PIM) by implementing in-DIMM processors (IDPs) located near memory banks. PIM can greatly accelerate in-memory join, whose performance is frequently bounded by main-memory accesses, by offloading the operations of join from host central processing units (CPUs) to the IDPs. As real PIM hardware has not been available until very recently, the prior PIM-assisted join algorithms have relied on PIM hardware simulators which assume fast shared memory between the IDPs and fast inter-IDP communication; however, on commodity PIM-enabled DIMMs, the IDPs do not share memory and demand the CPUs to mediate inter-IDP communication. Such discrepancies in the architectural characteristics make the prior studies incompatible with the DIMMs. Thus, to exploit the high potential of PIM on commodity PIM-enabled DIMMs, we need a new join algorithm designed and optimized for the DIMMs and their architectural characteristics. In this paper, we design and analyze Processing-In-DIMM Join (PID-Join), a fast in-memory join algorithm which exploits UPMEM DIMMs, currently the only publicly-available PIM-enabled DIMMs. The DIMMs impose several key challenges on efficient acceleration of join including the shared-nothing nature and limited compute capabilities of the IDPs, the lack of hardware support for fast inter-IDP communication, and the slow IDP-wise data transfers between the IDPs and the main memory. PID-Join overcomes the challenges by prototyping and evaluating hash, sort-merge, and nested-loop algorithms optimized for the IDPs, enabling fast inter-IDP communication using host CPU cache streaming and vector instructions, and facilitating fast rank-wise data transfers between the IDPs and the main memory. Our evaluation using a real system equipped with eight UPMEM DIMMs and 1,024 IDPs shows that PID-Join greatly improves the performance of in-memory join over various CPU-based in-memory join algorithms.
Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Jounghoo Lee, Seongyeon Park, Hanjun Kim 0001, Jinho Lee 0001, Youngsok Kim
Proc. ACM Manag. Data2
2022 GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUs
abstract
Analytical models can greatly help computer architects perform orders of magnitude faster early-stage design space exploration than using cycle-level simulators. To facilitate rapid design space exploration for graphics processing units (GPUs), prior studies have proposed GPU analytical models which capture first-order stall events causing performance degradation; however, the existing analytical models cannot accurately model modern GPUs due to their outdated and highly abstract GPU core microarchitecture assumptions. Therefore, to accurately evaluate the performance of modern GPUs, we need a new GPU analytical model which accurately captures the stall events incurred by the significant changes in the core microarchitectures of modern GPUs.
Jounghoo Lee, Yeonan Ha, Suhyun Lee 0002, Jinyoung Woo, Jinho Lee 0001, Hanhwi Jang, Youngsok Kim
ISCA3
2022 GuardiaNN: Fast and Secure On-Device Inference in TrustZone Using Embedded SRAM and Cryptographic Hardware
abstract
As more and more mobile/embedded applications employ Deep Neural Networks (DNNs) involving sensitive user data, mobile/embedded devices must provide a highly secure DNN execution environment to prevent privacy leaks. Aimed at securing DNN data, recent studies execute part of a DNN in a trusted execution environment (e.g., TrustZone) to isolate DNN execution from the other processes; however, as the trusted execution environments for mobile/embedded devices provide limited memory protection, DNN data remain unencrypted in DRAM and become vulnerable to physical attacks. The devices can prevent the physical attacks by keeping DNN data encrypted in DRAM; when DNN data get referenced during DNN execution, they get loaded to the SRAM and get decrypted by a CPU core. Unfortunately, using the SRAM with demand paging greatly increases DNN execution time due to the inefficient use of the SRAM and the high CPU consumption of data encryption/decryption.
Jinwoo Choi 0003, Jaeyeon Kim, Chaemin Lim, Suhyun Lee 0002, Jinho Lee 0001, Dokyung Song, Youngsok Kim
Middleware4