EDBT 2026 Demo / reviewers in the wild / expert
Seongyeon Park
dblp:324/7219
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0007-3480-6626ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime AdaptationabstractDynamic random walks are fundamental to various graph analysis applications, offering advantages by adapting to evolving graph properties. Their runtime-dependent transition probabilities break down the pre-computation strategy that underpins most existing CPU and GPU static random walk optimizations. This leaves practitioners suffering from suboptimal frameworks and having to write hand-tuned kernels that do not adapt to workload diversity. To handle this issue, we present FlexiWalker, the first GPU framework that delivers efficient, workload-generic support for dynamic random walks. Our design-space study shows that rejection sampling and reservoir sampling are more suitable than other sampling techniques under massive parallelism. Thus, we devise (i) new high-performance kernels for them that eliminate global reductions, redundant memory accesses, and random-number generation. Given the necessity of choosing the best-fitting sampling strategy at runtime, we adopt (ii) a lightweight first-order cost model that selects the faster kernel per node at runtime. To enhance usability, we introduce (iii) a compile-time component that automatically specializes user-supplied walk logic into optimized building blocks. On various dynamic random walk workloads with real-world graphs, FlexiWalker outperforms the best published CPU/GPU baselines by geometric means of 73.44× and 5.91×, respectively, while successfully executing workloads that prior systems cannot support. We open-source FlexiWalker in https://github.com/AIS-SNU/FlexiWalker. Seongyeon Park, Jaeyong Song 0002, Changmin Shin 0002, Sukjin Kim, Junguk Hong, Jinho Lee 0001 |
EuroSys | 1 |
| 2026 | LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMabstractLookup tables (LUTs) have recently gained attention as an alternative compute mechanism that maps input operands to precomputed results, eliminating the need for arithmetic logic. LUTs not only reduce logic complexity, but also naturally support diverse numerical precisions without requiring separate circuits for each bitwidth—an increasingly important feature in quantized DNNs. This creates a favorable tradeoff in PIM: memory capacity can be used in place of logic to increase computational throughput, aligning well with DRAM-PIM architectures that offer high bandwidth and easily available memory but limited logic density. In this work, we explore this capacity-computation tradeoff in LUT-based PIM designs, where memory capacity is traded for performance by packing multiple MAC operations into a single LUT lookup. Building on this insight, we propose LoCaLUT, a PIM-based design for efficient low-bit quantized DNN inference using operation-packed LUTs. First, we observe that these LUTs contain extensive redundancy and introduce LUT canonicalization, which eliminates duplicate entries to reduce LUT size. Second, we propose reordering LUT, a lightweight auxiliary LUT that remaps weight vectors to their canonical form required by LUT canonicalization with a simple LUT lookup. Third, we propose LUT slice streaming, a novel execution strategy that exploits the DRAM-buffer hierarchy by streaming only relevant LUT columns into the buffer and reusing them across multiple weight vectors. Evaluated on a real system based on UPMEM devices, we demonstrate a geometric mean speedup of$1.82 \times$across various numeric precisions and DNN models. We believe LoCaLUT opens a path toward scalable, low-logic PIM designs tailored for LUT-based DNN inference. Our implementation of LoCaLUT is available at https://github.com/AIS-SNU/LoCaLUT. Junguk Hong, Changmin Shin 0002, Sukjin Kim, Si Ung Noh, Taehee Kwon 0002, Seongyeon Park, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
HPCA | 6 |
| 2025 | G^3SA: A GPU-Accelerated Gold Standard Genomics Library for End-to-End Sequence AlignmentabstractSequence alignment is the first step of the bioinformatics pipelines for analyzing DNA sequences in genomics.As the DNA sequencer throughput is dramatically increasing, the pipeline bottleneck has now shifted to the massive amount of sequence alignment computation.GPU-based acceleration has emerged as a promising solution to this challenge, and several efforts have focused on accelerating core algorithmic components.However, when integrating existing GPUaccelerated components into BWA-MEM and Minimap2, two widely used tools for short and long reads respectively, we identify several unresolved issues.One could naively attempt to gather and/or fix those libraries, but such an approach would result in a severe slowdown despite non-trivial effort.To address these issues, we propose 𝐺 3 𝑆𝐴, the first GPU acceleration with efficient parallelization and optimization strategies for the end-to-end gold standard sequence alignment algorithms for short and long reads.We observe that the primary bottlenecks stem from irregular and redundant computation and memory access patterns.Therefore, we design our kernels to fully exploit the memory hierarchy and computational capabilities of GPUs, reducing costly computation and memory access.Furthermore, following the seed-chain-extend flow of standard sequence alignment algorithms, we propose techniques tailored to specific stages in the pipeline.Using these techniques, 𝐺 3 𝑆𝐴 achieves significant end-to-end speedup against the CPU and GPU baselines. Yeejoo Han, Seongyeon Park, Jinho Lee 0001 |
ICS | 3 |
| 2025 | PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
Sukjin Kim, Seongyeon Park, Si Ung Noh, Junguk Hong, Taehee Kwon 0002, Hunseong Lim, Jinho Lee 0001 |
USENIX ATC | 2 |
| 2024 | PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesabstractRecent dual in-line memory modules (DIMMs) are starting to support processing-in-memory (PIM) by associating their memory banks with processing elements (PEs), allowing applications to overcome the data movement bottleneck by offloading memory-intensive operations to the PEs. Many highly parallel applications have been shown to benefit from these PIM-enabled DIMMs, but further speedup is often limited by the huge overhead of inter-PE collective communication. This mainly comes from the slow CPU-mediated inter-PE communication methods, making it difficult for PIM-enabled DIMMs to accelerate a wider range of applications. Prior studies have tried to alleviate the communication bottleneck, but they lack enough flexibility and performance to be used for a wide range of applications. In this paper, we present PID-Comm, a fast and flexible inter-PE collective communication framework for commodity PIM-enabled DIMMs. The key idea of PID-Comm is to abstract the PEs as a multi-dimensional hypercube and allow multiple instances of inter-PE collective communication between the PEs belonging to certain dimensions of the hypercube. Leveraging this abstraction, PID-Comm first defines eight interPE collective communication patterns that allow applications to easily express their complex communication patterns. Then, PIDComm provides high-performance implementations of the interPE collective communication patterns optimized for the DIMMs. Our evaluation using 16 UPMEM DIMMs and representative parallel algorithms shows that PID-Comm greatly improves the performance by up to $5.19 \times$ compared to the existing inter-PE communication implementations. The implementation of PIDComm is available at https://github.com/AIS-SNU/PID-Comm. Si Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park, Jeehyun Kim, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
ISCA | 4 |
| 2024 | AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingabstractWith the advance in genome sequencing technology, the lengths of deoxyribonucleic acid (DNA) sequencing results are rapidly increasing at lower prices than ever. However, the longer lengths come at the cost of a heavy computational burden on aligning them. For example, aligning sequences to a human reference genome can take tens or even hundreds of hours. The current de facto standard approach for alignment is based on the guided dynamic programming method. Although this takes a long time and could potentially benefit from high-throughput graphic processing units (GPUs), the existing GPU-accelerated approaches often compromise the algorithm's structure, due to the GPU-unfriendly nature of the computational pattern. Unfortunately, such compromise in the algorithm is not tolerable in the field, because sequence alignment is a part of complicated bioinformatics analysis pipelines. In such circumstances, we propose AGAThA, an exact and efficient GPU-based acceleration of guided sequence alignment. We diagnose and address the problems of the algorithm being unfriendly to GPUs, which comprises strided/redundant memory accesses and workload imbalances that are difficult to predict. According to the experiments on modern GPUs, AGAThA achieves 18.8× speedup against the CPU-based baseline, 9.6× against the best GPU-based baseline, and 3.6× against GPU-based algorithms with different heuristics. Seongyeon Park, Junguk Hong, Jaeyong Song 0002, Hajin Kim, Youngsok Kim, Jinho Lee 0001 |
PPoPP | 1 |
| 2023 | Unsupervised Pre-Training for Data-Efficient Text-to-Speech on Low Resource LanguagesabstractNeural text-to-speech (TTS) models can synthesize natural human speech when trained on large amounts of transcribed speech. How-ever, collecting such large-scale transcribed data is expensive. This paper proposes an unsupervised pre-training method for a sequence-to-sequence TTS model by leveraging large untranscribed speech data. With our pre-training, we can remarkably reduce the amount of paired transcribed data required to train the model for the target downstream TTS task. The main idea is to pre-train the model to reconstruct de-warped mel-spectrograms from warped ones, which may allow the model to learn proper temporal assignment relation between input and output sequences. In addition, we propose a data augmentation method that further improves the data efficiency in finetuning. We empirically demonstrate the effectiveness of our proposed method in low-resource language scenarios, achieving outstanding performance compared to competing methods. The code and audio samples are available at: https://github.com/cnaigithub/SpeechDewarping Seongyeon Park, Myungseo Song, Bohyung Kim, Tae-Hyun Oh |
ICASSP | 1 |
| 2023 | Automatic Tuning of Loss Trade-offs without Hyper-parameter Search in End-to-End Zero-Shot Speech Synthesis
Seongyeon Park, Bohyung Kim, Tae-Hyun Oh |
INTERSPEECH | 1 |
| 2023 | Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMsabstractModern dual in-line memory modules (DIMMs) support processing-in-memory (PIM) by implementing in-DIMM processors (IDPs) located near memory banks. PIM can greatly accelerate in-memory join, whose performance is frequently bounded by main-memory accesses, by offloading the operations of join from host central processing units (CPUs) to the IDPs. As real PIM hardware has not been available until very recently, the prior PIM-assisted join algorithms have relied on PIM hardware simulators which assume fast shared memory between the IDPs and fast inter-IDP communication; however, on commodity PIM-enabled DIMMs, the IDPs do not share memory and demand the CPUs to mediate inter-IDP communication. Such discrepancies in the architectural characteristics make the prior studies incompatible with the DIMMs. Thus, to exploit the high potential of PIM on commodity PIM-enabled DIMMs, we need a new join algorithm designed and optimized for the DIMMs and their architectural characteristics. In this paper, we design and analyze Processing-In-DIMM Join (PID-Join), a fast in-memory join algorithm which exploits UPMEM DIMMs, currently the only publicly-available PIM-enabled DIMMs. The DIMMs impose several key challenges on efficient acceleration of join including the shared-nothing nature and limited compute capabilities of the IDPs, the lack of hardware support for fast inter-IDP communication, and the slow IDP-wise data transfers between the IDPs and the main memory. PID-Join overcomes the challenges by prototyping and evaluating hash, sort-merge, and nested-loop algorithms optimized for the IDPs, enabling fast inter-IDP communication using host CPU cache streaming and vector instructions, and facilitating fast rank-wise data transfers between the IDPs and the main memory. Our evaluation using a real system equipped with eight UPMEM DIMMs and 1,024 IDPs shows that PID-Join greatly improves the performance of in-memory join over various CPU-based in-memory join algorithms. Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Jounghoo Lee, Seongyeon Park, Hanjun Kim 0001, Jinho Lee 0001, Youngsok Kim |
Proc. ACM Manag. Data | 5 |
| 2022 | SALoBa: Maximizing Data Locality and Workload Balance for Fast Sequence Alignment on GPUsabstractSequence alignment forms an important backbone in many sequencing applications. A commonly used strategy for sequence alignment is an approximate string matching with a two-dimensional dynamic programming approach. Although some prior work has been conducted on GPU acceleration of a sequence alignment, we identify several shortcomings that limit exploiting the full computational capability of modern GPUs. This paper presents SALoBa, a GPU-accelerated sequence alignment library focused on seed extension. Based on the analysis of previous work with real-world sequencing data, we propose techniques to exploit the data locality and improve work-load balancing. The experimental results reveal that SALoBa significantly improves the seed extension kernel compared to state-of-the-art GPU-based methods. Seongyeon Park, Hajin Kim, Nauman Ahmed, Zaid Al-Ars, H. Peter Hofstee, Youngsok Kim, Jinho Lee 0001 |
IPDPS | 1 |