VLDB 2026 Research / reviewers in the wild / expert
Seongyoung Kang
dblp:271/6724
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0005-4751-7976ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA Acceleration
Seongyoung Kang, Se-Min Lim, Sang Woo Jun |
ISCA | 1 |
| 2025 | Bancroft: Genomics Acceleration Beyond On-Device MemoryabstractThis paper presents Bancroft, a computational genomics acceleration platform for processing datasets far exceeding accelerator memory capacity. Bancroft overcomes the capacity limitations of accelerator memory by storing genomic data in a novel genomic compression format on the host server, and decompressing it on-demand within the accelerator after fetching it over PCIe. The key innovation of Bancroft is algorithmic optimizations for reference-based compression and decompression of popular genomic file formats. The algorithm achieves high enough compression ratios to improve the effective bandwidth of the PCIe to DRAM-levels, while facilitating highly efficient hardware implementations. We evaluate a prototype implementation of Bancroft on an affordable Alveo U50 FPGA accelerator card equipped with 8 GB of High-Bandwidth Memory (HBM). Our evaluation demonstrates that Bancroft delivers speeds exceeding on-device DDR4 memory and over 30% of HBM performance, while incurring only 30% chip space overhead. This is an order of magnitude higher performance and efficiency compared to conventional PCIe-limited architectures. Using a real-world pre-alignment filtering application, Bancroft demonstrates over $7 \times$ performance improvement over conventional accelerators on scalable datasets. Se-Min Lim, Seongyoung Kang, Sang Woo Jun |
PACT | 2 |
| 2025 | Labidus: RISC-V Overlay with Streaming Asynchronous Custom InstructionsabstractHigh development complexity is one of the most critical issues preventing more widespread use of reconfigurable hardware accelerators such as FPGAs. While soft processor overlays allow productive development with high-level software tools, they suffer from low performance. We address this issue with Labidus, a parallel RISC-V soft processor overlay that addresses this performance gap while maintaining development simplicity. Labidus automatically generates custom instructions based on static analysis of user software. These custom instructions achieve high utilization through two key innovations: asynchronous semantics and sharing across four-core tiles. Our evaluation across four scientific computing applications shows that Labidus matches or exceeds the performance of even manually optimized FPGA accelerators for gigabyte-scale tasks, at a fraction of development effort. Gongjin Sun, Seongyoung Kang, Jane He, Se-Min Lim, Sang Woo Jun |
ASAP | 2 |
| 2024 | Sting: Near-storage accelerator framework for scalable triangle counting and beyondabstractOne of the most critical limitations to scalable graph mining is memory capacity, as graphs of interest continue to grow while the rate of DRAM scaling diminishes. While high-performance NVMe storage is cheap and dense enough to better support larger graphs, the relative performance limitations of secondary storage force a cost-performance trade-off. We present STING, which uses an asynchronous callback function to provide a general interface to in-storage graphs while allowing transparent near-storage acceleration. Using triangle counting, we show with transparent filtering and sorting acceleration, STING can improve state-of-the-art by 3x for cost and power efficiency. Seongyoung Kang, Sang Woo Jun |
DAC | 1 |
| 2022 | BunchBloomer: Cost-Effective Bloom Filter Accelerator for Genomics ApplicationsabstractBloom filters are a very important tool for many applications including genomics, where they are used as a compact data structure for counting k-mers, represent de Bruijn graphs, and more. Due to their random-access nature coupled with the large size required for genomics, Bloom filters for genomics can easily become bound by the random access performance of off-chip memory. This is especially true for accelerators such as FPGAs and GPUs, which can easily remove the computation overhead of the multiple hash functions. As a result, Bloom filter accelerators have typically focused either on small filters which can fit in fast on-chip memory, or require fast off-chip memory fabric such as Hybrid Memory Cubes. In this work, we present BunchBloomer, which improves the cost-effectiveness of FPGA Bloom filter accelerators by making better use of cheaper, lower-power DDR memory. BunchBloomer uses a multi-layer radix sorter to group table updates into bursts directed to the same 8 KiB memory region, which can be efficiently cached in on-chip memory. A single BunchBloomer device outperforms a costly 12-core server by over 2×, demonstrating an order of magnitude better power efficiency. It even achieves better power efficiency compared to published FPGA Bloom filter accelerators equipped with Hybrid Memory Cubes. Seongyoung Kang, Tarun Sai Ganesh Nerella, Shashank Uppoor, Sang Woo Jun |
FPL | 1 |
| 2022 | BurstZ+: Eliminating The Communication Bottleneck of Scientific Computing Accelerators via Accelerated CompressionabstractWe present BurstZ+, an accelerator platform that eliminates the communication bottleneck between PCIe-attached scientific computing accelerators and their host servers, via hardware-optimized compression. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once data is larger than its on-board memory capacity, and performance becomes limited by the communication bandwidth of moving data between the host memory and accelerator. Compression has not been very useful in solving this issue due to performance and efficiency issues of compressing floating point numbers, which scientific data often consists of. BurstZ+ is an FPGA-based prototype accelerator platform which addresses the bandwidth issue via a class of novel hardware-optimized floating point compression algorithm called ZFP-V. We demonstrate that BurstZ+ can completely remove the host-side communication bottleneck for accelerators, using multiple stencil kernels with a wide range of operational intensities. Evaluated against hand-optimized implementations of kernel accelerators of the same architecture, our single-pipeline BurstZ+ prototype outperforms an accelerator without compression by almost 4×, and even an accelerator with enough memory for the entire dataset by over 2×. Furthermore, the projected performance of BurstZ+ on a future, faster FPGA scales to almost 7× that of the same accelerator without compression, whose performance is still limited by the PCIe bandwidth. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2021 | : Near-Storage Accelerator for High-Performance Log AnalyticsabstractThis paper presents, a log analytics platform with near-storage accelerators for high-performance, cost- and power-efficient unstructured log processing. offloads log analytics queries to an efficient near-storage FPGA implementation of a token querying engine, which can take advantage of the high internal bandwidth of storage devices within the available chip resource limitations. This engine is flexible enough to handle complex queries including template search based on user-defined tree-based template libraries, as well as concurrent execution of multiple queries. also uses a log-optimized version of a simple, high-throughput compression algorithm in order to further improve the effective bandwidth of backing storage. Seongyoung Kang, Jiyoung An, Jinpyo Kim, Sang Woo Jun |
MICRO | 1 |
| 2020 | FPGA-Accelerated Time Series Mining on Low-Power IoT DevicesabstractWe present a case for FPGA-accelerated edge processing for low-power Internet-of-Things (IoT) devices, using time series similarity search as a driving application. As the data collection capabilities of low-power IoT device increase, the primary constraint on their capacity is becoming the resource requirements of wirelessly transferring collected data to a central repository. This work presents a solution to this limitation by augmenting the IoT device with a inexpensive, power-efficient FPGA accelerator, which can perform fairly complex edge mining operations and drastically reduce the wireless data transfer requirements. This approach reduces the total power consumption of the device despite the added FPGA component, while also reducing the computation requirements at the central server. We use the Dynamic Time Warping (DTW) algorithm as an example workload. Using a low-cost Lattice iCE40 UltraPlus FPGA, we demonstrate that the FPGA-augmented mining algorithm can both support significantly higher data collection rate while improving the computation power efficiency of the entire deployment by an order of magnitude. Seongyoung Kang, Jinyeong Moon, Sang Woo Jun |
ASAP | 1 |
| 2020 | BurstZ: a bandwidth-efficient scientific computing accelerator platform for large-scale dataabstractWe present BurstZ, a bandwidth-efficient accelerator platform for scientific computing. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once the working set becomes larger than the on-board memory capacity, causing the performance to become bottlenecked either by the communication bandwidth between the host and the accelerator. Compression has not been very useful in solving this issue due to the difficulty of efficiently compressing floating point numbers, which scientific data often consists of. Most compression algorithms are either ineffective with floating point numbers, or has a high performance overhead. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ICS | 2 |