Se-Min Lim

dblp:233/7104 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-9810-0485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA Acceleration
Seongyoung Kang, Se-Min Lim, Sang Woo Jun
ISCA2
2025 Bancroft: Genomics Acceleration Beyond On-Device Memory
abstract
This paper presents Bancroft, a computational genomics acceleration platform for processing datasets far exceeding accelerator memory capacity. Bancroft overcomes the capacity limitations of accelerator memory by storing genomic data in a novel genomic compression format on the host server, and decompressing it on-demand within the accelerator after fetching it over PCIe. The key innovation of Bancroft is algorithmic optimizations for reference-based compression and decompression of popular genomic file formats. The algorithm achieves high enough compression ratios to improve the effective bandwidth of the PCIe to DRAM-levels, while facilitating highly efficient hardware implementations. We evaluate a prototype implementation of Bancroft on an affordable Alveo U50 FPGA accelerator card equipped with 8 GB of High-Bandwidth Memory (HBM). Our evaluation demonstrates that Bancroft delivers speeds exceeding on-device DDR4 memory and over 30% of HBM performance, while incurring only 30% chip space overhead. This is an order of magnitude higher performance and efficiency compared to conventional PCIe-limited architectures. Using a real-world pre-alignment filtering application, Bancroft demonstrates over $7 \times$ performance improvement over conventional accelerators on scalable datasets.
Se-Min Lim, Seongyoung Kang, Sang Woo Jun
PACT1
2025 Labidus: RISC-V Overlay with Streaming Asynchronous Custom Instructions
abstract
High development complexity is one of the most critical issues preventing more widespread use of reconfigurable hardware accelerators such as FPGAs. While soft processor overlays allow productive development with high-level software tools, they suffer from low performance. We address this issue with Labidus, a parallel RISC-V soft processor overlay that addresses this performance gap while maintaining development simplicity. Labidus automatically generates custom instructions based on static analysis of user software. These custom instructions achieve high utilization through two key innovations: asynchronous semantics and sharing across four-core tiles. Our evaluation across four scientific computing applications shows that Labidus matches or exceeds the performance of even manually optimized FPGA accelerators for gigabyte-scale tasks, at a fraction of development effort.
Gongjin Sun, Seongyoung Kang, Jane He, Se-Min Lim, Sang Woo Jun
ASAP4
2024 Durin: CPU-FPGA Heterogeneous Platform for Scalable Low-Dimensional Data Clustering
abstract
Reconfigurable hardware accelerators, known for their high performance and power efficiency, have yet to be fully leveraged for clustering low-dimensional data at realistic scales. In this work, we identify and address two hurdles for accelerating this important class of applications: the overhead of implementing indexing data structures in hardware, and the PCIe bottleneck when the data capacity spills over from accelerator memory to host storage. We overcome these hurdles for the first time with Durin, a CPU-FPGA heterogeneous system with a hardware-software codesigned index structure, which minimizes the hardware resource overhead of high-performance neighbor search. It also minimizes PCIe overhead by facilitating asynchronous acceleration of distance calculation, and block floating-point compression. We show that a desktop computer with Durin implemented on a mid-range Alveo U50 FPGA can outperform a 32-thread Xeon server by almost 20× , with an order of magnitude power efficiency improvements. Furthermore, Durin outperforms even the best-case projection of a conventional standalone accelerator design, which implements the entirety of the clustering algorithm in the FPGA and its High Bandwidth Memory (HBM), by 2×.
Se-Min Lim, Esmerald Aliaj, Sang Woo Jun
IEEE Big Data1
2024 FlexForge: Efficient Reconfigurable Cloud Acceleration via Peripheral Resource Disaggregation
abstract
Reconfigurable hardware acceleration in the cloud using Field-Programmable Gate Arrays (FPGAs) is an increasingly popular solution for scaling performance and cost-effectiveness. For efficient utilization of FPGA resources, cloud platforms typically support elastic FPGA resource allocation. However, FPGAs are usually allocated in a homogeneous unit consisting of logic, memory, and PCIe bandwidth. Because user kernels have a wide and varying combination of resource requirements, this can result in high internal fragmentation and underutilization of each resource. To address this issue, we present FlexForge, a platform facilitating high-performance disaggregation of peripheral resources over a network of potentially untrusted FPGAs, aided by a secondary inter-FPGA network. Evaluated on a mix of representative accelerator applications deployed on a prototype cluster, FlexForge improves the overall performance of the cloud by up to 70 % and 20 % on average across all possible combinations without significant additional hardware resource requirements.
Se-Min Lim, Sang Woo Jun
DATE1
2024 Morbius: Platform-Adaptive Hardware Accelerator for Scalable Sequence Motif Discovery
abstract
Motif finding is one of the fundamental tools of computational biology. Unfortunately, the benefits of high-performance, low-power acceleration have not been available to this application due to the limited parallelism available within efficient classes of algorithms, such as probabilistic Gibbs sampling. In this work, we present Morbius, which demonstrates the benefits of reconfigurable application-specific hardware acceleration using FPGAs as a solution to the execution time and cost overhead of motif finding. Morbius employs a novel, hardware-optimized data structure called base pair matrix to minimize off-chip data movement, and implements a small number of deep hardware pipelines to achieve high sequential performance. Furthermore, we develop performance and chip space prediction models based on microarchitectural parameters of the accelerator, to facilitate optimal performance on a wide range of accelerators potential users may already have. We compare Morbius on a wide range of FPGA platforms spanning low-profile $250 M.2 FPGA cards to Amazon F1, and demonstrate up to orders of magnitude performance improvements compared to costly server machines, and even higher cost and power efficiency.
Se-Min Lim, Esmerald Aliaj, Sang Woo Jun
e-Science1