EDBT 2026 Demo / reviewers in the wild / expert
Gongjin Sun
dblp:248/4972
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0002-5420-1361ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Labidus: RISC-V Overlay with Streaming Asynchronous Custom InstructionsabstractHigh development complexity is one of the most critical issues preventing more widespread use of reconfigurable hardware accelerators such as FPGAs. While soft processor overlays allow productive development with high-level software tools, they suffer from low performance. We address this issue with Labidus, a parallel RISC-V soft processor overlay that addresses this performance gap while maintaining development simplicity. Labidus automatically generates custom instructions based on static analysis of user software. These custom instructions achieve high utilization through two key innovations: asynchronous semantics and sharing across four-core tiles. Our evaluation across four scientific computing applications shows that Labidus matches or exceeds the performance of even manually optimized FPGA accelerators for gigabyte-scale tasks, at a fraction of development effort. Gongjin Sun, Seongyoung Kang, Jane He, Se-Min Lim, Sang Woo Jun |
ASAP | 1 |
| 2025 | Revisiting Memory Hierarchies with CMM-H: Use Device-side Caching to Integrate DRAM and SSD for a Hybrid CXL MemoryabstractEmerging data-intensive applications increasingly demand large-scale, cost-effective, and high-performance memory solutions. Samsung's CXL Memory Module-Hybrid (CMM-H) uniquely integrates DDR DRAM and NAND flash storage within a single CXL-attached memory device for higher capacity while keeping still high performance. This paper provides a comprehensive exploration of the CMM-H module, detailing its architectural design, operational workflow, and caching mechanisms. We evaluate the performance of CMM-H through extensive experiments, highlighting its benefits and limitations. Our study contributes insights into hybrid CXL memory architectures and provides valuable guidance for future CXL memory design improvements. Mohammadreza Soltaniyeh, Gongjin Sun, Xuebin Yao, Amir Beygi, Ramdas Kachare, Dongwan Zhao, Hingkwan Huen, Senthil Murugesapandian, Caroline Kahn |
HotStorage | 2 |
| 2024 | ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture ModelabstractCompute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coherent memory space between host and CXL storage devices such as such as Solid-state drive (SSD) for memory expansion, with a corresponding DRAM implemented as the device cache. However, this introduces challenges such as substantial cache miss penalties, sub-optimal caching due to data access granularity mismatch between the DRAM "cache" and SSD "memory", and inefficient hardware cache management. To address these issues, we propose a novel solution, named ICGMM, which optimizes caching and eviction directly on hardware, employing a Gaussian Mixture Model (GMM)-based approach. We prototype our solution on an FPGA board, which demonstrates a noteworthy improvement compared to the classic Least Recently Used (LRU) cache strategy. We observe a decrease in the cache miss rate ranging from 0.32% to 6.14%, leading to a substantial 16.23% to 39.14% reduction in the average SSD access latency. Furthermore, when compared to the state-of-the-art Long Short-Term Memory (LSTM)-based cache policies, our GMM algorithm on FPGA showcases an impressive latency reduction of over 10,000 times. Remarkably, this is achieved while demanding much fewer hardware resources. Hanqiu Chen, Yitu Wang, Vitorio Cargnini, Mohammadreza Soltaniyeh, Gongjin Sun, Pradeep Subedi, Yiran Chen 0001, Cong Hao |
DAC | 6 |
| 2022 | BurstZ+: Eliminating The Communication Bottleneck of Scientific Computing Accelerators via Accelerated CompressionabstractWe present BurstZ+, an accelerator platform that eliminates the communication bottleneck between PCIe-attached scientific computing accelerators and their host servers, via hardware-optimized compression. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once data is larger than its on-board memory capacity, and performance becomes limited by the communication bandwidth of moving data between the host memory and accelerator. Compression has not been very useful in solving this issue due to performance and efficiency issues of compressing floating point numbers, which scientific data often consists of. BurstZ+ is an FPGA-based prototype accelerator platform which addresses the bandwidth issue via a class of novel hardware-optimized floating point compression algorithm called ZFP-V. We demonstrate that BurstZ+ can completely remove the host-side communication bottleneck for accelerators, using multiple stencil kernels with a wide range of operational intensities. Evaluated against hand-optimized implementations of kernel accelerators of the same architecture, our single-pipeline BurstZ+ prototype outperforms an accelerator without compression by almost 4×, and even an accelerator with enough memory for the entire dataset by over 2×. Furthermore, the projected performance of BurstZ+ on a future, faster FPGA scales to almost 7× that of the same accelerator without compression, whose performance is still limited by the PCIe bandwidth. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2020 | BurstZ: a bandwidth-efficient scientific computing accelerator platform for large-scale dataabstractWe present BurstZ, a bandwidth-efficient accelerator platform for scientific computing. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once the working set becomes larger than the on-board memory capacity, causing the performance to become bottlenecked either by the communication bandwidth between the host and the accelerator. Compression has not been very useful in solving this issue due to the difficulty of efficiently compressing floating point numbers, which scientific data often consists of. Most compression algorithms are either ineffective with floating point numbers, or has a high performance overhead. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ICS | 1 |
| 2019 | Combining Prefetch Control and Cache Partitioning to Improve Multicore PerformanceabstractModern commercial multi-core processors are equipped with multiple hardware prefetchers on each core. The prefetchers can significantly improve application performance. However, shared resources, such as last-level cache (LLC) and off-chip memory bandwidth and controller, can lead to prefetch interference. Multiple techniques have been proposed to reduce such interference and improve the performance isolation across cores, such as coordinated control among prefetchers and cache partitioning (CP). Each of them has its advantages and disadvantages. This paper proposes combining these two techniques in a coordinated way. Prefetchers and LLC are treated as separate resources and a multi-resource management mechanism is proposed to control prefetching and cache partitioning. This control mechanism is implemented as a Linux kernel module and can be applied to a wide variety of prefetch architectures. An implementation on Intel Xeon E5 v4 processor shows that combining LLC partitioning and prefetch throttling provides a significant improvement in performance and fairness. Gongjin Sun, Junjie Shen 0001, Alexander V. Veidenbaum |
IPDPS | 1 |