EDBT 2026 Demo / reviewers in the wild / expert
Jiyoung An
dblp:240/0912
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Barad-dur: Near-Storage Accelerator for Training Large Graph Neural NetworksabstractGraph Neural Networks (GNNs) enable effective machine learning on graph-structured data, but their performance and scalability are often limited by the irregular structure and large size of real-world graphs. Conventional accelerators such as GPUs suffer a sharp performance loss when the target graph exceeds its fast random-access memory, because the overhead of partitioning graphs into memory-size chunks quickly dominates performance. We address this issue with Barad-dur, a near-storage GNN accelerator which processes the entire graph in cost-efficient solid-state storage (SSD) to remove the partitioning overhead. Barad-dur minimizes the performance impact of using relatively slow SSDs instead of DRAM, via storage-optimized graph encodings, hardware-accelerated access reordering, and the higher internal storage bandwidth available to near-storage accelerators. We demonstrate that Barad-dur can better maintain computational throughput on larger graphs compared to in-memory CPU and GPU systems. On even moderately large graphs such as the Twitter graph, Barad-dur achieves an order of magnitude higher performance compared to a highly optimized Pytorch Geometric implementation running on the NVIDIA V100 GPU, while requiring a fraction of capital cost and energy. Jiyoung An, Esmerald Aliaj, Sang Woo Jun |
PACT | 1 |
| 2023 | PreCog: Near-Storage Accelerator for Heterogeneous CNN InferenceabstractComputational Storage Devices (CSD) with near-storage acceleration is gaining popularity for data-intensive applications, by moving power-efficient hardware acceleration closer to data. However, because the power-constrained near-storage accelerator is often not powerful enough by itself to handle all computation requirements of an application, it must intelligently cooperate with other computation and acceleration units in the system. In this work, we explore how a near-storage accelerator can best fit into a larger computer system in the context of CNN inference. We demonstrate that an attractive configuration is using the near-storage accelerator to offload only the first convolution and pooling layers, where the accelerator can achieve almost an order of magnitude better performance compared to a general convolution accelerator. Targeting only the first layer allows some FPGA-specific floating-point computation optimizations such as pre-determine the range of output exponents, performing a costly floating point normalization task only once, as well as pack more input into the datapath to mitigate the performance impact of wide strides of the convolution filter. We package these optimizations into a flexible library we call Static Range Float (SRFloat), and construct a prototype system called PreCog. We evaluate PreCog implemented on the Samsung SmartSSD platform, and demonstrate over$\mathbf{3}\times$performance efficiency compared to conventional convolution accelerators on prominent CNN models without introducing a communications bottleneck. Jiyoung An, Esmerald Aliaj, Sang Woo Jun |
ASAP | 1 |
| 2021 | : Near-Storage Accelerator for High-Performance Log AnalyticsabstractThis paper presents, a log analytics platform with near-storage accelerators for high-performance, cost- and power-efficient unstructured log processing. offloads log analytics queries to an efficient near-storage FPGA implementation of a token querying engine, which can take advantage of the high internal bandwidth of storage devices within the available chip resource limitations. This engine is flexible enough to handle complex queries including template search based on user-defined tree-based template libraries, as well as concurrent execution of multiple queries. also uses a log-optimized version of a simple, high-throughput compression algorithm in order to further improve the effective bandwidth of backing storage. Seongyoung Kang, Jiyoung An, Jinpyo Kim, Sang Woo Jun |
MICRO | 2 |