EDBT 2026 Demo / reviewers in the wild / expert
Kyomin Sohn
dblp:84/2909
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-8094-9843ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AME-PIM: Can Memory be Your Next Tensor Accelerator?abstractHigh Bandwidth Memory with Processing-in-Memory (HBM-PIM) offers an opportunity to reduce data movement by executing computation directly inside memory, but current commercial platforms expose limited instruction sets and require specialized software stacks. In this work, we investigate whether HBM-PIM can serve as a backend for ISA-level matrix acceleration, using the RISC-V Attached Matrix Extension (AME) as a semantic reference. We propose a PEP-based execution model that maps AME element-wise and matrix instructions to HBM-PIM micro-kernels and data instructions in memory operations. Differently from SoA HBM-PIM, we introduce a reduction-free outer-product dataflow that enables accumulation entirely within memory despite the lack of native reduction support. Our approach supports end-to-end execution of element-wise operations, GEMV, and GEMM in PIM mode, minimizing host involvement and off-chip transfers. An experimental evaluation on Samsung Aquabolt-XL shows that AME matrix tile multiplication achieves up to 14.9 GFLOP/s (59.4 FLOP/cycle) on a single HBM pseudo-channel. Emanuele Venieri, Simone Manoni, Alberto Florian, Jaehyun Park 0006, Kyomin Sohn, Andrea Bartolini |
CF | 5 |
| 2025 | MAPLE: Flexible-Precision Processing-In-Memory Architecture for Efficient On-Device ML
Jaewon Park, Quang Anh Hoang, Jonathan Ta, Shinhaeng Kang, Kyomin Sohn, Sang Woo Jun |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform QuantizationabstractNon-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference. Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo |
ISCAS | 5 |
| 2025 | Accelerating Confidential Recommendation Model Inference With Near-Memory ProcessingabstractTrusted Executing Environments (TEEs) in hardware designs protect program execution from other untrusted software programs in the processor as well as untrusted off-chip hardware components. Meanwhile, Near-Memory Processing (NMP) has shown performance and energy benefits on memory-intensive workloads. Recently, novel memory encryption schemes have been proposed to allow TEEs to leverage the benefits of NMP without requiring trust in the NMP components. In this paper, we present a system design of confidential computing with NMP that can be directly used in Intel SGX, a TEE platform available in commercial processors today. We develop the full software stack and evaluate the results on commercial processors with the emulated AxDIMM, an FPGA-based NMP platform. In our case study on personalized Deep Learning Recommendation Model (DLRM) inference, the proposed confidential computing in NMP achieves up to 1.51× latency reduction and up to 2.57× throughput improvement. Wenjie Xiong 0001, Liu Ke 0001, Maxim Ostapenko, Yongmin Tai, Yeongon Cho, Joon-Ho Song, Jinin So, Kyungsoo Kim 0003, Yongsuk Kwon, Jin Jung, Byeongho Kim, Shinhaeng Kang, Sukhan Lee 0002, Jeonghyeon Cho, Kyomin Sohn, Xuan Zhang 0001, Hsien-Hsin S. Lee, G. Edward Suh |
IEEE Trans. Dependable Secur. Comput. | 16 |
| 2024 | Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingabstractLarge language models (LLMs) have emerged due to their capability to generate high-quality content across diverse contexts. To reduce their explosively increasing demands for computing resources, a mixture of experts (MoE) has emerged. The MoE layer enables exploiting a huge number of parameters with less computation. Applying state-of-the-art continuous batching increases throughput; however, it leads to frequent DRAM access in the MoE and attention layers. We observe that conventional computing devices have limitations when processing the MoE and attention layers, which dominate the total execution time and exhibit low arithmetic intensity (Op/B). Processing MoE layers only with devices targeting low-Op/B such as processing-in-memory (PIM) architectures is challenging due to the fluctuating Op/B in the MoE layer caused by continuous batching, To address these challenges, we propose Duplex, which comprises xPU tailored for high-Op/B and Logic-PIM to effectively perform low-Op/B operation within a single device. Duplex selects the most suitable processor based on the Op/B of each layer within LLMs. As the Op/B of the MoE layer is at least 1 and that of the attention layer has a value of 4–8 for grouped query attention, prior PIM architectures are not efficient, which place processing units inside DRAM dies and only target extremely low-Op/B (under one) operations. Based on recent trends, Logic-Pimadds more through-silicon vias (TSVs) to enable high-bandwidth communication between the DRAM die and the logic die and place powerful processing units on the logic die, which is best suited for handling low-Op/B operations ranging from few to a few dozens. To maximally utilize the xPU and Logic-Pim,we propose expert and attention co-processing. By exploiting proper processing units for MoE and attention layers, Duplex shows up to 2.67 × higher throughput and consumes 42.0% less energy compared to GPU systems for LLM inference. Sungmin Yun 0001, Kwanhee Kyung, Juhwan Cho, Jaewan Choi, Jongmin Kim 0007, Byeongho Kim, Sukhan Lee 0002, Kyomin Sohn, Jung Ho Ahn |
MICRO | 8 |
| 2023 | Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn |
HCS | 21 |
| 2022 | An FPGA-based RNN-T Inference Accelerator with PIM-HBMabstractIn this paper, we implemented a world-first RNN-T inference accelerator using FPGA with PIM-HBM that can multiply the internal bandwidth of the memory. The accelerator offloads matrix-vector multiplication (GEMV) operations of LSTM layers in RNN-T into PIM-HBM, and PIM-HBM reduces the execution time of GEMV significantly by exploiting HBM internal bandwidth. To ensure that the memory commands are issued in a pre-defined order, which is one of the most important constraints in exploiting PIM-HBM, we implement a direct memory access (DMA) module and change configuration of the on-chip memory controller by utilizing the flexibility and reconfigurability of the FPGA. In addition, we design the other hardware modules for acceleration such as non-linear functions (i.e., sigmoid and hyperbolic tangent), element-wise operation, and ReLU module, to operate these compute-bound RNN-T operations on FPGA. For this, we prepare FP16 quantized weight and MLPerf input datasets, and modify the PCIe device driver and C++ based control codes. On our evaluation, our accelerator with PIM-HBM reduces the execution time of RNN-T by 2.5 × on average with 11.09% reduced LUT size and improves energy efficiency up to 2.6 × compared to the baseline. Shinhaeng Kang, Sukhan Lee 0002, Byeongho Kim, Hweesoo Kim, Kyomin Sohn, Nam Sung Kim, Eojin Lee |
FPGA | 5 |
| 2021 | Aquabolt-XL: Samsung HBM2-PIM with in-memory processing for ML accelerators and beyondabstractUsing PIM to overcome memory bottleneck • Although various bandwidth increase methods have been proposed, it is physically impossible to achieve a breakthrough increase. - Limited by # of PCB wires, # of CPU ball, and thermal constraints • PIM has been proposed to improve performance of bandwidth-intensive workloads and improve energy efficiency by reducing computing-memory data movement. Jin Hyun Kim, Shinhaeng Kang, Sukhan Lee 0002, Woongjae Song, Yuhwan Ro, Seungwon Lee 0006, David Wang 0003, Hyunsung Shin, BengSeng Phuah, Jihyun Choi, Jinin So, Yeongon Cho, Joon-Ho Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Young-Soo Sohn, Kwang-Il Park, Nam Sung Kim |
HCS | 17 |
| 2021 | Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology : Industrial ProductabstractEmerging applications such as deep neural network demand high off-chip memory bandwidth. However, under stringent physical constraints of chip packages and system boards, it becomes very expensive to further increase the bandwidth of off-chip memory. Besides, transferring data across the memory hierarchy constitutes a large fraction of total energy consumption of systems, and the fraction has steadily increased with the stagnant technology scaling and poor data reuse characteristics of such emerging applications. To cost-effectively increase the bandwidth and energy efficiency, researchers began to reconsider the past processing-in-memory (PIM) architectures and advance them further, especially exploiting recent integration technologies such as 2.5D/3D stacking. Albeit the recent advances, no major memory manufacturer has developed even a proof-of-concept silicon yet, not to mention a product. This is because the past PIM architectures often require changes in host processors and/or application code which memory manufacturers cannot easily govern. In this paper, elegantly tackling the aforementioned challenges, we propose an innovative yet practical PIM architecture. To demonstrate its practicality and effectiveness at the system level, we implement it with a 20nm DRAM technology, integrate it with an unmodified commercial processor, develop the necessary software stack, and run existing applications without changing their source code. Our evaluation at the system level shows that our PIM improves the performance of memory-bound neural network kernels and applications by 11.2× and 3.5×, respectively. Atop the performance improvement, PIM also reduces the energy per bit transfer by 3.5×, and the overall energy efficiency of the system running the applications by 3.2×. Sukhan Lee 0002, Shinhaeng Kang, Jaehoon Lee 0005, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee 0006, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, Seongil O, Anand Iyer, David Wang 0003, Kyomin Sohn, Nam Sung Kim |
ISCA | 15 |