Eui-Cheol Lim

dblp:85/5500 · also Euicheol Lim · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-8910-533XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 9 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference System
abstract
The expansion of long-context Large Language Models (LLMs) creates significant memory system challenges. While Processing-in-Memory (PIM) is a promising accelerator, we identify that it suffers from critical inefficiencies when scaled to long contexts: severe channel underutilization, performancelimiting I/O bottlenecks, and massive memory waste from static KV cache management. In this work, we propose PIMphony, a PIM orchestrator that systematically resolves these issues with three co-designed techniques. First, Token-Centric PIM Partitioning (TCP) ensures high channel utilization regardless of batch size. Second, Dynamic PIM Command Scheduling (DCS) mitigates the I/O bottleneck by overlapping data movement and computation. Finally, a Dynamic PIM Access (DPA) controller enables dynamic memory management to eliminate static memory waste. Implemented via an MLIR-based compiler and evaluated on a cycle-accurate simulator, PIMphony significantly improves throughput for long-context LLM inference (up to 72B parameters and 1M context length). Our evaluations show performance boosts of up to 11.3× on PIM-only systems and 8.4× on xPU+PIM systems, enabling more efficient deployment of LLMs in real-world long-context applications.
Hyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee, Gyeonggeun Jung, Hyungdeok Lee, Yousub Jung, Jaehan Park, Yosub Song, Byeongsu Yang, Haerang Choi, Guhyun Kim, Jongsoon Won, Woojae Shin, Gyeongcheol Shin, Yongkee Kwon, Ilkon Kim, Eui-Cheol Lim, John Kim 0001, Jungwook Choi
HPCA20
2026 A New NB-LDPC Code Construction Method With Arbitrary Generalized Girth
abstract
!Non-binary low density parity check (NB-LDPC) codes are known to have higher correction capability than their binary counterparts, low density parity check (LDPC) codes. The latest technique for analyzing NB-LDPC codes is to measure generalized girth. The generalized girth cannot be measured in a binary LDPC code, and can be considered an indicator for measuring the performance of NB-LDPC code. In this paper, we propose a new analysis method to represent NB-LDPC code as a graph. This graph is called a permutation graph. In this paper, we present and prove a reasoning for analyzing generalized girth using the permutation graph. This is a faster and simpler method than the existing generalized girth analysis method. The proposed code construction algorithm constructs NB-LDPC codes to have an arbitrary generalized girth length. Because the permutation graph-based analysis is applied, which is more effective than existing methods, the complexity of code construction is significantly reduced. This enables more efficient code construction. The decoding simulations show that the proposed method improves the correction performance of NB-LDPC codes compared to the existing state-of-the-art.
Jaeil Lim, Jaewon Chung, Donghun Jeong, Daegeun Jee, Eui-Cheol Lim
IEEE Trans. Commun.5
2025 A New ECC Configuration Method for DRAM System Considering Metadata
abstract
In this paper, a new ECC (error correcting code) solution for DRAM (dynamic random access memory) in computing systems is proposed. Existing papers on ECC for DRAM systems do not consider storage space for metadata. The methodology proposed in this paper considers storing metadata attached to a cacheline data in DRAM. We infer the maximum number of single-chip error correction cases that a linear code can support while considering metadata storage space. This can be said to be the maximum theoretical correction probability for a single chip error. A methodology to construct a code with maximum single-chip error correction is presented. A decoding methodology for the code is proposed. The proposed ECC solution can correct not only single chip failure but also additional small bit errors. We calculate the correction capability of the proposed methodology and verified it through simulation. The encoder and decoder hardware were synthesized and compared with existing methodologies.
Jaeil Lim, Jaewon Chung, Donghun Jeong, Daegeun Jee, Eui-Cheol Lim
IEEE Trans. Computers5
2024 SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference
abstract
•Recap Accelerator-in-Memory (AiM) & AiMX •System Extensions of AiMX Card for Datacenter •AiM & AiMX for On-device AI •Design Choices for Future AiM/AiMX •Conclusion
Guhyun Kim, Jinkwon Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim 0001, Ilkon Kim, Jaehan Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Yonghoon Park, Yousub Jung, Gi-Ho Park, Eui-Cheol Lim
HCS24
2024 Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP.
Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim
MICRO10
2023 Memory-Centric Computing with SK Hynix's Domain-Specific Memory
Yongkee Kwon, Guhyun Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim 0001, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Seongju Lee, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, John Kim 0001, Eui-Cheol Lim, Junhyun Chun
HCS26
2022 QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioning
abstract
We have seen many successful deployments of deep learning accelerator designs on different platforms and technologies, e.g., FPGA, ASIC, and Processing In-Memory platforms. However, the size of the deep learning models keeps increasing, making computations a burden on the accelerators. A naive approach to resolve this issue is to design larger accelerators; however, it is not scalable due to high resource requirements, e.g., power consumption and off-chip memory sizes. A promising solution is to utilize multiple accelerators and use them as needed, similar to conventional multiprocessing. For example, for smaller networks, we may use a single accelerator, while we may use multiple accelerators with proper network partitioning for larger networks. However, partitioning DNN models into multiple parts leads to large communication overheads due to inter-layer communications. In this paper, we propose a scalable solution to accelerate DNN models on multiple devices by devising a new model partitioning technique. Our technique transforms a DNN model into layer-wise partitioned models using an autoencoder. Since the autoencoder encodes a tensor output into a smaller dimension, we can split the neural network model into multiple pieces while significantly reducing the communication overhead to pipeline them. Our evaluation results conducted on state-of-the-art deep learning models show that the proposed technique significantly improves performance and energy efficiency. Our solution increases performance and energy efficiency by up to 30.5% and 28.4% with minimal accuracy loss as compared to running the same model on pipelined multi-block accelerators without the autoencoder.
Hyukjun Kwon, Seowoo Kim, Minho Ha, Eui-Cheol Lim, Mohsen Imani, Yeseong Kim
DAC6
2022 System Architecture and Software Stack for GDDR6-AiM
abstract
This poster presents system architecture, software stack, and performance analysis for SK hynix’s very first GDDR6-based processing-in-memory (PIM) product sample, called Accelerator-in-Memory (AiM).AiM is designed for the in-memory acceleration of matrix-vector product operations, which are commonly found in machine learning applications. The strength of AiM primarily comes from the two design factors, which are 1) all-bank operation support and 2) extended DRAM command set. All-bank operations allow AiM to fully utilize the abundant internal DRAM bandwidth, which makes it an attractive solution for memory-bound applications. The extended command set allows the host to address these new operations efficiently and provides a clean separation of concerns between the AiM architecture and its software stack design.We present a dedicated FPGA-based reference platform with a software stack, which is used to validate AiM design and evaluate its system-level performance. We also demonstrate FMC-based AiM extension cards that are compatible with the off-the-shelf FPGA boards and serve as an open research platform allowing potential collaborators and academic institutes to access our hardware and software systems.
Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeongbin Kim 0001, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seong Ju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gimoon Hong, Dongyoon Ka, Kyudong Hwang, Jeongje Park, Kyeong Pil Kang, Jungyeon Kim, Junyeol Jeon, Myeongjun Lee, Minyoung Shin, Minhwan Shin, Jaekyung Cha, Changson Jung, Kijoon Chang, Chunseok Jeong, Eui-Cheol Lim, Il Park 0001, Junhyun Chun
HCS39
2022 Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems
abstract
Modern deep learning (DL) training is memory-consuming, constrained by the memory capacity of each computation component and cross-device communication bandwidth. In response to such constraints, current approaches include increasing parallelism in distributed training and optimizing inter-device communication. However, model parameter communication is becoming a key performance bottleneck in distributed DL training. To improve parameter communication performance, we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols and MPI-like collective communication for synchronization, to allow low-latency and parallel access to training data and model parameters shared among worker GPUs. To enable high bandwidth transfers between GPUs and the disaggregated memory system, we propose a decentralized parameter communication scheme to decouple and localize parameter synchronization traffic. Furthermore, we propose dynamic tensor routing and partitioning to fully utilize the non-uniform serial bus bandwidth varied across different cloud computing systems. Finally, we design a deadlock avoidance and dual synchronization to ensure high-performance parameter synchronization. Our evaluation shows that COARSE achieves up to 48.3% faster DL training compared to the state-of-the-art MPI AllReduce communication.
Zixuan Wang 0027, Joonseop Sim, Eui-Cheol Lim, Jishen Zhao
HPCA3
2022 A Lightweight and Efficient GPU for NDP Utilizing Data Access Pattern of Image Processing
abstract
As the demand for image applications with high resolution increases, the importance of the system for image processing is growing. Graphics processing units (GPUs) can increase computational capacity with massive parallelism, but are still subject to limited memory bandwidth. Near-data-processing (NDP) is expected to mitigate the performance and energy overhead caused as a result of data transfer by performing computations on the logic die of 3D-stacked memory. Although prior studies have demonstrated the advantages of NDP, a NDP solution focused on image processing has not yet been developed. This article proposes a GPU-based NDP architecture and well-matched optimization strategies considering both the characteristics of image applications and NDP constraints. First, data allocation to the processing unit is addressed to maintain the data locality and data access pattern. Second, a lightweight yet efficient NDP GPU architecture is proposed. By applying a prefetcher that leverages the pattern-aware data allocation, the number of active warps and the on-chip SRAM size of the NDP are significantly reduced. This enables the NDP constraints to be satisfied and a greater number of processing units to be integrated on a logic die. The evaluation results show that the proposed NDP GPU improves the performance by 1.85× and consumes 82.7 percent energy compared to the baseline NDP GPU.
Jungwoo Choi, Boyeal Kim, Ji-Ye Jeon, Eui-Cheol Lim, Chae-Eun Rhee
IEEE Trans. Computers5
2019 POSTER: GPU Based Near Data Processing for Image Processing with Pattern Aware Data Allocation and Prefetching
abstract
The following topics are dealt with: parallel processing; multiprocessing systems; graphics processing units; cache storage; storage management; data structures; shared memory systems; learning (artificial intelligence); program compilers; memory architecture.
Jungwoo Choi, Boyeal Kim, Ji-Ye Jeon, Eui-Cheol Lim, Chae-Eun Rhee
PACT5
2019 Exploration of a PIM Design Configuration for Energy-Efficient Task Offloading
abstract
Processing in memory (PIM) has been proposed to overcome the structural difficulties associated with conventional types of computing architecture and to realize a breakthrough for applications with high data requirements. There have been numerous attempts to utilize the PIM concept. Among them, PIM offloading takes advantage of high memory bandwidths, and various sophisticated conditions have been proposed to offload certain jobs to PIM on an instruction, task or application basis. However, schemes thus far are complicated and difficult to use. This can hinder the widespread use of PIM. Moreover, a performance improvement is not guaranteed in all cases even with these complicated schemes. This paper focuses on the energy efficiency of PIM technology. The potential-based conditions are defined to offload as many tasks as possible. This is justified in that PIM takes absolute advantage over the host in terms of power consumption. A PIM configuration favorable to the selected tasks is then devised. Simulation results show that the combination of potential-based task selection and the associated design configuration can effectively speed up the process while also reducing the energy use in most cases.
Byoung-Hak Kim, Eui-Cheol Lim, Chae-Eun Rhee
ISCAS2