EDBT 2026 Demo / reviewers in the wild / expert
Jeonghyeon Cho
dblp:75/8787
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Confidential Recommendation Model Inference With Near-Memory ProcessingabstractTrusted Executing Environments (TEEs) in hardware designs protect program execution from other untrusted software programs in the processor as well as untrusted off-chip hardware components. Meanwhile, Near-Memory Processing (NMP) has shown performance and energy benefits on memory-intensive workloads. Recently, novel memory encryption schemes have been proposed to allow TEEs to leverage the benefits of NMP without requiring trust in the NMP components. In this paper, we present a system design of confidential computing with NMP that can be directly used in Intel SGX, a TEE platform available in commercial processors today. We develop the full software stack and evaluate the results on commercial processors with the emulated AxDIMM, an FPGA-based NMP platform. In our case study on personalized Deep Learning Recommendation Model (DLRM) inference, the proposed confidential computing in NMP achieves up to 1.51× latency reduction and up to 2.57× throughput improvement. Wenjie Xiong 0001, Liu Ke 0001, Maxim Ostapenko, Yongmin Tai, Yeongon Cho, Joon-Ho Song, Jinin So, Kyungsoo Kim 0003, Yongsuk Kwon, Jin Jung, Byeongho Kim, Shinhaeng Kang, Sukhan Lee 0002, Jeonghyeon Cho, Kyomin Sohn, Xuan Zhang 0001, Hsien-Hsin S. Lee, G. Edward Suh |
IEEE Trans. Dependable Secur. Comput. | 15 |
| 2024 | An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language ModelsabstractTransformer-based large language models (LLMs) such as Generative Pre-trained Transformer (GPT) have become popular due to their remarkable performance across diverse applications, including text generation and translation. For LLM training and inference, the GPU has been the predominant accelerator with its pervasive software development ecosystem and powerful computing capability. However, as the size of LLMs keeps increasing for higher performance and/or more complex applications, a single GPU cannot efficiently accelerate LLM training and inference due to its limited memory capacity, which demands frequent transfers of the model parameters needed by the GPU to compute the current layer(s) from the host CPU memory/storage. A GPU appliance may provide enough aggregated memory capacity with multiple GPUs, but it suffers from frequent transfers of intermediate values among GPU devices, each accelerating specific layers of a given LLM. As the frequent transfers of these model parameters and intermediate values are performed over relatively slow device-to-device interconnects such as PCIe or NVLink, they become the key bottleneck for efficient acceleration of LLMs. Focusing on accelerating LLM inference, which is essential for many commercial services, we develop CXL-PNM, a processing near memory (PNM) platform based on the emerging interconnect technology, Compute eXpress Link (CXL). Specifically, we first devise an LPDDR5X-based CXL memory architecture with 512GB of capacity and 1.1TB/s of bandwidth, which boasts 16× larger capacity and 10× higher bandwidth than GDDR6and DDR5-based CXL memory architectures, respectively, under a module form-factor constraint. Second, we design a CXLPNM controller architecture integrated with an LLM inference accelerator, exploiting the unique capabilities of such CXL memory to overcome the disadvantages of competing technologies such as HBM-PIM and AxDIMM. Lastly, we implement a CXLPNM software stack that supports seamless and transparent use of CXL-PNM for Python-based LLM programs. Our evaluation shows that a CXL-PNM appliance with 8 CXL-PNM devices offers 23% lower latency, 31% higher throughput, and 2.8× higher energy efficiency at 30% lower hardware cost than a GPU appliance with 8 GPU devices for an LLM inference service. Sangsoo Park, Kyungsoo Kim 0003, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim 0006, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Yeongon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, Nam Sung Kim |
HPCA | 15 |
| 2023 | Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn |
HCS | 19 |
| 2022 | Improving In-Memory Database Operations with Acceleration DIMM (AxDIMM)abstractThe significant overhead needed to transfer the data between CPUs and memory devices is one of the hottest issues in many areas of computing, such as database management systems. Disaggregated computing on the memory devices is being highlighted as one promising approach. In this work, we introduce a new near-memory acceleration scheme for in-memory database operations, called Acceleration DIMM (AxDIMM). It behaves like a normal DIMM through the standard DIMM-compatible interface, but has embedded computing units for data-intensive operations. With the minimized data transfer overhead, it reduces CPU resource consumption, relieves the memory bandwidth bottleneck, and boosts energy efficiency. We implement scan operations, one of the most data-intensive database operations, within AxDIMM and compare its performance with SIMD (Single Instruction Multiple Data) implementation on CPU. Our investigation shows that the acceleration achieves 6.8x more throughput than the SIMD implementation. Donghun Lee 0001, Jinin So, Minseon Ahn, Jong-Geon Lee, Jeonghyeon Cho, Oliver Rebholz, Vishnu Charan Thummala, Ravi Shankar JV, Sachin Suresh Upadhya, Mohammed Ibrahim Khan, Jin Hyun Kim |
DaMoN | 6 |
| 2021 | Aquabolt-XL: Samsung HBM2-PIM with in-memory processing for ML accelerators and beyondabstractUsing PIM to overcome memory bottleneck • Although various bandwidth increase methods have been proposed, it is physically impossible to achieve a breakthrough increase. - Limited by # of PCB wires, # of CPU ball, and thermal constraints • PIM has been proposed to improve performance of bandwidth-intensive workloads and improve energy efficiency by reducing computing-memory data movement. Jin Hyun Kim, Shinhaeng Kang, Sukhan Lee 0002, Woongjae Song, Yuhwan Ro, Seungwon Lee 0006, David Wang 0003, Hyunsung Shin, BengSeng Phuah, Jihyun Choi, Jinin So, Yeongon Cho, Joon-Ho Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Young-Soo Sohn, Kwang-Il Park, Nam Sung Kim |
HCS | 16 |
| 2021 | Industry's First 7.2 Gbps 512GB DDR5 ModuleabstractSpurred by the increasing market needs for big data and cloud services, global server suppliers and hyper- scalers are looking to adopt high-speed and large-capacity memory modules. To fulfill this trend, the brand- new low-voltage operable DDR5 (double data rate 5th generation) memory can be an appropriate solution, with the highest speed of 7.2 Gbps and the largest capacity of 512 GB. However, some critical obstacles, such as increased capacity and high-speed I/O requirements, unstable power noise occurrences, high power consumption, and increase in operating temperature, must be overcome. This poster will cover various technical pathfinding solutions for world's first DDR5 512 GB module with an advanced DRAM process and I/O schemes, package technology, and module architecture regarding improvements in the following four aspects: performance, speed, capacity, and power. This will unveil the industry's first high-performance and large-capacity memory product with 8-stacked DDR5 DRAMs. Samsung believes that this product will pave the way for achieving both higher bandwidth and lower power consumption to inaugurate the era of terabyte DRAM modules for next-gen servers. Sung Joo Park, Jonghoon J. Kim, Kun Joo, Kyoungsun Kim, Young-Tae Kim, Woo-Jin Na, IkJoon Choi, Hye-Seung Yu, Wonyoung Kim, Ju-Yeon Jung, Young-Uk Chang, Gong-Heum Han, Hangi-Jung, Sunwon Kang, Jeonghyeon Cho, Hoyoung Song, Tae-Young Oh, Young-Soo Sohn, Sangjoon Hwang 0001 |
HCS | 18 |