Kyungsoo Kim 0003

dblp:08/60-3 · also KyungSoo Kim 0003 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-8927-1530ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 78% Hardware accelerators and domain-specific architectures · 17% Cloud and datacenter computing · 5%
Network and information security
1 paper
Hardware security and side channels · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › processing-in-memory
near-memory processing
1.622025
Accelerating Confidential Recommendation Model Inference With Near-Memory Processing · IEEE Trans. Dependable Secur. Comput. 2025
An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models · HPCA 2024
Memory systems
processing-in-memory
1.622025
Accelerating Confidential Recommendation Model Inference With Near-Memory Processing · IEEE Trans. Dependable Secur. Comput. 2025
An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models · HPCA 2024
Hardware security and side channels › trusted execution environments
confidential computing
0.912025
Accelerating Confidential Recommendation Model Inference With Near-Memory Processing · IEEE Trans. Dependable Secur. Comput. 2025
Hardware security and side channels
trusted execution environments
0.912025
Accelerating Confidential Recommendation Model Inference With Near-Memory Processing · IEEE Trans. Dependable Secur. Comput. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.812024
An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models · HPCA 2024
Recommender systems
DLRM inference
0.312025
Accelerating Confidential Recommendation Model Inference With Near-Memory Processing · IEEE Trans. Dependable Secur. Comput. 2025

Methods — techniques the papers use, named apart from their topics

memory encryption · 2.6Intel SGX · 2.6FPGA emulation · 2.6software stack · 0.8LPDDR5X · 0.8CXL interconnect · 0.8
YearPublicationVenuePosition
2025 Accelerating Confidential Recommendation Model Inference With Near-Memory Processing
abstract
Trusted Executing Environments (TEEs) in hardware designs protect program execution from other untrusted software programs in the processor as well as untrusted off-chip hardware components. Meanwhile, Near-Memory Processing (NMP) has shown performance and energy benefits on memory-intensive workloads. Recently, novel memory encryption schemes have been proposed to allow TEEs to leverage the benefits of NMP without requiring trust in the NMP components. In this paper, we present a system design of confidential computing with NMP that can be directly used in Intel SGX, a TEE platform available in commercial processors today. We develop the full software stack and evaluate the results on commercial processors with the emulated AxDIMM, an FPGA-based NMP platform. In our case study on personalized Deep Learning Recommendation Model (DLRM) inference, the proposed confidential computing in NMP achieves up to 1.51× latency reduction and up to 2.57× throughput improvement.
Wenjie Xiong 0001, Liu Ke 0001, Maxim Ostapenko, Yongmin Tai, Yeongon Cho, Joon-Ho Song, Jinin So, Kyungsoo Kim 0003, Yongsuk Kwon, Jin Jung, Byeongho Kim, Shinhaeng Kang, Sukhan Lee 0002, Jeonghyeon Cho, Kyomin Sohn, Xuan Zhang 0001, Hsien-Hsin S. Lee, G. Edward Suh
IEEE Trans. Dependable Secur. Comput.8
2024 An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models
abstract
Transformer-based large language models (LLMs) such as Generative Pre-trained Transformer (GPT) have become popular due to their remarkable performance across diverse applications, including text generation and translation. For LLM training and inference, the GPU has been the predominant accelerator with its pervasive software development ecosystem and powerful computing capability. However, as the size of LLMs keeps increasing for higher performance and/or more complex applications, a single GPU cannot efficiently accelerate LLM training and inference due to its limited memory capacity, which demands frequent transfers of the model parameters needed by the GPU to compute the current layer(s) from the host CPU memory/storage. A GPU appliance may provide enough aggregated memory capacity with multiple GPUs, but it suffers from frequent transfers of intermediate values among GPU devices, each accelerating specific layers of a given LLM. As the frequent transfers of these model parameters and intermediate values are performed over relatively slow device-to-device interconnects such as PCIe or NVLink, they become the key bottleneck for efficient acceleration of LLMs. Focusing on accelerating LLM inference, which is essential for many commercial services, we develop CXL-PNM, a processing near memory (PNM) platform based on the emerging interconnect technology, Compute eXpress Link (CXL). Specifically, we first devise an LPDDR5X-based CXL memory architecture with 512GB of capacity and 1.1TB/s of bandwidth, which boasts 16× larger capacity and 10× higher bandwidth than GDDR6and DDR5-based CXL memory architectures, respectively, under a module form-factor constraint. Second, we design a CXLPNM controller architecture integrated with an LLM inference accelerator, exploiting the unique capabilities of such CXL memory to overcome the disadvantages of competing technologies such as HBM-PIM and AxDIMM. Lastly, we implement a CXLPNM software stack that supports seamless and transparent use of CXL-PNM for Python-based LLM programs. Our evaluation shows that a CXL-PNM appliance with 8 CXL-PNM devices offers 23% lower latency, 31% higher throughput, and 2.8× higher energy efficiency at 30% lower hardware cost than a GPU appliance with 8 GPU devices for an LLM inference service.
Sangsoo Park, Kyungsoo Kim 0003, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim 0006, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Yeongon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, Nam Sung Kim
HPCA2
2023 Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn
HCS9