Jaehong Cho

dblp:371/8693 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0008-8973-5164ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 61% Memory systems · 30% GPUs and heterogeneous computing · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024
Memory systems
processing-in-memory
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024

Methods — techniques the papers use, named apart from their topics

GEMV · 0.8GEMM · 0.8
YearPublicationVenuePosition
2026 LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
Jaehong Cho, Hyunmin Choi, Guseul Heo, Jongse Park
ISPASS1
2024 NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
abstract
Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head attention, and (3) feed-forward networks. In batched processing, QKV generation and feed-forward networks involve compute-intensive matrix-matrix multiplications (GEMM), while multi-head attention requires bandwidth-heavy matrix-vector multiplications (GEMV). Machine learning accelerators like TPUs or NPUs are proficient in handling GEMM but are less efficient for GEMV computations. Conversely, Processing-in-Memory (PIM) technology is tailored for efficient GEMV computation, while it lacks the computational power to handle GEMM effectively.
Guseul Heo, Jaehong Cho, Hyunmin Choi, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan 0001, Jongse Park
ASPLOS (3)3