Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sanghun Oh

dblp:131/3870 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0002-7669-112XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 46% Hardware accelerators and domain-specific architectures · 46% Storage systems · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › processing-in-memory
in-flash processing
0.912025
AiF: Accelerating On-Device LLM Inference Using In-Flash Processing · ISCA 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.912025
AiF: Accelerating On-Device LLM Inference Using In-Flash Processing · ISCA 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
AiF: Accelerating On-Device LLM Inference Using In-Flash Processing · ISCA 2025
Memory systems
processing-in-memory
0.912025
AiF: Accelerating On-Device LLM Inference Using In-Flash Processing · ISCA 2025
Storage systems
flash and SSD
0.312025
AiF: Accelerating On-Device LLM Inference Using In-Flash Processing · ISCA 2025

Methods — techniques the papers use, named apart from their topics

matrix-vector multiplication · 0.9error correction · 0.9
YearPublicationVenuePosition
2025 AiF: Accelerating On-Device LLM Inference Using In-Flash Processing
abstract
While large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint.
Jae Yong Lee 0004, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001
ISCA3