Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Seonah Yoo

dblp:402/3607 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 87% Hardware accelerators and domain-specific architectures · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory controller
DRAM address mapping
0.912025
FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM Inference · HPCA 2025
Memory systems
processing-in-memory
0.912025
FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM Inference · HPCA 2025
Hardware accelerators and domain-specific architectures › efficient inference
on-device LLM inference
0.312025
FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM Inference · HPCA 2025

Methods — techniques the papers use, named apart from their topics

user-level library · 0.9memory controller design · 0.9
YearPublicationVenuePosition
2025 FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM Inference
abstract
The rise of on-device inference of large language models (LLMs) is rapidly escalating the demand for memory-intensive operations on edge devices. While DRAMbased processing-in-memory (PIM) is a promising solution for overcoming the memory wall, edge devices require PIM to function both as a compute unit and a memory device due to their limited memory capacity. Such PIM-enabled memory complicates the partition and placement of a tensor into DRAM banks in a PIM-operable manner. Notably, we highlight that LLM weights need to be accessible by both PIM and system-on-chip (SoC) processors, as the same weights are used for both SoC-favorable GEMM and PIM-favorable GEMV operations. This necessitates different memory mappings for PIM and SoC processors, leading to potential re-layout costs when switching between the two. To address this challenge, we propose FACIL, a flexible DRAM address mapping solution that efficiently places tensors in DRAM for PIM operations while allowing SoC processors to access the same data using contiguous virtual addresses. FACIL consists of (i) a memory controller that assigns different DRAM address mapping to the page offset bits of each huge page and (ii) a user-level library that determines the appropriate DRAM address mapping. We demonstrate that enabling re-layout-free access of both PIM and SoC processor benefits LLM inference on various on-device LLM tasks, including short conversation and code autocompletion, reducing the time-to-first-token by $2.37 \times$ and $2.63 \times$, respectively, over the SoC-PIM baseline.
Seong Hoon Seo, Junghoon Kim 0008, Donghyun Lee 0005, Seonah Yoo, Seokwon Moon, Yeonhong Park, Jae W. Lee
HPCA4