Jeageun Jung

dblp:280/0433 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0004-3622-6199ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 33% Hardware accelerators and domain-specific architectures · 29% Hardware reliability and fault tolerance · 19%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
DRAM
0.712023
Predicting Future-System Reliability with a Component-Level DRAM Fault Model · MICRO 2023
Hardware reliability and fault tolerance
memory reliability
0.712023
Predicting Future-System Reliability with a Component-Level DRAM Fault Model · MICRO 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN inference
0.512021
Accelerating bandwidth-bound deep learning inference with main-memory accelerators · SC 2021
High-performance computing › numerical linear algebra
GEMM
0.512021
Accelerating bandwidth-bound deep learning inference with main-memory accelerators · SC 2021
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
Accelerating bandwidth-bound deep learning inference with main-memory accelerators · SC 2021
Memory systems
processing-in-memory
0.512021
Accelerating bandwidth-bound deep learning inference with main-memory accelerators · SC 2021
High-performance computing › numerical linear algebra
matrix multiplication
0.112021
Accelerating bandwidth-bound deep learning inference with main-memory accelerators · SC 2021

Methods — techniques the papers use, named apart from their topics

empirical analysis · 0.7memory-side address generation · 0.5PIM kernels · 0.5
YearPublicationVenuePosition
2026 ECC Enabled Reliable and Performant Processing-in-Memory
Jeageun Jung, Margaret Lee, Mattan Erez
ISCA1
2023 Predicting Future-System Reliability with a Component-Level DRAM Fault Model
abstract
We introduce a new fault model for recent and future DRAM systems that uses empirical analysis to derive DRAM internal-component level fault models. This modeling level offers higher fidelity and greater predictive capability than prior models that rely on logical-address based characterization and modeling. We show how to derive the model, overcoming several challenges of using a publicly-available dataset of memory error logs. We then demonstrate the utility of our model by scaling it and analyzing the expected reliability of DDR5, HBM3, and LPDDR5 based systems. In addition to the novelty of the analysis and the model itself, we draw several insights regarding on-die ECC design and tradeoffs and the efficacy of repair/retirement mechanisms.
Jeageun Jung, Mattan Erez
MICRO1
2021 Accelerating bandwidth-bound deep learning inference with main-memory accelerators
abstract
Matrix-matrix multiplication operations (GEMMs) are important in many HPC and machine-learning applications. They are often mapped to discrete accelerators (e.g., GPUs) to improve performance. However, we find that large tall/skinny and fat/short matrices benefit little from discrete acceleration and also do not perform well on a CPU. Such matrices are prevalent in important workloads, such as deep-learning inference within large-scale datacenters. We demonstrate the large potential of accelerating these GEMMs with processing in the main CPU memory, where processing in memory units (PIMs) take advantage of otherwise untapped bandwidth without requiring data copies. We develop a novel GEMM execution flow and corresponding memory-side address-generation logic that exploits GEMM locality and enables long-running PIM kernels despite the complex address-mapping functions employed by the CPU. Our evaluation of StepStone variants at the channel, device, and within-device PIM levels demonstrate 12X better minimum latency than a CPU and 2.8X greater throughput for strict query latency constraints. End-to-end performance analysis of recent recommendation and language models shows that StepStone outperforms a fast CPU by up to 16X and also the best prior main-memory acceleration approaches by up to 2.4X.
Benjamin Y. Cho, Jeageun Jung, Mattan Erez
SC2