Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wonhyuk Yang

dblp:372/2525 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0003-1918-8445ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 47% GPUs and heterogeneous computing · 19% Performance modeling and evaluation · 17%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation
0.912025
PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.912025
PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025
Memory systems
cache management
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Memory systems › cache
DRAM cache
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
GPUs and heterogeneous computing › GPU memory management
GPU memory oversubscription
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Memory systems › processing-in-memory
near-data processing
0.812024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Memory systems
processing-in-memory
0.812024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Memory systems › non-volatile memory
storage class memory
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Performance modeling and evaluation › simulation › processor simulation
instruction set simulation
0.312025
PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025
Performance modeling and evaluation
simulation
0.312025
PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025
Interconnection networks and networks-on-chip › high-speed interconnect
CXL interconnect
0.212024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Energy-efficient computing › power management
memory power management
0.212024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024

Methods — techniques the papers use, named apart from their topics

spike · 0.9gem5 · 0.9RISC-V ISA · 0.9MLIR · 0.9LLVM · 0.9simulation · 0.8low-latency offloading · 0.8
YearPublicationVenuePosition
2026 A Programming Model for Efficient Inter-Kernel Control-Flow on Memory-Mapped Near-Data Processing Architecture (WIP)
abstract
As the memory wall problem worsens, Near-Data Processing (NDP) has emerged to reduce data movement by computing close to data. Recently proposed memory-mapped NDP (M2NDP) enables general-purpose NDP with low hardware overhead by extending RISC-V ISA and maximizing data parallelism with lightweight µthreads. However, a high-level programming model for this architecture—particularly one that naturally expresses control flow across kernels—has not yet been established.
Seungheon Lee, Wonhyuk Yang, Seonyeong Heo, Gwangsun Kim
LCTES2
2025 PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework
abstract
Deep Neural Networks (DNNs) have continuously increasing demands for the performance and efficiency of Neural Processing Units (NPUs).While analytical models enable rapid exploration of high-level aspects (e.g., tiling), later stages of NPU design require a cycle-accurate simulator that supports various scenarios.However, existing NPU simulators are limited in several aspects, including support for high-speed, multi-core, multi-model tenancy, generic ISA (with vector operations), compiler, data-dependent timing model, and enabling both inference and training.To address these challenges, we propose PyTorchSim, 1 a novel NPU simulation framework integrated with PyTorch 2. PyTorchSim models NPUs with a custom RISC-V-based ISA extended to support various acceleration units (e.g., systolic array).Our custom backend for PyTorch 2 compiles a given DNN using this ISA through lowering passes with MLIR and LLVM.Then, our extended Gem5 and Spike simulators execute the machine code to accurately model the DNN's timing and functional aspects on the NPU.However, as such a conventional Instruction-Level Simulation (ILS) inevitably runs slowly, we propose Tile-Level Simulation (TLS) to improve speed without sacrificing accuracy.It uses tile-granularity operation latencies from offline ILS runs for high speed while still modeling DRAM and interconnect with cycle-accurate simulators.Furthermore, TLS can also be employed for sparse tensor operations using auxiliary * These authors contributed equally to this work.
Wonhyuk Yang, Yunseon Shin, Okkyun Woo, Geonwoo Park, Hyungkyu Ham, Jeehoon Kang, Jongse Park, Gwangsun Kim
MICRO1
2024 Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory
abstract
We propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that mandate memory oversubscription, resulting in substantial speedups. However, the DRAM cache needs to be carefully designed to address the latency and bandwidth limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can easily thrash the DRAM cache and degrade performance, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by G PU s with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probe traffic and increase effective DRAM BW with minimal cost overhead, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache implementation with Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power consumption and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, the HMS improves performance by up to 12.5× (2.9× overall) and reduces energy by up to 89.3% (48.1 % overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91–93 % and 57–75 %, respectively.
Jeongmin Hong 0001, Sungjun Cho, Geonwoo Park, Wonhyuk Yang, Young-Ho Gong, Gwangsun Kim
HPCA4
2024 Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP.
Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim
MICRO6