VLDB 2026 Research / reviewers in the wild / expert
Wonhyuk Yang
dblp:372/2525
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0003-1918-8445ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 47% GPUs and heterogeneous computing · 19% Performance modeling and evaluation · 17% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation |
0.9 | 1 | 2025 | PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.9 | 1 | 2025 | PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025 |
Memory systems
cache management |
0.8 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
Memory systems › cache
DRAM cache |
0.8 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.8 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
GPUs and heterogeneous computing › GPU memory management
GPU memory oversubscription |
0.8 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
Memory systems › processing-in-memory
near-data processing |
0.8 | 1 | 2024 | Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024 |
Memory systems
processing-in-memory |
0.8 | 1 | 2024 | Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024 |
Memory systems › non-volatile memory
storage class memory |
0.8 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
Performance modeling and evaluation › simulation › processor simulation
instruction set simulation |
0.3 | 1 | 2025 | PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025 |
Performance modeling and evaluation
simulation |
0.3 | 1 | 2025 | PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation Framework · MICRO 2025 |
Interconnection networks and networks-on-chip › high-speed interconnect
CXL interconnect |
0.2 | 1 | 2024 | Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024 |
Energy-efficient computing › power management
memory power management |
0.2 | 1 | 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024 |
Methods — techniques the papers use, named apart from their topics
spike · 0.9gem5 · 0.9RISC-V ISA · 0.9MLIR · 0.9LLVM · 0.9simulation · 0.8low-latency offloading · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Programming Model for Efficient Inter-Kernel Control-Flow on Memory-Mapped Near-Data Processing Architecture (WIP)abstractAs the memory wall problem worsens, Near-Data Processing (NDP) has emerged to reduce data movement by computing close to data. Recently proposed memory-mapped NDP (M2NDP) enables general-purpose NDP with low hardware overhead by extending RISC-V ISA and maximizing data parallelism with lightweight µthreads. However, a high-level programming model for this architecture—particularly one that naturally expresses control flow across kernels—has not yet been established. Seungheon Lee, Wonhyuk Yang, Seonyeong Heo, Gwangsun Kim |
LCTES | 2 |
| 2025 | PyTorchSim: A Comprehensive, Fast, and Accurate NPU Simulation FrameworkabstractDeep Neural Networks (DNNs) have continuously increasing demands for the performance and efficiency of Neural Processing Units (NPUs).While analytical models enable rapid exploration of high-level aspects (e.g., tiling), later stages of NPU design require a cycle-accurate simulator that supports various scenarios.However, existing NPU simulators are limited in several aspects, including support for high-speed, multi-core, multi-model tenancy, generic ISA (with vector operations), compiler, data-dependent timing model, and enabling both inference and training.To address these challenges, we propose PyTorchSim, 1 a novel NPU simulation framework integrated with PyTorch 2. PyTorchSim models NPUs with a custom RISC-V-based ISA extended to support various acceleration units (e.g., systolic array).Our custom backend for PyTorch 2 compiles a given DNN using this ISA through lowering passes with MLIR and LLVM.Then, our extended Gem5 and Spike simulators execute the machine code to accurately model the DNN's timing and functional aspects on the NPU.However, as such a conventional Instruction-Level Simulation (ILS) inevitably runs slowly, we propose Tile-Level Simulation (TLS) to improve speed without sacrificing accuracy.It uses tile-granularity operation latencies from offline ILS runs for high speed while still modeling DRAM and interconnect with cycle-accurate simulators.Furthermore, TLS can also be employed for sparse tensor operations using auxiliary * These authors contributed equally to this work. Wonhyuk Yang, Yunseon Shin, Okkyun Woo, Geonwoo Park, Hyungkyu Ham, Jeehoon Kang, Jongse Park, Gwangsun Kim |
MICRO | 1 |
| 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryabstractWe propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that mandate memory oversubscription, resulting in substantial speedups. However, the DRAM cache needs to be carefully designed to address the latency and bandwidth limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can easily thrash the DRAM cache and degrade performance, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by G PU s with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probe traffic and increase effective DRAM BW with minimal cost overhead, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache implementation with Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power consumption and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, the HMS improves performance by up to 12.5× (2.9× overall) and reduces energy by up to 89.3% (48.1 % overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91–93 % and 57–75 %, respectively. Jeongmin Hong 0001, Sungjun Cho, Geonwoo Park, Wonhyuk Yang, Young-Ho Gong, Gwangsun Kim |
HPCA | 4 |
| 2024 | Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersabstractEmerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP. Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim |
MICRO | 6 |