Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jeongmin Hong 0001

dblp:154/3027-1 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0002-5492-5346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 66% GPUs and heterogeneous computing · 26% Energy-efficient computing · 4%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Memory systems › cache
DRAM cache
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
GPUs and heterogeneous computing › GPU memory management
GPU memory oversubscription
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Memory systems › processing-in-memory
near-data processing
0.812024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Memory systems
processing-in-memory
0.812024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Memory systems › non-volatile memory
storage class memory
0.812024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024
Interconnection networks and networks-on-chip › high-speed interconnect
CXL interconnect
0.212024
Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders · MICRO 2024
Energy-efficient computing › power management
memory power management
0.212024
Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory · HPCA 2024

Methods — techniques the papers use, named apart from their topics

simulation · 0.8low-latency offloading · 0.8
YearPublicationVenuePosition
2024 Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory
abstract
We propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that mandate memory oversubscription, resulting in substantial speedups. However, the DRAM cache needs to be carefully designed to address the latency and bandwidth limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can easily thrash the DRAM cache and degrade performance, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by G PU s with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probe traffic and increase effective DRAM BW with minimal cost overhead, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache implementation with Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power consumption and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, the HMS improves performance by up to 12.5× (2.9× overall) and reduces energy by up to 89.3% (48.1 % overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91–93 % and 57–75 %, respectively.
Jeongmin Hong 0001, Sungjun Cho, Geonwoo Park, Wonhyuk Yang, Young-Ho Gong, Gwangsun Kim
HPCA1
2024 Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP.
Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim
MICRO2