Seunghyun Jin 0002

dblp:43/5656-2 · also Seugn Hyun Jin 0002 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-3883-9608ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Slice: A Selective Local Inference Framework with Codec Exploitation for Accelerating Video Super-Resolution
Mingu Jung, Sungbin Kim, Seunghyun Jin 0002, Hyunwuk Lee, Won Woo Ro
ISCA4
2025 PIMFY: Eliminating Remote Page Walks in MCM GPUs
abstract
Multi-Chip Module (MCM) GPUs are suffering from non-uniform memory access (NUMA) due to communication through in-package interconnects between chiplets. As chiplet-to-chiplet communication increases, the performance saturates even with scaled hardware resources in MCM GPUs. Previous works have attempted to mitigate the NUMA caused by data page access between chiplets. However, we observe that page table page access for address translation requests also generates significant NUMA and degrades overall performance of MCM GPUs. In this paper, we analyze the impact of remote page table page access that causes significant challenges in MCM GPUs. Based on our analysis, we propose Page tables In My Front Yard (PIMFY), a technique that eliminates remote page table page access and accelerates page walks, thus improving overall performance in MCM GPUs. Exploiting static address mapping nature during GPU kernel execution, PIMFY replicates page tables onto all chiplets and prevents remote page table access. Our evaluation shows that PIMFY reduces average page walk latency by 27.07%, and enhances overall performance by$1.21 \times$.
Junsung Kim 0002, Sungwoo Kim 0003, Seunghyun Jin 0002, Won Woo Ro
ICCD3
2024 GUMSO: Gating Unnecessary On-Chip Memory Slices for Power Optimization on GPUs
abstract
The importance of power efficiency in GPUs has grown significantly for data centers, as it directly impacts costs and sustainability. While there have been many works on power optimization for GPUs, their approaches often target applications consuming high parallelism for their entire application sequence. However, graph applications, characterized by their large data size and inherent irregularity in graph structures, present distinct challenges as they suffer from kernel launches with limited parallelism, resulting in under-utilization of GPU hardware components. Even with this under-utilization, last-level cache (LLC) and network-on-chip (NoC) of the GPU are fully activated even with a single thread running, incurring large leakage power. In our analysis, we find out that LLC and NoC contribute up to 39.9% of the total GPU power consumption during graph applications. To address this power inefficiency, we propose GUMSO, an energy-efficient design that enables adaptive power-gating of LLC slices for small kernel executions in GPUs. By managing the utilization of LLC slices, our approach reduces GPU energy consumption by an average of 18.3% across various graph applications with minimum performance overheads.
Seunghyun Jin 0002, Hyunwuk Lee, Won Woo Ro
ISLPED1
2024 SHREG: Mitigating register redundancy in GPUs
Seunghyun Jin 0002, Hyunwuk Lee, Junsung Kim 0002, Won Woo Ro
J. Syst. Archit.1
2023 INTERPRET: Inter-Warp Register Reuse for GPU Tensor Core
abstract
Tensor cores in the recent NVIDIA GPUs are under the spotlight due to their superior computation throughput for general matrix-matrix multiplication (GEMM) that has been widely used for deep learning applications. For massive-scale GEMMs, the entire matrix is practically divided into sub-matrices and assigned to multiple thread blocks and warps, and then processed by the tensor cores. Meanwhile, the same sub-matrix is regularly reused as an input to different sub-GEMMs, which causes redundant load operations from different warps and waste of register file spaces. To tackle this issue, we propose INTERPRET, a novel tensor core microarchitecture designed to minimize unnecessary accesses to the cache/memory hierarchy by leveraging the inter-warp data reuse characteristics. INTERPRET adopts a register renaming scheme to reduce the redundant load requests as well as the waste of register files, resulting in the reduction of the effective data load latency. INTERPRET further improves performance via non-speculative tensor preloading by leveraging the register file space saved by the register renaming. As INTERPRET is implemented based on the data access patterns of tensor core operations exhibiting a high level of regularity, the synergistic integration of the register renaming and tensor preloading can significantly improve the processing efficiency. Our experiments show that the proposed design achieves an average speedup of 34.1% and reduces energy consumption by 27.9%.
Jae Seok Kwak, Myung Kuk Yoon, Ipoom Jeong, Seunghyun Jin 0002, Won Woo Ro
PACT4