EDBT 2026 Demo / reviewers in the wild / expert
Jihun Yoon
dblp:267/9556
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0003-2777-2262ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting LLC Bandwidth Utilization in GPUs through Adaptive Fine-Grained Data MigrationabstractModern server-grade GPUs (e.g., NVIDIA A100) integrate hundreds of cores and tens of memory partitions, providing massive compute capability and memory bandwidth. However, the increased scale amplifies interconnect overhead between cores and memory partitions. To mitigate this, NVIDIA A100 clusters multiple cores and memory partitions into two large groups, thereby simplifying interconnect complexity. Unfortunately, this partitioning introduces a new limitation: remote partition accesses. A core accessing a remote memory partition incurs higher latency and lower bandwidth compared to local accesses.In this paper, we propose a cache-line migration mechanism across partitions to alleviate remote memory access overhead. Our design is motivated by two key observations: (1) conventional GPUs employ limited and often ineffective optimizations for remote access handling, and (2) GPU applications typically exhibit high temporal locality, where a specific partition of cores makes frequent memory accesses for the same data within short time intervals. Leveraging these insights, we dynamically migrate cache-lines to the local partition where the requesting core resides. Experimental results demonstrate that our approach achieves up to 1.24× speedup over the baseline with NVIDIA A100-like replication, highlighting its effectiveness in reducing remote access penalties. Jihun Yoon, Sungbin Jang, Seokin Hong |
DATE | 1 |
| 2025 | Zebra: Leveraging Diagonal Attention Pattern for Vision Transformer AcceleratorabstractVision Transformers (ViTs) have achieved remarkable performance in computer vision, but their computational complexity and challenges in optimizing memory bandwidth limit hardware acceleration. A major bottleneck lies in the self-attention mechanism, which leads to excessive data movement and unnecessary computations despite high input sparsity and low computational demands. To address this challenge, existing transformer accelerators have leveraged sparsity in attention maps. However, their performance gains are limited due to low hardware utilization caused by the irregular distribution of nonzero values in the sparse attention maps. Self-attention often exhibits strong diagonal patterns in the attention map, as the diagonal elements tend to have higher values than others. To exploit this, we introduce Zebra, a hardware accelerator framework optimized for diagonal attention patterns. A core component of Zebra is the Striped Diagonal (SD) pruning technique, which prunes the attention map by preserving only the diagonal elements at runtime. This reduces computational load without requiring offline pre-computation or causing significant accuracy loss. Zebra features a reconfigurable accelerator architecture that supports optimized matrix multiplication method, called Striped Diagonal Matrix Multiplication (SDMM), which computes only the diagonal elements of matrices. With this novel method, Zebra addresses low hardware utilization, a key barrier to leveraging the diagonal patterns. Experimental results demonstrate that Zebra achieves a 57 × speedup over a CPU and 1.7 × over the state-of-the-art ViT accelerator with similar inference accuracy. Sukhyun Han, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Seokin Hong |
DATE | 4 |
| 2025 | Minimizing Read Disturb via Localized Page Allocation for Modern NAND Flash-Based SSDsabstractTo meet the increasing demand for higher storage density, modern NAND flash-based SSDs employ Quad-Level Cell (QLC) technology, which stores four bits per memory cell. However, it exacerbates the read disturb issue, where repeated read operations gradually shift the threshold voltages of unselected NAND flash cells. This voltage drift leads to frequent read retries and accelerates device wear, ultimately degrading performance and endurance. To address this issue, we propose Localized Page Allocation (LPA), a novel data mapping scheme that allocates each page to a restricted region of the wordline. LPA assigns four consecutive bits in a single page to a single cell. This approach confines read operations to a limited portion of the bitlines, thereby reducing the number of NAND flash cells exposed to read disturb. However, this localized access increases the complexity of sensing operations. To address this challenge, we introduce Progressive Voltage Convergence (PVC) method that adaptively determines the next sensing voltage for each bitline based on its previous sensing result. This fine-grained control reduces the number of sensing steps during read operations. For additional optimization, we employ lossless compression, which reduces the number of sensing steps for the compressed pages. Experimental evaluations using real-world workloads demonstrate that our design improves I/O performance by up to$\mathbf{1 1 \%}$and extends endurance by 79% compared to the state-of-the-art techniques. Our design requires minor hardware modifications to the conventional NAND flash chips and the SSD controller, which leads to small area and power overhead. Joonseong Hwang, Minjin Park, Jihun Yoon, Yoonho Jang, Seokin Hong |
ICCD | 4 |
| 2024 | Towards Precise Pose Estimation in Robotic Surgery: Introducing Occlusion-Aware Loss
Jiuk Hong, Jihun Yoon, Bokyung Park, Min-Kook Choi, Heechul Jung |
MICCAI (6) | 3 |
| 2024 | A Case for Speculative Address Translation with Rapid Validation for GPUsabstractA unified address space is vital for heterogeneous systems as it enables efficient data sharing between CPUs and GPUs. However, GPU address translation faces challenges due to high TLB pressure, particularly with irregular and memory-intensive applications. Compared to an ideal scenario, we observe that address translation overheads cause a slowdown of up to 34.5% in modern heterogeneous systems. This paper introduces Avatar, a novel framework to accelerate address translation in GPUs. Avatar comprises two key components: Contiguity-Aware Speculative Translation (CAST) and In-Cache Validation (CAVA) mechanisms. Avatar identifies the potential for predicting virtual-to-physical address mapping by monitoring contiguous pages that lie in both virtual and physical address spaces. Leveraging this insight, CAST speculatively translates virtual addresses into physical addresses. This speculative address translation enables immediate data fetching into GPUs while addressing translation occurs in the background, reducing TLB-miss overhead. Unfortunately, modern GPUs lack support for speculative execution, which limits CAST's performance gain. Data fetched from speculated physical addresses is unusable until validation. CAVA addresses this limitation by quickly validating speculated physical addresses. To this end, CAVA embeds page mapping information into each 32B sector of 128B cache lines. Thus, CAVA enables fetching a sector block from memory for a speculated address and rapidly validating the speculative translation using the embedded mapping information. Our experiments show that Avatar achieves a 90.3% (high) speculation accuracy and improves GPU performance by 37.2% (on average). Junhyeok Park 0001, Osang Kwon, Seongwook Kim, Gwangeun Byeon, Jihun Yoon, Prashant J. Nair, Seokin Hong |
MICRO | 6 |
| 2022 | Surgical Scene Segmentation Using Semantic Image Synthesis with a Virtual Surgery Environment
Jihun Yoon, SeulGi Hong, Seungbum Hong, Soyeon Shin, Bokyung Park, Nakjun Sung, Hayeong Yu, Sungjae Kim, Woo Jin Hyung, Min-Kook Choi |
MICCAI (8) | 1 |
| 2022 | Self-Supervised Knowledge Transfer via Loosely Supervised Auxiliary TasksabstractKnowledge transfer using convolutional neural networks (CNNs) can help efficiently train a CNN with fewer parameters or maximize the generalization performance under limited supervision. To enable a more efficient transfer of pretrained knowledge under relaxed conditions, we propose a simple yet powerful knowledge transfer methodology without any restrictions regarding the network structure or dataset used, namely self-supervised knowledge transfer (SSKT), via loosely supervised auxiliary tasks. For this, we devise a training methodology that transfers previously learned knowledge to the current training process as an auxiliary task for the target task through self-supervision using a soft label. The SSKT is independent of the network structure and dataset, and is trained differently from existing knowledge transfer methods; hence, it has an advantage in that the prior knowledge acquired from various tasks can be naturally transferred during the training process to the target task. Furthermore, it can improve the generalization performance on most datasets through the proposed knowledge transfer between different problem domains from multiple source networks. SSKT outperforms the other transfer learning methods (KD, DML, and MAXL) through experiments under various knowledge transfer settings. The source code will be made available to the public1. Seungbum Hong, Jihun Yoon, Min-Kook Choi, Junmo Kim 0002 |
WACV | 2 |
| 2021 | Semi-Supervised Object Detection With Sparsely Annotated DatasetabstractWhen training an anchor-based object detector with a sparsely annotated dataset, the effort required to locate positive examples can cause performance degradation. Because anchor-based object detection models collect positive examples under IoU between anchors and ground-truth bounding boxes, in a sparsely annotated image, some objects that are not annotated can be assigned as negative examples, such as backgrounds. We attempt to solve this problem with two approaches: 1) using an anchor-less object detector and 2) using a single-object tracker for semi-supervised learning-based object detection. The proposed technique performs bidirectional single-object tracking from sparsely annotated bounding boxes as starting points in videos to obtain dense annotations. On applying our method to the EPIC-KITCHENS-55 dataset, we were able to achieve runner-up performance in the Unseen section, while achieving the first place in the Seen section of the EPIC-KITCHENS 2020 object detection challenge under IoU > 0.5 on the EPIC-KITCHENS 2020 object detection challenge. Jihun Yoon, Seungbum Hong, Min-Kook Choi |
ICIP | 1 |
| 2021 | hSDB-instrument: Instrument Localization Database for Laparoscopic and Robotic Surgeries
Jihun Yoon, Sunghwan Heo, Hayeong Yu, Jayeon Lim, Chihyun Song, SeulGi Hong, Seungbum Hong, Bokyung Park, Woo Jin Hyung, Min-Kook Choi |
MICCAI (4) | 1 |