EDBT 2026 Demo / reviewers in the wild / expert
Young-Ho Gong
dblp:130/7755
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-8270-7875ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Heterogeneity-Aware Optimizations for Resource Efficient Edge RecommendationabstractRecommendation systems are widely deployed on edge devices to enable personalized user experience. While recommendation inference has traditionally been performed on centralized servers, recent advances in mobile SoCs have motivated a shift toward on-device execution. However, achieving efficient recommendation inference on edge devices remains challenging due to edge-specific execution characteristics and heterogeneity. In this paper, we characterize resource inefficiencies under realistic edge constraints and propose optimization strategies. Yerin Lee, Gyudong Kim, Eunjin Lee, Jeff Zhang 0001, Young-Ho Gong, Carole-Jean Wu |
DATE | 5 |
| 2026 | LayUp: Layer-wise Parallelization for Energy-Efficient Edge LLM Training Exploiting Unified Memory Characteristics
Bang-San Lee, Young-Ho Gong |
ISLPED | 2 |
| 2025 | SHIFT ECC: A Value Converting HBM ECC Approach for Refresh Energy Efficient Integer Quantized DNN InferenceabstractAs the parameter size of deep neural networks (DNNs) increases, high bandwidth memory (HBM) is widely adopted to satisfy the growing demand for memory bandwidth. However, due to the shorter retention time caused by higher on-chip temperature, HBM requires more frequent refresh operations, resulting in significant refresh energy and performance overhead. In this paper, we propose SHIFT ECC, a lightweight and robust ECC scheme for INT8 quantized DNNs on HBM, to reduce refresh operations while maintaining inference accuracy. SHIFT ECC enhances DNN reliability by converting negative weights into positive weights, eventually mitigating the primary cause of retention errors (mostly 1→0 bit errors). Additionally, SHIFT ECC applies stronger ECC to the upper bits (more important bits) of DNN weights while protecting the lower bits (less important bits) with weaker ECC, which further enhances the robustness of DNNs with the same number of parity bits. Our evaluation results show that when the proportion of 1→0 bit errors is 100% and 99%, SHIFT ECC reduces average refresh energy by 32.6% and 35.0%, respectively, reducing average memory read latency by 21.7% compared to the state-of-the-art refresh reduction technique. Jae Yoon Lee, Young Seo Lee, Young-Ho Gong, Seon Wook Kim, Sung Woo Chung |
ISLPED | 3 |
| 2024 | Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryabstractWe propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that mandate memory oversubscription, resulting in substantial speedups. However, the DRAM cache needs to be carefully designed to address the latency and bandwidth limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can easily thrash the DRAM cache and degrade performance, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by G PU s with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probe traffic and increase effective DRAM BW with minimal cost overhead, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache implementation with Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power consumption and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, the HMS improves performance by up to 12.5× (2.9× overall) and reduces energy by up to 89.3% (48.1 % overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91–93 % and 57–75 %, respectively. Jeongmin Hong 0001, Sungjun Cho, Geonwoo Park, Wonhyuk Yang, Young-Ho Gong, Gwangsun Kim |
HPCA | 5 |
| 2024 | Sparrow ECC: A Lightweight ECC Approach for HBM Refresh Reduction towards Energy-efficient DNN InferenceabstractExponential growth in deep neural network (DNN) model size has resulted in significant demands for memory bandwidth, leading to the extensive adoption of high bandwidth memory (HBM) in DNN inference. However, with the shorter retention time due to high operating temperature, HBM requires more frequent refresh operations, suffering larger refresh energy/performance overhead. In this paper, we propose Sparrow ECC, a lightweight but stronger HBM ECC technique for less refresh operations while preserving inference accuracy. Sparrow ECC exploits the dominant exponent pattern (i.e., value similarity) in pre-trained DNN weights, limiting the exponent value range of the pre-trained weights to prevent anomalously large weight value change due to the errors. In addition, through duplication and single error correction (SEC) code, Sparrow ECC strongly protects the critical bits in DNN weights. In our evaluation, when the proportion of 1→0 bit errors is 100% and 99%, Sparrow ECC reduces the refresh energy consumption by 90.40% and 93.22%, on average, respectively, compared to the state-of-the-art (RS(19,17)+ZEM [22]) refresh reduction technique, while preserving inference accuracy. Hoseok Kim, Seung Hun Choi, Joonho Kong, Young-Ho Gong, Sung Woo Chung |
ISLPED | 4 |
| 2023 | Twin ECC: A Data Duplication Based ECC for Strong DRAM Error ResilienceabstractWith the continuous scaling of process technology, DRAM reliability has become a critical challenge in modern memory systems. Currently, DRAM memory systems for servers employ ECC DIMMs with a single error correction and double error detection (SECDED) code. However, the SECDED code is insufficient to ensure DRAM reliability since memory systems become more susceptible to errors. Though various studies have proposed multi-bit correctable ECC schemes, such ECC schemes cause performance and/or storage overhead. To minimize performance degradation while providing strong error resilience, in this paper, we propose Twin ECC, a low-cost memory protection scheme through data duplication. In a 512-bit data, Twin ECC duplicates meaningful data into meaningless zeros. Since ‘1’$\rightarrow$‘0’ error pattern is dominant in DRAM cells, Twin ECC provides strong error resilience by performing bitwise OR operations between the original meaningful data and duplicated data. After the bitwise OR operations, Twin ECC adopts the SECDED code for further enhancing data protection. Our evaluations show that Twin ECC reduces the system failure probability by average 64.8%, 56.9%, and 49.5%, when the portion of '1 ‘$\rightarrow$‘0’ error is 100%, 90%, and 80%, respectively, while causing only 0.7% performance overhead and no storage overhead compared to the baseline ECC DIMM with SECDED code. Hyeong Kon Bae, Myung Jae Chung, Young-Ho Gong, Sung Woo Chung |
DATE | 3 |
| 2023 | Scale-CIM: Precision-scalable computing-in-memory for energy-efficient quantized neural networks
Young Seo Lee, Young-Ho Gong, Sung Woo Chung |
J. Syst. Archit. | 2 |
| 2022 | Stealth ECC: A Data-Width Aware Adaptive ECC Scheme for DRAM Error ResilienceabstractAs DRAM process technology scales down and DRAM density continues to grow, DRAM errors have become a primary concern in modern data centers. Typically, data centers have adopted memory systems with a single error correction double error detection (SECDED) code. However, the SECDED code is not sufficient to satisfy DRAM reliability demands as memory systems get more vulnerable. Though the servers in data centers employ strong ECC schemes, such ECC schemes lead to substantial performance and/or storage overhead. In this paper, we propose Stealth ECC, a cost-effective memory protection scheme providing stronger error correctability than the conventional SECDED code, with negligible performance overhead and without storage overhead. Depending on the data-width (either narrow-width or full-width), Stealth ECC adaptively selects ECC schemes. For narrow-width values, Stealth ECC provides multi-bit error correctability by storing more parity bits in MSB side, instead of zeros. Furthermore, with bitwise interleaved data placement between x4 DRAM chips, Stealth ECC is robust to a single DRAM chip error for narrow-width values. On the other hand, for full-width values, Stealth ECC adopts the SECDED code, which maintains DRAM reliability comparable to the conventional SECDED code. As a result, thanks to the reliability improvement of narrow-width values, Stealth ECC enhances overall DRAM reliability, while incurring negligible performance overhead as well as no storage overhead. Our simulation results show that Stealth ECC reduces the probability of system failure (caused by DRAM errors) by 47.9%, on average, with only 0.9% performance overhead compared to the conventional SECDED code. Young Seo Lee, Gunjae Koo, Young-Ho Gong, Sung Woo Chung |
DATE | 3 |
| 2021 | Monolithic 3D stacked multiply-accumulate units
Young Seo Lee, Ji Heon Lee, Young-Ho Gong, Seon Wook Kim, Sung Woo Chung |
Integr. | 4 |
| 2019 | Exploring the Relation between Monolithic 3D L1 GPU Cache Capacity and Warp Scheduling EfficiencyabstractThe warp scheduler plays an important role in the GPU for efficient utilization of hardware resources. However, the efficiency of the warp scheduler is often limited by the L1 cache (especially, L1 data cache) capacity; providing large capacity for an L1 cache is challenging due to the increased latency. In this paper, we adopt Monolithic 3D (M3D) technology to design a large capacity L1 cache for GPU performance enhancement, not deteriorating the latency. Our evaluation results show that the M3D L1 cache improves GPU performance by 2.18~2.24× on average, compared to the 2D conventional L1 cache. Cong Thuan Do, Young-Ho Gong, Cheol Hong Kim, Seon Wook Kim, Sung Woo Chung |
ISLPED | 2 |
| 2017 | Architecting large-scale SRAM arrays with monolithic 3D integrationabstractIn this paper, we architect large-scale SRAM arrays with monolithic 3D (M3D) integration technology. We introduce M3D-based SRAM arrays with three different ways of integration: M3D-R (vertical routing-only), M3D-VBL (vertical bitline), and M3D-VWL (vertical wordline). We also apply M3D-based SRAM arrays to last-level caches: tag arrays for eDRAM LLCs and data arrays for SRAM LLCs. The proposed LLCs with M3D-based SRAM arrays lead to better performance and lower energy by 0.02%∼1.7% and 49.1%∼79.1%, respectively, compared to that with TSV-based 3D SRAM arrays. Joonho Kong, Young-Ho Gong, Sung Woo Chung |
ISLPED | 2 |
| 2016 | Exploiting Refresh Effect of DRAM Read Operations: A Practical Approach to Low-Power RefreshabstractDynamic random access memory (DRAM) requires periodic refresh operations to retain its data. In practice, DRAM retention times are normally distributed from 64 ms to several seconds. However, the conventional refresh method uses 64 ms as the refresh interval, since it applies the same refresh interval to all DRAM rows. Thus, the conventional refresh method results in unnecessary refresh operations (eventually, energy waste) to the DRAM rows whose retention times are longer than 64 ms. In this paper, we propose a practical refresh scheme that exploits refresh effect of DRAM read operations to reduce refresh overhead. Our proposed scheme applies a refresh interval longer than the conventional refresh interval (64 ms) to the DRAM chip. In this case, weak DRAM rows (DRAM rows whose retention times are shorter than the refresh interval of the DRAM chip) cannot retain their data. In order to retain the data stored in the weak DRAM rows, the memory controller issues read operations to the weak DRAM rows every required refresh interval for the weak DRAM rows. Our evaluation results show that our proposed scheme with 192 ms refresh interval reduces average refresh energy consumption up to 66.0 percent, which in turn reduces average DRAM energy consumption up to 31.8 percent, compared to the conventional refresh method (64 ms). Our proposed scheme requires no modification to internal DRAM chip structures, but it only adds a small weak row buffer (the buffer for the weak row information) to the memory controller, which has a negligible area overhead. Young-Ho Gong, Sung Woo Chung |
IEEE Trans. Computers | 1 |