VLDB 2026 Research / reviewers in the wild / expert
Sung Woo Chung
dblp:46/2413
· DBLP profile ↗
50ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-5347-9586ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SHIFT ECC: A Value Converting HBM ECC Approach for Refresh Energy Efficient Integer Quantized DNN InferenceabstractAs the parameter size of deep neural networks (DNNs) increases, high bandwidth memory (HBM) is widely adopted to satisfy the growing demand for memory bandwidth. However, due to the shorter retention time caused by higher on-chip temperature, HBM requires more frequent refresh operations, resulting in significant refresh energy and performance overhead. In this paper, we propose SHIFT ECC, a lightweight and robust ECC scheme for INT8 quantized DNNs on HBM, to reduce refresh operations while maintaining inference accuracy. SHIFT ECC enhances DNN reliability by converting negative weights into positive weights, eventually mitigating the primary cause of retention errors (mostly 1→0 bit errors). Additionally, SHIFT ECC applies stronger ECC to the upper bits (more important bits) of DNN weights while protecting the lower bits (less important bits) with weaker ECC, which further enhances the robustness of DNNs with the same number of parity bits. Our evaluation results show that when the proportion of 1→0 bit errors is 100% and 99%, SHIFT ECC reduces average refresh energy by 32.6% and 35.0%, respectively, reducing average memory read latency by 21.7% compared to the state-of-the-art refresh reduction technique. Jae Yoon Lee, Young Seo Lee, Young-Ho Gong, Seon Wook Kim, Sung Woo Chung |
ISLPED | 5 |
| 2025 | Thermal Challenges and Opportunities for Off-the-shelf 3D-stacked CPUsabstractIn recent years, 3D stacking has emerged as a promising technology for high-performance CPUs, as it offers higher yield and improved inter-die bandwidth. However, 3D CPUs are more vulnerable to thermal problems than conventional 2D CPUs due to increased power density and limited heat dissipation capabilities. In this paper, we analyze the thermal characteristics of off-the-shelf 2D and 3D CPUs, comparing them in terms of performance and on-chip temperature. In addition, to mitigate the thermal vulnerability of 3D CPUs, we introduce two thermal-aware scheduling techniques: 1) floorplan-based thermal-aware scheduling, and 2) adaptive voltage scaling (AVS)-based thermal-aware scheduling. Floorplan-based thermal-aware scheduling primarily assigns tasks to cores that are advantageous for heat dissipation based on the floorplan to mitigate thermal hotspots. AVS-based thermal-aware scheduling prioritizes the allocation of tasks to cores with lower power consumption, considering different power consumption due to process variation. Our evaluation results demonstrate that floorplan-based and AVS-based thermal-aware scheduling reduce energy consumption by 10.3% and 12.4%, respectively, compared to the legacy Linux scheduler, while maintaining performance. Jae Yoon Lee, Chae Young Sim, Seung Hun Choi, Sung Woo Chung |
ISLPED | 4 |
| 2024 | ComBoost: An Instruction Complexity Aware DTM Technique for Edge DevicesabstractRecent edge devices show high power density in CPUs, resulting in excessive heat generation. Since mechanical cooling solutions are impractical in edge devices due to their small form factor, software-controlled dynamic thermal management (DTM) plays a crucial role in resolving thermal problems. In state-of-the-art edge devices, proactive DTM techniques such as ARM intelligent power allocation (IPA) mainly exploit the current CPU status (e.g., on-chip temperature, core utilization, and frequency) to estimate the current power consumption which eventually affects the future on-chip temperature. However, they overlook the impact of instruction complexity on thermal behaviors, which results in too conservative or aggressive voltage and frequency control. Even with the same frequency and core utilization, the on-chip temperature increases with different gradients depending on the instruction complexity of workloads. In this paper, we propose an instruction complexity aware DTM technique for edge devices, called ComBoost. Based on the real-time monitoring of on-chip temperature, utilization, and frequency, ComBoost examines the instruction complexity as well as the current CPU status to determine the target frequency. ComBoost then proactively adjusts the voltage and frequency of cores to minimize the performance degradation from thermal throttling. In the off-the-shelf edge device, ComBoost improves performance by 16.8%, 18.6%, and 15.5%, on average, compared to the legacy, IPA, and prior RL-based technique, respectively. Seung Hun Choi, Joonho Kong, Sung Woo Chung |
ISLPED | 3 |
| 2024 | Sparrow ECC: A Lightweight ECC Approach for HBM Refresh Reduction towards Energy-efficient DNN InferenceabstractExponential growth in deep neural network (DNN) model size has resulted in significant demands for memory bandwidth, leading to the extensive adoption of high bandwidth memory (HBM) in DNN inference. However, with the shorter retention time due to high operating temperature, HBM requires more frequent refresh operations, suffering larger refresh energy/performance overhead. In this paper, we propose Sparrow ECC, a lightweight but stronger HBM ECC technique for less refresh operations while preserving inference accuracy. Sparrow ECC exploits the dominant exponent pattern (i.e., value similarity) in pre-trained DNN weights, limiting the exponent value range of the pre-trained weights to prevent anomalously large weight value change due to the errors. In addition, through duplication and single error correction (SEC) code, Sparrow ECC strongly protects the critical bits in DNN weights. In our evaluation, when the proportion of 1→0 bit errors is 100% and 99%, Sparrow ECC reduces the refresh energy consumption by 90.40% and 93.22%, on average, respectively, compared to the state-of-the-art (RS(19,17)+ZEM [22]) refresh reduction technique, while preserving inference accuracy. Hoseok Kim, Seung Hun Choi, Joonho Kong, Young-Ho Gong, Sung Woo Chung |
ISLPED | 5 |
| 2023 | Twin ECC: A Data Duplication Based ECC for Strong DRAM Error ResilienceabstractWith the continuous scaling of process technology, DRAM reliability has become a critical challenge in modern memory systems. Currently, DRAM memory systems for servers employ ECC DIMMs with a single error correction and double error detection (SECDED) code. However, the SECDED code is insufficient to ensure DRAM reliability since memory systems become more susceptible to errors. Though various studies have proposed multi-bit correctable ECC schemes, such ECC schemes cause performance and/or storage overhead. To minimize performance degradation while providing strong error resilience, in this paper, we propose Twin ECC, a low-cost memory protection scheme through data duplication. In a 512-bit data, Twin ECC duplicates meaningful data into meaningless zeros. Since ‘1’$\rightarrow$‘0’ error pattern is dominant in DRAM cells, Twin ECC provides strong error resilience by performing bitwise OR operations between the original meaningful data and duplicated data. After the bitwise OR operations, Twin ECC adopts the SECDED code for further enhancing data protection. Our evaluations show that Twin ECC reduces the system failure probability by average 64.8%, 56.9%, and 49.5%, when the portion of '1 ‘$\rightarrow$‘0’ error is 100%, 90%, and 80%, respectively, while causing only 0.7% performance overhead and no storage overhead compared to the baseline ECC DIMM with SECDED code. Hyeong Kon Bae, Myung Jae Chung, Young-Ho Gong, Sung Woo Chung |
DATE | 4 |
| 2023 | Scale-CIM: Precision-scalable computing-in-memory for energy-efficient quantized neural networks
Young Seo Lee, Young-Ho Gong, Sung Woo Chung |
J. Syst. Archit. | 3 |
| 2023 | Aggressive GPU cache bypassing with monolithic 3D-based NoC
Cong Thuan Do, Cheol Hong Kim, Sung Woo Chung |
J. Supercomput. | 3 |
| 2023 | Correction to: Aggressive GPU cache bypassing with monolithic 3D-based NoC
Cong Thuan Do, Cheol Hong Kim, Sung Woo Chung |
J. Supercomput. | 3 |
| 2022 | Stealth ECC: A Data-Width Aware Adaptive ECC Scheme for DRAM Error ResilienceabstractAs DRAM process technology scales down and DRAM density continues to grow, DRAM errors have become a primary concern in modern data centers. Typically, data centers have adopted memory systems with a single error correction double error detection (SECDED) code. However, the SECDED code is not sufficient to satisfy DRAM reliability demands as memory systems get more vulnerable. Though the servers in data centers employ strong ECC schemes, such ECC schemes lead to substantial performance and/or storage overhead. In this paper, we propose Stealth ECC, a cost-effective memory protection scheme providing stronger error correctability than the conventional SECDED code, with negligible performance overhead and without storage overhead. Depending on the data-width (either narrow-width or full-width), Stealth ECC adaptively selects ECC schemes. For narrow-width values, Stealth ECC provides multi-bit error correctability by storing more parity bits in MSB side, instead of zeros. Furthermore, with bitwise interleaved data placement between x4 DRAM chips, Stealth ECC is robust to a single DRAM chip error for narrow-width values. On the other hand, for full-width values, Stealth ECC adopts the SECDED code, which maintains DRAM reliability comparable to the conventional SECDED code. As a result, thanks to the reliability improvement of narrow-width values, Stealth ECC enhances overall DRAM reliability, while incurring negligible performance overhead as well as no storage overhead. Our simulation results show that Stealth ECC reduces the probability of system failure (caused by DRAM errors) by 47.9%, on average, with only 0.9% performance overhead compared to the conventional SECDED code. Young Seo Lee, Gunjae Koo, Young-Ho Gong, Sung Woo Chung |
DATE | 4 |
| 2021 | Monolithic 3D stacked multiply-accumulate units
Young Seo Lee, Ji Heon Lee, Young-Ho Gong, Seon Wook Kim, Sung Woo Chung |
Integr. | 6 |
| 2021 | Thermal-aware adaptive VM allocation considering server locations in heterogeneous data centers
Younggeun Kim 0001, Seon Young Kim, Seung Hun Choi, Sung Woo Chung |
J. Syst. Archit. | 4 |
| 2020 | An Adaptive Thermal Management Framework for Heterogeneous Multi-Core ProcessorsabstractOff-the-shelf embedded systems have adopted heterogeneous multi-core processors which have high-performance big cores and low-power small cores. Though there are two different types of cores in heterogeneous multi-core processors, conventional DVFS (Dynamic Voltage and Frequency Scaling)-based DTM (Dynamic Thermal Management) techniques do not utilize the different types of cores to cool down hot cores. Rather, they primarily reduce the voltage and frequency of the hot cores, leading to performance degradation. In this article, we propose a novel adaptive DTM framework for heterogeneous multi-core processors, which utilizes the big and small cores to prevent performance degradation. Our proposed framework exploits two migration-based DTM techniques: 1) a technique (denoted as Migrationbig↔big) that migrates applications from hot big cores (big cores whose temperature is above a pre-defined threshold) to cold big cores (big cores whose temperature is below the threshold) and 2) a technique (denoted as Migrationbig↔small) that migrates all applications from the big cores to the small cores. In case of thermal emergency of the big cores, our proposed framework checks the number of cold big cores. When there exist available cold big cores, our proposed framework employs Migrationbig↔big to cool down the hot big cores while not reducing the big core frequency. On the other hand, when there does not exist any available cold big core, our proposed framework employs one between Migrationbig↔small and a DVFS-based DTM technique, which is expected to result in better performance. In our experiments on an embedded development board, our proposed framework improves the average performance by 8.9 percent, compared to ARM's DVFS-based IPA (Intelligent Power Allocation), satisfying thermal constraints. Our framework also improves the average performance by 10.4 percent, compared to a state-of-the-art predictive DVFS-based DTM technique. Younggeun Kim 0001, Minyong Kim, Joonho Kong, Sung Woo Chung |
IEEE Trans. Computers | 4 |
| 2020 | Signal Strength-Aware Adaptive Offloading with Local Image Preprocessing for Energy Efficient Mobile DevicesabstractTo prolong battery life of mobile devices, image processing applications often exploit offloading techniques which run some or all of the computations on remote servers. Unfortunately, the existing offloading techniques do not consider the fact that data transmission time and energy consumption of wireless network interfaces exponentially increase when signal strength decreases. In this paper, we propose an adaptive offloading for image processing applications, which considers wireless signal strength. To improve performance and energy efficiency of offloading, we also propose to adaptively exploit local preprocessing (executing image preprocessing on local mobile devices), considering wireless signal strength; the local preprocessing usually reduces the size of transmission image in offloading. Our proposed technique estimates performance and energy consumption of the following three methods, depending on the wireless signal strength: 1) local execution (executing all the computations on the local mobile devices), 2) offloading without local preprocessing, and 3) offloading with local preprocessing. Based on the estimated performance and energy consumption, our technique employs one among the three methods, which is expected to result in the best performance or energy efficiency. In our evaluation on an off-the-shelf smartphone, when a user prefers performance to energy, our proposed technique improves performance by 27.1 percent, compared to the conventional offloading technique that does not consider the signal strength. On the other hand, when a user prefers energy to performance, our proposed technique saves system-wide (not just CPU nor wireless network interface) energy consumption by 26.3 percent, on average, compared to the conventional offloading technique. Younggeun Kim 0001, Young Seo Lee, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2020 | A novel warp scheduling scheme considering long-latency operations for high-performance GPUs
Cong Thuan Do, Hong Jun Choi, Sung Woo Chung, Cheol Hong Kim |
J. Supercomput. | 3 |
| 2019 | A High-Performance Processing-in-Memory Accelerator for Inline Data DeduplicationabstractIn data centers, inline data deduplication which eliminates redundant data on the fly, is crucial to significantly reduce storage cost. However, it causes substantial performance and energy overhead due to a large number of memory accesses in the conventional GPU. In this paper, we propose a highperformance processing-in-memory accelerator for inline data deduplication, called Deduplication Unit (DU) to reduce the latency and power consumption. We place the DUs in a base die or core dies of a 3D stacked memory to improve performance. Our simulation results show that the DUs in the base die reduce the latency and processing unit power consumption by 17.3% and 45.5%, on average, respectively, compared to the conventional GPU. In addition, in our thermal simulation, peak temperature of the DU is still lower than the threshold temperature. Young Seo Lee, Ji Heon Lee, Jeong Hwan Choi, Sung Woo Chung |
ICCD | 5 |
| 2019 | Exploring the Relation between Monolithic 3D L1 GPU Cache Capacity and Warp Scheduling EfficiencyabstractThe warp scheduler plays an important role in the GPU for efficient utilization of hardware resources. However, the efficiency of the warp scheduler is often limited by the L1 cache (especially, L1 data cache) capacity; providing large capacity for an L1 cache is challenging due to the increased latency. In this paper, we adopt Monolithic 3D (M3D) technology to design a large capacity L1 cache for GPU performance enhancement, not deteriorating the latency. Our evaluation results show that the M3D L1 cache improves GPU performance by 2.18~2.24× on average, compared to the 2D conventional L1 cache. Cong Thuan Do, Young-Ho Gong, Cheol Hong Kim, Seon Wook Kim, Sung Woo Chung |
ISLPED | 5 |
| 2019 | Temperature-aware Adaptive VM Allocation in Heterogeneous Data CentersabstractVirtualized data centers usually consist of heterogeneous servers which have different specifications (performance). Though there are usually a number of unused servers with different performance in such heterogeneous data centers, conventional DVFS (Dynamic Voltage and Frequency Scaling)-based DTM (Dynamic Thermal Management) techniques do not exploit the unused servers to cool down hot servers. In this paper, we propose a novel DTM technique which adaptively exploits external computing resources (unused servers with different performance) as well as internal computing resources (unused CPU cores in the server) available in heterogeneous data centers. When the temperature of a CPU core in a server exceeds a pre-defined thermal threshold, our proposed technique first identifies memory intensiveness and usage of VMs (Virtual Machines). Depending on the memory intensiveness and usage of VMs, our technique adaptively employs the following three methods: 1) a method that migrates a VM to another server with different performance, 2) a method that migrates VMs among CPU cores in the server, and 3) a DVFS-based method. In our experiments, our proposed technique improves performance by 9.6% and saves system-wide EDP by 12.9%, on average (by up to 17.1% and 24.5%, respectively), compared to a conventional DVFS-based DTM technique, satisfying thermal constraints. Younggeun Kim 0001, Jeong In Kim, Seung Hun Choi, Seon Young Kim, Sung Woo Chung |
ISLPED | 5 |
| 2018 | A Survey on Recent OS-Level Energy Management Techniques for Mobile Processing UnitsabstractTo improve mobile experience of users, recent mobile devices have adopted powerful processing units (CPUs and GPUs). Unfortunately, the processing units often consume a considerable amount of energy, which in turn shortens battery life of mobile devices. For energy reduction of the processing units, mobile devices adopt energy management techniques based on software, especially OS (Operating Systems), as well as hardware. In this survey paper, we summarize recent OS-level energy management techniques for mobile processing units. We categorize the energy management techniques into three parts, according to main operations of the summarized techniques: 1) techniques adjusting power states of processing units, 2) techniques exploiting other computing resources, and 3) techniques considering interactions between displays and processing units. We believe this comprehensive survey paper will be a useful guideline for understanding recent OS-level energy management techniques and developing more advanced OS-level techniques for energy-efficient mobile processing units. Younggeun Kim 0001, Joonho Kong, Sung Woo Chung |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Signal strength-aware adaptive offloading for energy efficient mobile devicesabstractTo prolong battery life of mobile devices, applications often exploit offloading techniques which run computations on remote servers. Unfortunately, the existing offloading techniques do not consider the fact that data transmission time and energy consumption of wireless network interfaces exponentially increase when signal strength decreases. In this paper, we propose an adaptive offloading technique that considers signal strength. Our technique estimates gain (reduced computation time and energy of mobile devices) and loss (increased data transmission time and energy of network interfaces) of offloading depending on signal strength. Based on the estimated gain and loss, our technique determines whether it offloads computations to a server or not. In evaluation, our proposed technique improves performance by 30.1% and saves system-wide energy consumption by 25.0%, on average, compared to the conventional offloading technique that does not consider signal strength. Younggeun Kim 0001, Sung Woo Chung |
ISLPED | 2 |
| 2017 | Architecting large-scale SRAM arrays with monolithic 3D integrationabstractIn this paper, we architect large-scale SRAM arrays with monolithic 3D (M3D) integration technology. We introduce M3D-based SRAM arrays with three different ways of integration: M3D-R (vertical routing-only), M3D-VBL (vertical bitline), and M3D-VWL (vertical wordline). We also apply M3D-based SRAM arrays to last-level caches: tag arrays for eDRAM LLCs and data arrays for SRAM LLCs. The proposed LLCs with M3D-based SRAM arrays lead to better performance and lower energy by 0.02%∼1.7% and 49.1%∼79.1%, respectively, compared to that with TSV-based 3D SRAM arrays. Joonho Kong, Young-Ho Gong, Sung Woo Chung |
ISLPED | 3 |
| 2017 | Enhancing Energy Efficiency of Multimedia Applications in Heterogeneous Mobile Multi-Core ProcessorsabstractRecent smart devices have adopted heterogeneous multi-core processors which have high-performance big cores and low-power small cores. Unfortunately, the conventional task scheduler for heterogeneous multi-core processors does not provide appropriate amount of CPU resources for multimedia applications (whose QoS is important to users), resulting in energy waste; it often executes multimedia applications and non-multimedia applications on the same core. In this paper, we propose an advanced task scheduler for heterogeneous multi-core processors, which provides appropriate amount of CPU resources for multimedia applications. Our proposed task scheduler isolates multimedia applications from non-multimedia applications at runtime, exploiting the fact that multimedia applications have a specific thread for video/audio playback (to play video/audio, a multimedia application should use a function that generates the specific thread). Since multimedia applications usually require a smaller amount of CPU resources than non-multimedia applications due to dedicated hardware decoders, our proposed task scheduler allocates the former to the small cores and the latter to the big cores. In our experiments on an Android-based development board, our proposed task scheduler saves system-wide (not just CPU) energy consumption by 8.9 percent, on average, compared to the conventional task scheduler, preserving QoS of multimedia applications. In addition, it improves performance of non-multimedia applications by 13.7 percent, on average, compared to the conventional task scheduler. Younggeun Kim 0001, Minyong Kim, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2016 | Exploiting Refresh Effect of DRAM Read Operations: A Practical Approach to Low-Power RefreshabstractDynamic random access memory (DRAM) requires periodic refresh operations to retain its data. In practice, DRAM retention times are normally distributed from 64 ms to several seconds. However, the conventional refresh method uses 64 ms as the refresh interval, since it applies the same refresh interval to all DRAM rows. Thus, the conventional refresh method results in unnecessary refresh operations (eventually, energy waste) to the DRAM rows whose retention times are longer than 64 ms. In this paper, we propose a practical refresh scheme that exploits refresh effect of DRAM read operations to reduce refresh overhead. Our proposed scheme applies a refresh interval longer than the conventional refresh interval (64 ms) to the DRAM chip. In this case, weak DRAM rows (DRAM rows whose retention times are shorter than the refresh interval of the DRAM chip) cannot retain their data. In order to retain the data stored in the weak DRAM rows, the memory controller issues read operations to the weak DRAM rows every required refresh interval for the weak DRAM rows. Our evaluation results show that our proposed scheme with 192 ms refresh interval reduces average refresh energy consumption up to 66.0 percent, which in turn reduces average DRAM energy consumption up to 31.8 percent, compared to the conventional refresh method (64 ms). Our proposed scheme requires no modification to internal DRAM chip structures, but it only adds a small weak row buffer (the buffer for the weak row information) to the memory controller, which has a negligible area overhead. Young-Ho Gong, Sung Woo Chung |
IEEE Trans. Computers | 2 |
| 2015 | M-DTM: migration-based dynamic thermal management for heterogeneous mobile multi-core processors
Younggeun Kim 0001, Minyong Kim, Jae Min Kim, Sung Woo Chung |
DATE | 4 |
| 2015 | Stabilizing CPU Frequency and Voltage for Temperature-Aware DVFS in Mobile DevicesabstractRecent mobile devices adopt high-performance processors to support various functions. As a side effect, higher performance inevitably leads to power density increase, eventually resulting in thermal problems. In order to alleviate the thermal problems, off-the-shelf mobile devices rely on dynamic voltage-frequency scaling (DVFS)-based dynamic thermal management (DTM) schemes. Unfortunately, in the DVFS-based DTM schemes, an excessive number of DTM operations worsen not only performance but also power efficiency. In this paper, we propose a temperature-aware DVFS scheme for Android-based mobile devices to optimize power or performance depending on the option. We evaluate our scheme in the off-the-shelf mobile device. Our evaluation results show that our scheme saves energy consumption by 12.7%, on average, when we use the power optimizing option. Our scheme also enhances the performance by 6.3%, on average, by using the performance optimizing scheme, still reducing the energy consumption by 6.7%. Jae Min Kim, Younggeun Kim 0001, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2015 | An Energy-Efficient Last-Level Cache Architecture for Process Variation-Tolerant 3D MicroprocessorsabstractAs process technologies evolves, tackling process variation problems is becoming more challenging in 3D (i.e., die-stacked) microprocessors. Process variation adversely affects performance, power, and reliability of the 3D microprocessors, which in turn results in yield losses. In particular, last-level caches (LLCs: L2 or L3 caches) are known as the most vulnerable component to process variation in 3D microprocessors. In this paper, we propose a novel cache architecture that exploits narrow-width values for yield improvement of LLCs (in this paper, L2 caches) in 3D microprocessors. Our proposed architecture disables faulty cache subparts and turns on only the portions that store meaningful data in the cache arrays, which results in high energy-efficiency as well as high cache yield. In an energy-/performance-efficient manner, our proposed architecture significantly recovers not only SRAM cell failure-induced yield losses but also leakage-induced yield losses. Joonho Kong, Farinaz Koushanfar, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2014 | Leveraging Process Variation for Performance and Energy: In the Perspective of OverclockingabstractProcess variation is one of the most important factors to be considered in recent microprocessor design, since it negatively affects performance, power, and yield of microprocessors. However, by leveraging process variation, overclocking techniques can improve performance. As microprocessors have substantial clock cycle time margin for yield, there is enough room for performance improvement by overclocking techniques. In this paper, we adopt the F-overclocking technique, which increases clock frequency without changing supply voltage. Our experimental results show that the F-overclocking technique significantly improves performance as well as energy consumption. In addition, the F-overclocking technique is superior to the conventional overclocking technique which increases clock frequency and supply voltage together in the perspective of energy efficiency and reliability, showing similar performance improvement. Furthermore, we propose an adaptive overclocking controller which dynamically applies the F-overclocking technique based on the application characteristics. By adopting our adaptive overclocking controller, we further minimize the reliability loss caused by the F-overclocking technique. Hyung Beom Jang, Junhee Lee 0004, Joonho Kong, Taeweon Suh, Sung Woo Chung |
IEEE Trans. Computers | 5 |
| 2013 | Exploiting Application/System-Dependent Ambient Temperature for Accurate Microarchitectural SimulationabstractIn the early design stage of processors, Dynamic Thermal Management (DTM) schemes should be evaluated to avoid excessively high temperature, while minimizing performance overhead. In this paper, we show that conventional thermal simulations using the fixed ambient temperature may lead to the wrong conclusions in terms of temperature, performance, reliability, and leakage power. Though ambient temperature converges to a steady-state value after hundreds of seconds when we run SPEC CPU2000 benchmark suite, the steady-state ambient temperature is significantly different depending on applications and system configuration. To overcome inaccuracy of conventional thermal simulations, we propose that microarchitectural thermal simulations should exploit application/system-dependent ambient temperature. Our evaluation results reveal that performance, thermal behavior, reliability, and leakage power of the same DTM scheme are different when we use the application/system-dependent ambient temperature instead of the fixed ambient temperature. For accurate simulation results, future microarchitectural thermal researchers are expected to evaluate their proposed DTM schemes based on application/system-dependent ambient temperature. Hyung Beom Jang, Jinhang Choi, Ikroh Yoon, Sung-Soo Lim, Seungwon Shin 0002, Naehyuck Chang, Sung Woo Chung |
IEEE Trans. Computers | 7 |
| 2012 | Exploiting narrow-width values for process variation-tolerant 3-D microprocessorsabstractProcess variation is a challenging problem in 3D microprocessors, since it adversely affects performance, power, and reliability of 3D microprocessors, which in turn results in yield losses. In this paper, we propose a novel architectural scheme that exploits the narrow-width value for yield improvement of last-level caches in 3D microprocessors. In a energy-/performance-efficient manner, our proposed scheme improves cache yield by 58.7% and 17.3% compared to the baseline and the naïve way-reduction scheme (that simply discards faulty cache lines), respectively. Joonho Kong, Sung Woo Chung |
DAC | 2 |
| 2012 | Fine-Grain Voltage Tuned Cache Architecture for Yield Management Under Process VariationsabstractProcess variations cause large fluctuations in performance and power consumption in the manufactured chips, which eventually results in yield losses. In this paper, to mitigate access time failures and excessive leakage in caches, we propose a novel selective wordline boosting mechanism combined with SRAM cell arrays voltage lowering. Based on our evaluation, the proposed approach recovers up to 83.1% of the yield losses. Joonho Kong, Yan Pan 0010, Serkan Ozdemir, Anitha Mohan, Gokhan Memik, Sung Woo Chung |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2010 | Exploiting application-dependent ambient temperature for accurate architectural simulationabstractIn the early stage of processor design, Dynamic Thermal Management (DTM) schemes should be evaluated to avoid excessively high temperature, while minimizing performance overhead as small as possible. In this paper, we show that conventional thermal simulations using fixed ambient temperature may lead to wrong conclusion in terms of performance and temperature; though ambient temperature converges to a steady state after hundreds of seconds, the steady state ambient temperature is significantly different depending on applications. To overcome the inaccuracy of conventional thermal simulations, we propose that architectural thermal simulation should exploit application-dependent ambient temperature. Our evaluation results show that the performance of the same DTM scheme is different, when application-dependent ambient temperature (compared to fixed temperature) is used. For accurate simulation, future architectural thermal researchers are expected to evaluate their proposed DTM schemes, reflecting application- dependent ambient temperature. Hyung Beom Jang, Jinhang Choi, Ikroh Yoon, Sung-Soo Lim, Seungwon Shin 0002, Naehyuck Chang, Sung Woo Chung |
ICCD | 7 |
| 2010 | Architecture/OS Support for Embedded Multi-core SystemsabstractSung Woo Chung, Hsien-Hsin S. Lee, Woo Hyong Lee; Architecture/OS Support for Embedded Multi-core Systems, The Computer Journal, Volume 53, Issue 8, 1 Octo Sung Woo Chung, Hsien-Hsin S. Lee, Woo Hyong Lee |
Comput. J. | 1 |
| 2010 | A fast and simple system performance emulator for enhanced solid state disks: a case study of long read operationsabstractIn this paper, we propose a fast and simple system emulator, called a system performance emulator (SPE), to evaluate long read operations. The SPE estimates how much system-wide performance is enhanced by using a faster solid state disk (SSD). By suspending a CPU for a certain time during direct memory access (DMA) transfer and subtracting this suspended time from the total DMA time, the SPE estimates the improvement in system performance expected from an enhanced SSD prior to its manufacture. We also examine the relation between storage performance and system performance using the SPE. Do Yeun Kim, Chanik Park, Eui-Young Chung, Sung Woo Chung |
J. Zhejiang Univ. Sci. C | 4 |
| 2010 | Load Unbalancing Strategy for Multicore Embedded ProcessorsabstractLoad balancing has been known as an essential feature for enhancing the performance of distributed systems. For embedded systems, however, this is not always true since load balancing leads to lavish power consumption by fully utilizing all the embedded cores even for a small number of tasks. Furthermore, the previously proposed load unbalancing strategies do not concern much about the characteristics of the embedded system's real workload. In this paper, to resolve this problem, we propose a novel load unbalancing strategy based on the task characteristics: periodic and aperiodic. In the proposed strategy, the periodic tasks that are more likely to be executed repeatedly are concentrated on the minimum number of cores, whereas the aperiodic tasks that are not likely to occur again soon are distributed to the maximum number of cores. The experimental results on an ARM11MPCore test chip show that the proposed strategy reduces power consumption and mean waiting time of the aperiodic tasks by up to 26 percent and 82 percent, respectively, compared to the load balancing strategy. As compared to the aggressive load unbalancing strategy, the proposed strategy also reduces mean waiting time of the aperiodic tasks by 92 percent with similar power efficiency. Hyeran Jeon, Woo Hyong Lee, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2010 | Predictive Temperature-Aware DVFSabstractIn this paper, we propose predictive temperature-aware dynamic voltage and frequency scaling (DVFS) using the performance counters that are already embedded in commercial microprocessors. By using the performance counters and simple regression analysis, we can predict the localized temperature and efficiently scale the voltage/frequency. When localized thermal problems that were not detected by thermal sensors are found after layout (or fabrication), the thermal problems can be avoided by the proposed software solution without delaying time-to-market. The evaluation results show that in a Linux-based laptop with the Intel Core2 Duo processor, DVFS using the performance counters performs comparable to DVFS using the thermal sensor. Jong Sung Lee, Kevin Skadron, Sung Woo Chung |
IEEE Trans. Computers | 3 |
| 2010 | On the Thermal Attack in Instruction CachesabstractThe instruction cache has been recognized as one of the least hot units in microprocessors, which leaves the instruction cache largely ignored in on-chip thermal management. Consequently, thermal sensors are not allocated near the instruction cache. However, malicious codes can exploit the deficiency in this empirical design and heat up fine-grain localized hotspots in the instruction cache, which might lead to physical damages. In this paper, we show how instruction caches can be thermally attacked by malicious codes and how simple techniques can be utilized to protect instruction caches from the thermal attack. Joonho Kong, Johnsy K. John, Eui-Young Chung, Sung Woo Chung, Jie S. Hu |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2010 | Energy-Optimal Dynamic Thermal Management: Computation and Cooling Power Co-OptimizationabstractConventional dynamic thermal management (DTM) assumes that the thermal resistance of a heat-sink is a given constant determined at design time. However, the thermal resistance of a common forced-convection heat sink is inversely proportional to the flow rate of the air or coolant at the expense of the cooling power consumption. The die temperature of the silicon devices strongly affects its leakage power consumption and reliability, and it can be changed by adjusting the thermal resistance of the cooling devices. Different from conventional DTM which aims to avoid the thermal emergency, our proposed DTM regards the thermal resistance of a forced-convection heat sink as a control variable, and minimize the total power consumption both for computation and cooling. We control the cooling power consumption together with the microprocessor clock frequency and supply voltage, and track the energy-optimal die temperature. Consequently, we reduce a significant amount of the temperature-dependent leakage power consumption of the microprocessor while spending a bit higher cooling power than conventional DTM, and eventually consume less total power. Experimental results show the proposed DTM saves up to 8.2% of the total energy compared with a baseline DTM approach. Our proposed DTM also enhances the Failures in Time (FIT) up to 80% in terms of the electromigration lifetime reliability. Donghwa Shin, Sung Woo Chung, Eui-Young Chung, Naehyuck Chang |
IEEE Trans. Ind. Informatics | 2 |
| 2009 | Selective wordline voltage boosting for caches to manage yield under process variationsabstractOne of the most important hurdles of technology scaling is process variations, i.e., variations in device characteristics. Process variations cause large fluctuations in performance and power consumption in the manufactured chips. In addition, these fluctuations cause reductions in the chip yields. In this work, we present an analysis of a representative high-performance processor architecture and show that the caches have the highest probability of causing yield losses under process variations. We then propose a novel selective wordline voltage boosting mechanism that aims at reducing the latency of the cache lines that are affected by process variations. We show that our approach can eliminate over 80% of the yield losses under medium level of variations, while incurring less than 1% per-access energy overhead on average and less than 4.5% area overhead. Yan Pan 0010, Joonho Kong, Serkan Ozdemir, Gokhan Memik, Sung Woo Chung |
DAC | 5 |
| 2009 | Exploiting narrow-width values for thermal-aware register file designsabstractLocalized heating-up creates thermal hotspots across the chip, with the integer register file ranked as the hottest unit in high-performance microprocessors. In this paper, we perform a detailed study on the thermal behavior of a low-power value-aware register file (VARF) that is subjected to internal fine-grain hotspots. To further optimize its thermal behavior, we propose and evaluate three thermal-aware control schemes, thermal sensor (TS), access counter (AC), and register-id (ID) based, to balance the access activity and thus the temperature across different partitions in the VARF. The simulation results using SPEC CINT2000 benchmarks show that the register-id controlled VARF (ID-VARF) scheme achieves optimized thermal behavior at minimum cost as compared to the other schemes. We further evaluate the performance impact of the thermal-aware VARF design with the dynamic thermal management (DTM). The experimental results show that the ID-VARF can improve the performance by 26.1% and 7.2% over the conventional register file and the original VARF design, respectively. Shuai Wang 0006, Jie S. Hu, Sotirios G. Ziavras, Sung Woo Chung |
DATE | 4 |
| 2009 | Energy-optimal dynamic thermal management for green computingabstractExisting thermal management systems for microprocessors assume that the thermal resistance of the heat-sink is constant and that the objective of the cooling system is simply to avoid thermal emergencies. But in fact the thermal resistance of the usual forced-convection heat-sink is inversely proportional to the fan speed, and a more rational objective is to minimize the total power consumption of both processor and cooling system. Our new method of dynamic thermal management uses both the fan speed and the voltage/frequency of the microprocessor as control variables. Experiments show that tracking the energy-optimal steady-state temperature can saves up to 17.6% of the overall energy, when compared with a conventional approach that merely avoids over-heating. Donghwa Shin, Naehyuck Chang, Jinhang Choi, Sung Woo Chung, Eui-Young Chung |
ICCAD | 5 |
| 2009 | The impact of liquid cooling on 3D multi-core processorsabstractRecently, 3D integration has been regarded as one of the most promising techniques due to its abilities of reducing global wire lengths and lowering power consumption. However, 3D integrated processors inevitably cause higher power density and lower thermal conductivity, since the closer proximity of heat generating dies makes existing thermal hotspots more severe. Without an efficient cooling method inside the package, 3D integrated processors should suffer severe performance degradation by dynamic thermal management as well as reliability problems. In this paper, we analyze the impact of the liquid cooling on a 3D multi-core processor compared to the conventional air cooling. We also evaluate the leakage power consumption and the lifetime reliability depending on the temperature of each functional unit in the 3D multi-core processor. The simulation results show that the liquid cooling reduces the temperature of the L1 instruction cache (the hottest block in this evaluation) by as much as 45 degrees, resulting in 12.8% leakage reduction, on average, compared to the conventional air cooling. Moreover, the reduced temperature of the L1 instruction cache also improves the reliability of electromigration, stress migration, time-dependent dielectric breakdown, thermal cycling, and negative bias temperature instability significantly. Hyung Beom Jang, Ikroh Yoon, Cheol Hong Kim, Seungwon Shin 0002, Sung Woo Chung |
ICCD | 5 |
| 2008 | Energy-Effective Instruction Fetch Unit for Embedded ProcessorsabstractFor energy-aware embedded and mobile processors, this paper proposes a new energy-effective design of the instruction fetch unit that exploits the fact that the per-access energy consumption decreases as the cache size decreases. Cheol Hong Kim, Intae Hwang, Changhyeon Chae, Daewon Choi, Taejin Jung, Sung Woo Chung |
CCNC | 6 |
| 2008 | On-Demand Solution to Minimize I-Cache Leakage Energy with Maintaining PerformanceabstractThis paper describes a new on-demand wake-up prediction policy for reducing leakage power. The key insight is that branch prediction can be used to selectively wake up only the needed cache line. This achieves better leakage savings than the best prior policies while avoiding the performance overheads of those policies, without needing an extra prediction structure. The proposed policy reduces leakage energy by 92.7 percent with only 0.08 percent performance overhead on average. The branch-prediction-based approach requires an extra pipeline stage for wake up, which adds to the branch misprediction penalty. Fortunately, this cost is mitigated because the extra wake-up stage is overlapped with misprediction recovery. This paper assumes the superdrowsy leakage control technique using reduced supply voltage because it is well suited to the instruction cache's criticality. However, the proposed policy can be also applied to other leakage-saving circuit techniques. Sung Woo Chung, Kevin Skadron |
IEEE Trans. Computers | 1 |
| 2008 | Energy and Performance Optimization of Demand Paging With OneNAND FlashabstractNew fusion memory devices consisting of multiple heterogeneous memory components in a single die or package offer efficient ways to optimize embedded systems in terms of energy, performance, and cost. Samsung Electronics recently announced the OneNAND fusion memory, in which a NAND flash array is integrated with dual SRAM buffers to provide a nor-type I/O interface. OneNAND has the low cost and large capacity of a NAND flash but also permits eXecution-in-Place (XIP) like a nor flash. The deployment of such devices requires careful system-level resource management because of their impact on energy consumption and performance, and existing memory optimization techniques, such as the demand paging used with NAND flash, may no longer be appropriate for systems with a fusion memory. We introduce a new online demand paging scheme that fully exploits the XIP capability of OneNAND flash by classifying pages as load preferred (residing in the on-chip SRAM) and XIP preferred (accessed directly from the OneNAND flash and discarded after use). This achieves, on average, a 26% reduction in energy consumption and a 19% increase in performance, compared with conventional NAND flash demand paging. Yongsoo Joo, Yongseok Choi, Jaehyun Park 0005, Chanik Park, Sung Woo Chung, Eui-Young Chung, Naehyuck Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2006 | Practice and Experience of an Embedded Processor Core Modeling
Gi-Ho Park, Sung Woo Chung, Han-Jong Kim, Jung-Bin Im, Jung-Wook Park, Shin-Dug Kim, Sung-Bae Park |
HPCC | 2 |
| 2006 | A Novel Software Solution for Localized Thermal Problems
Sung Woo Chung, Kevin Skadron |
ISPA | 1 |
| 2005 | An Accurate Architectural Simulator for ARM1136
Hyo-Joong Suh, Sung Woo Chung |
EUC | 2 |
| 2005 | Distance-aware L2 cache organizations for scalable multiprocessor systems
Sung Woo Chung, Hyong-Shik Kim, Chu Shik Jhon |
J. Syst. Archit. | 1 |
| 2003 | Distance-aware L2 Cache Organizations for Scalable Multiprocessor SystemsabstractIn this paper, we suggest an LRU/distance-aware combined second-level (L2) cache for scalable CC-NUMA multiprocessors, which is composed of a traditional LRU cache and an additional cache maintaining the distance information of individual cache blocks. The LRU cache selects a victim using age information, while the distance-aware cache does this using distance information. Both work together to reduce the overall distance effectively upon cache misses by keeping long-distance blocks as well as recently used blocks. It has been observed that the proposed cache outperforms the traditional LRU cache by up to 28% in the execution time. It is also found to perform even better than an LRU cache of twice the size. Sung Woo Chung, Hyong-Shik Kim, Chu Shik Jhon |
DSD | 1 |
| 2001 | Accelerating the Continuous Data in a SCI-Based Multimedia SystemabstractThe SCI (Scalable Coherent Interface) is widely used in high performance clustered systems these days. However, there has not been any consideration for scheduling policies of improving transmission performance of continuous data in a SCI-based multimedia server. The scheduling policy that does not take into account the property of continuous data may hurt the overall performance. We propose two scheduling policies. The first policy is to maintain separate queues in a node in order to provide differentiated transmission services to continuous data and discrete data. The second one is to let continuous data even preempt discrete data that is being transmitted, which may bring more timely service to continuous data. We measured how the continuous data latency is reduced by simulations, which show that 8-70% reduction is expected, on an average, with scheduling policy and observed even better reduction with the second scheduling policy. Sung Woo Chung, Hyong-Shik Kim, Chu Shik Jhon |
ICPADS | 1 |
| 1998 | PANDA: Ring-Based Multiprocessor System Using New Snooping ProtocolabstractThe PANDA is a ring-based Cache Coherent Non-Uniform Memory Access (CC-NUMA) multiprocessor system under implementation at the Seoul National University. Its main goal is to ameliorate the data miss latency by using the unidirectional point-to-point interconnection network. We introduce the PANDA architecture and present a new snooping protocol for this system. We evaluate the performance of the PANDA for a small to medium scale multiprocessor system using analytical models and a program-driven simulator. We compare the proposed system to other alternatives of point-to-point connected machines, such as the Express Ring and full map directory based system. The simulation results show up to 29% performance improvement against the Express Ring. They also show that the PANDA performs no worse than the full map directory based system, which has the additional hardware costs for the directory management. Sung Woo Chung, Chu Shik Jhon |
ICPADS | 1 |