Dong Hyuk Woo

dblp:07/6993 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Memory systems · 40% GPUs and heterogeneous computing · 18% Processor architecture and microarchitecture · 18%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 28 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
3d-stacked memory
0.322015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
An optimized 3D-stacked memory architecture by exploiting excessive, high-density TSV bandwidth · HPCA 2010
Processor architecture and microarchitecture › chip multiprocessor
3d chip multiprocessor
0.212015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
Processor architecture and microarchitecture
chip multiprocessor
0.212015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
GPUs and heterogeneous computing › GPU cache
GPU cache hierarchy
0.212015
GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015
Energy-efficient computing › memory energy efficiency
low-power cache design
0.212015
GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015
Memory systems › cache design
non-inclusive cache
0.212015
GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015
Memory systems
non-volatile memory
0.222010
SAFER: Stuck-At-Fault Error Recovery for Memories · MICRO 2010
Security refresh: prevent malicious wear-out and increase durability for phase-change memory with dynamically randomized address mapping · ISCA 2010
Integrated circuit design › 3d integration
through-silicon via
0.212015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
GPUs and heterogeneous computing
control flow divergence
0.212013
SIMD divergence optimization through intra-warp compaction · ISCA 2013
Processor architecture and microarchitecture
SIMD
0.212013
SIMD divergence optimization through intra-warp compaction · ISCA 2013
Memory systems › hybrid memory
hybrid main memory
0.112012
Hybrid DRAM/PRAM-based main memory for single-chip CPU/GPU · DAC 2012
Memory systems
cache
0.112010
Chameleon: Virtualizing idle acceleration cores of a heterogeneous multicore processor for caching and prefetching · ACM Trans. Archit. Code Optim. 2010
Hardware reliability and fault tolerance
error-correcting codes for memory
0.112010
SAFER: Stuck-At-Fault Error Recovery for Memories · MICRO 2010
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.112010
Chameleon: Virtualizing idle acceleration cores of a heterogeneous multicore processor for caching and prefetching · ACM Trans. Archit. Code Optim. 2010
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.112010
Chameleon: Virtualizing idle acceleration cores of a heterogeneous multicore processor for caching and prefetching · ACM Trans. Archit. Code Optim. 2010
Memory systems › non-volatile memory
phase change memory
0.112010
Security refresh: prevent malicious wear-out and increase durability for phase-change memory with dynamically randomized address mapping · ISCA 2010
Memory systems › cache
prefetching
0.112010
COMPASS: a programmable data prefetcher using idle GPU shaders · ASPLOS 2010
Memory systems › non-volatile memory
resistive memory
0.112010
SAFER: Stuck-At-Fault Error Recovery for Memories · MICRO 2010
Memory systems
software prefetching
0.112010
COMPASS: a programmable data prefetcher using idle GPU shaders · ASPLOS 2010
Hardware reliability and fault tolerance › memory reliability
stuck-at fault recovery
0.112010
SAFER: Stuck-At-Fault Error Recovery for Memories · MICRO 2010
Storage systems › flash and SSD › flash memory management
wear leveling
0.112010
Security refresh: prevent malicious wear-out and increase durability for phase-change memory with dynamically randomized address mapping · ISCA 2010
Performance modeling and evaluation
benchmarking
0.112015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
Energy-efficient computing › power management › memory power management
cache energy reduction
0.112015
GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015
Performance modeling and evaluation › benchmarking
parallel benchmark
0.112015
Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory) · IEEE Trans. Computers 2015
Parallel and multicore computing › data parallelism
data-parallel applications
0.012013
SIMD divergence optimization through intra-warp compaction · ISCA 2013
Hardware security and side channels
memory security
0.012010
Security refresh: prevent malicious wear-out and increase durability for phase-change memory with dynamically randomized address mapping · ISCA 2010
Memory systems › memory hierarchy
cache hierarchy
0.012010
An optimized 3D-stacked memory architecture by exploiting excessive, high-density TSV bandwidth · HPCA 2010
Memory systems › cache › prefetching
data prefetching
0.012010
Chameleon: Virtualizing idle acceleration cores of a heterogeneous multicore processor for caching and prefetching · ACM Trans. Archit. Code Optim. 2010

Methods — techniques the papers use, named apart from their topics

simulation · 0.3write buffer management · 0.1hot data management · 0.1access scheduling · 0.1wear leveling · 0.1vertical l2 fetch/write-back network · 0.1hamming coding · 0.1false sharing management · 0.1error correcting pointers · 0.1compute shader · 0.1
YearPublicationVenuePosition
2015 Design and Analysis of 3D-MAPS (3D Massively Parallel Processor with Stacked Memory)
abstract
This paper describes the architecture, design, analysis, and simulation and measurement results of the 3D-MAPS (3D massively parallel processor with stacked memory) chip built with a 1.5 V, 130 nm process technology and a two-tier 3D stacking technology using 1.2$\micro\hbox{m}$-diameter, 6$\micro \hbox{m}$-height through-silicon vias (TSVs) and$3.4\nbsp\micro\hbox{m}$-diameter face-to-face bond pads. 3D-MAPS consists of a core tier containing 64 cores and a memory tier containing 64 memory blocks. Each core communicates with its dedicated 4KB SRAM block using face-to-face bond pads, which provide negligible data transfer delay between the core and the memory tiers. The maximum operating frequency is 277 MHz and the maximum memory bandwidth is 70.9 GB/s at 277 MHz. The peak measured memory bandwidth usage is 63.8 GB/s and the peak measured power is approximately 4 W based on eight parallel benchmarks.
Dae Hyun Kim 0004, Krit Athikulwongse, Michael B. Healy, Mohammad M. Hossain, Moongon Jung, Ilya Khorosh, Gokul Kumar, Young-Joon Lee, Dean L. Lewis, Tzu-Wei Lin, Chang Liu 0034, Shreepad Panth, Mohit Pathak, Minzhen Ren, Guanhao Shen, Taigon Song, Dong Hyuk Woo, Xin Zhao 0001, Joungho Kim, Ho Choi, Gabriel H. Loh, Hsien-Hsin S. Lee, Sung Kyu Lim
IEEE Trans. Computers17
2015 GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs
abstract
As various graphics processing unit architectures are deployed across broad computing spectrum from a hand-held or embedded device to a high-performance computing server, OpenCL becomes the de facto standard programming environment for general-purpose computing on graphics processing units. Unlike its CPU counterpart, OpenCL has several distinct features such as its disciplined memory model, which is partially inherited from conventional 3D graphics programming models. On the other hand, due to ever increasing memory bandwidth pressure and low power requirement, the capacity of on-chip caches in GPUs keeps increasing overtime. Given such trends, we believe that we have interesting programming model/architecture co-optimization opportunities, in particular, how to energy-efficiently utilize large on-chip caches for GPUs. In this paper, as a showcase, we study the characteristics of the OpenCL memory model and propose a technique called GPU Region-aware energy-efficient non-inclusive cache hierarchy, or GREEN cache hierarchy. With the GREEN cache, our simulation results show that we can save 56 percent of dynamic energy in the L1 cache, 39 percent of dynamic energy in the L2 cache, and 50 percent of leakage energy in the L2 cache with practically no performance degradation and off-chip access increases.
Jaekyu Lee, Dong Hyuk Woo, Hyesoon Kim, Mani Azimi
IEEE Trans. Computers2
2013 SIMD divergence optimization through intra-warp compaction
abstract
SIMD execution units in GPUs are increasingly used for high performance and energy efficient acceleration of general purpose applications. However, SIMD control flow divergence effects can result in reduced execution efficiency in a class of GPGPU applications, classified as divergent applications. Improving SIMD efficiency, therefore, has the potential to bring significant performance and energy benefits to a wide range of such data parallel applications.
Aniruddha S. Vaidya, Anahita Shayesteh, Dong Hyuk Woo, Roy Saharoy, Mani Azimi
ISCA3
2013 Pragmatic Integration of an SRAM Row Cache in Heterogeneous 3-D DRAM Architecture Using TSV
abstract
As scaling DRAM cells becomes more challenging and energy-efficient DRAM chips are in high demand, the DRAM industry has started to undertake an alternative approach to address these looming issues-that is, to vertically stack DRAM dies with through-silicon-vias (TSVs) using 3-D-IC technology. Furthermore, this emerging integration technology also makes heterogeneous die stacking in one DRAM package possible. Such a heterogeneous DRAM chip provides a unique, promising opportunity for computer architects to contemplate a new memory hierarchy for future system design. In this paper, we study how to design such a heterogeneous DRAM chip for improving both performance and energy efficiency. In particular, we found that, if we want to design an SRAM row cache in a DRAM chip, simple stacking alone cannot address the majority of traditional SRAM row cache design issues. In this paper, to address these issues, we propose a novel floorplan and several architectural techniques that fully exploit the benefits of 3-D stacking technology. Our multi-core simulation results with memory-intensive applications suggest that, by tightly integrating a small row cache with its corresponding DRAM array, we can improve performance by 30% while saving dynamic energy by 31%.
Dong Hyuk Woo, Nak Hee Seong, Hsien-Hsin S. Lee
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Acceleration of bulk memory operations in a heterogeneous multicore architecture
abstract
In this paper, we present a novel approach of using the integrated GPU to accelerate conventional operations that are normally performed by the CPUs, the bulk memory operations, such as memcpy or memset. Offloading the bulk memory operations to the GPU has many advantages, i) the throughput driven GPU outperforms the CPU on the bulk memory operations; ii) for on-die GPU with unified cache between the GPU and the CPU, the GPU private caches can be leveraged by the CPU for storing moved data and reducing the CPU cache bottleneck; iii) with additional lightweight hardware, asynchronous offload can be supported as well; and iv) different from the prior arts using dedicated hardware copy engines (e.g., DMA), our approach leverages the exiting GPU hardware resources as much as possible. The performance results based on our solution showed that offloaded bulk memory operations outperform CPU up to 4.3 times in micro benchmarks while still using less resources. Using eight real world applications and a cycle based full system simulation environment, the results showed 30% speedup for five, more than 20% speedup for two of the eight applications.
Jong-Hyuk Lee, Ziyi Liu 0002, Xiaonan Tian, Dong Hyuk Woo, Larry Shi, Dainis Boumber, Yonghong Yan 0001, Kyeong-An Kwon
PACT4
2012 Hybrid DRAM/PRAM-based main memory for single-chip CPU/GPU
abstract
Single-chip CPU/GPU architecture is being adopted in high-end (embedded) systems, e.g., smartphones and tablet PCs. Main memory subsystem is expected to consist of hybrid DRAM and phase-change RAM (PRAM) due to the difficulties in DRAM scaling. In this work, we address the performance optimization of the hybrid DRAM/PRAM main memory for single chip CPU/GPU. Based on the tight requirements of low latency from CPU and the relative tolerance to long latency from GPU, DRAM is first allocated to CPU while PRAM with longer write latency is allocated to GPU. Then, in order to improve the write performance of GPU traffic, we propose (1) an in-DRAM write buffer to accommodate GPU write traffics, (2) dynamic hot data management to improve the efficiency of write buffer, (3) runtime-adaptive adjustment of write buffer size to meet the given CPU performance bound, and (4) CPU-aware DRAM access scheduling to give low latency to CPU traffics. The experiments show that the proposed method gives 1.02~44.2 times performance improvement in GPU performance with modest (negligible) CPU performance overhead (when compute-intensive CPU programs run).
Dongki Kim, Sungkwang Lee, Jaewoong Chung, Daehyun Kim 0001, Dong Hyuk Woo, Sungjoo Yoo, Sunggu Lee
DAC5
2010 COMPASS: a programmable data prefetcher using idle GPU shaders
abstract
A traditional fixed-function graphics accelerator has evolved into a programmable general-purpose graphics processing unit over the last few years. These powerful computing cores are mainly used for accelerating graphics applications or enabling low-cost scientific computing. To further reduce the cost and form factor, an emerging trend is to integrate GPU along with the memory controllers onto the same die with the processor cores. However, given such a system-on-chip, the GPU, while occupying a substantial part of the silicon, will sit idle and contribute nothing to the overall system performance when running non-graphics workloads or applications lack of data-level parallelism. In this paper, we propose COMPASS, a compute shader-assisted data prefetching scheme, to leverage the GPU resource for improving single-threaded performance on an integrated system. By harnessing the GPU shader cores with very lightweight architectural support, COMPASS can emulate the functionality of a hardware-based prefetcher using the idle GPU and successfully improve the memory performance of single-thread applications. Moreover, thanks to its flexibility and programmability, one can implement the best performing prefetch scheme to improve each specific application as demonstrated in this paper. With COMPASS, we envision that a future application vendor can provide a custom-designed COMPASS shader bundled with its software to be loaded at runtime to optimize the performance. Our simulation results show that COMPASS can improve the single-thread performance of memory-intensive applications by 68% on average.
Dong Hyuk Woo, Hsien-Hsin S. Lee
ASPLOS1
2010 An optimized 3D-stacked memory architecture by exploiting excessive, high-density TSV bandwidth
abstract
Memory bandwidth has become a major performance bottleneck as more and more cores are integrated onto a single die, demanding more and more data from the system memory. Several prior studies have demonstrated that this memory bandwidth problem can be addressed by employing a 3D-stacked memory architecture, which provides a wide, high frequency memory-bus interface. Although previous 3D proposals already provide as much bandwidth as a traditional L2 cache can consume, the dense through-silicon-vias (TSVs) of 3D chip stacks can provide still more bandwidth. In this paper, we contest that we need to re-architect our memory hierarchy, including the L2 cache and DRAM interface, so that it can take full advantage of this massive bandwidth. Our technique, SMART-3D, is a new 3D-stacked memory architecture with a vertical L2 fetch/write-back network using a large array of TSVs. Simply stated, we leverage the TSV bandwidth to hide latency behind very large data transfers. We analyze the design trade-offs for the DRAM arrays, careful enough to avoid compromising the DRAM density because of TSV placement. Moreover, we propose an efficient mechanism to manage the false sharing problem when implementing SMART-3D in a multi-socket system. For single-threaded memory-intensive applications, the SMART-3D architecture achieves speedups from 1.53 to 2.14 over planar designs and from 1.27 to 1.72 over prior 3D designs. We achieve similar speedups for multi-program and multi-threaded workloads on multi-core and multi-socket processors. Furthermore, SMART-3D can even lower the energy consumption in the L2 cache and 3D DRAM for it reduces the total number of row buffer misses.
Dong Hyuk Woo, Nak Hee Seong, Dean L. Lewis, Hsien-Hsin S. Lee
HPCA1
2010 Security refresh: prevent malicious wear-out and increase durability for phase-change memory with dynamically randomized address mapping
abstract
Phase change memory (PCM) is an emerging memory technology for future computing systems. Compared to other non-volatile memory alternatives, PCM is more matured to production, and has a faster read latency and potentially higher storage density. The main roadblock precluding PCM from being used, in particular, in the main memory hierarchy, is its limited write endurance. To address this issue, recent studies proposed to either reduce PCM's write frequency or use wear-leveling to evenly distribute writes. Although these techniques can extend the lifetime of PCM, most of them will not prevent deliberately designed malicious codes from wearing it out quickly. Furthermore, all the prior techniques did not consider the circumstances of a compromised OS and its security implication to the overall PCM design. A compromised OS will allow adversaries to manipulate processes and exploit side channels to accelerate wear-out.
Nak Hee Seong, Dong Hyuk Woo, Hsien-Hsin S. Lee
ISCA2
2010 SAFER: Stuck-At-Fault Error Recovery for Memories
abstract
As technology scaling poses a threat to DRAM scaling due to physical limitations such as limited charge, alternative memory technologies including several emerging non-volatile memories are being explored as possible DRAM replacements. One main roadblock for wider adoption of these new memories is the limited write endurance, which leads to wear-out related permanent failures. Furthermore, technology scaling increases the variation in cell lifetime resulting in early failures of many cells. Existing error correcting techniques are primarily devised for recovering from transient faults and are not suitable for recovering from permanent stuck-at faults, which tend to increase gradually with repeated write cycles. In this paper, we propose SAFER, a novel hardware-efficient multi-bit stuck-at fault error recovery scheme for resistive memories, which can function in conjunction with existing wear-leveling techniques. SAFER exploits the key attribute that a failed cell with a stuck-at value is still readable, making it possible to continue to use the failed cell to store data, thereby reducing the hardware overhead for error recovery. SAFER partitions a data block dynamically while ensuring that there is at most one fail bit per partition and uses single error correction techniques per partition for fail recovery. SAFER increases the number of recoverable fails and achieves better lifetime improvement with smaller hardware overhead relative to recently proposed Error Correcting Pointers and even ideal hamming coding scheme.
Nak Hee Seong, Dong Hyuk Woo, Vijayalakshmi Srinivasan, Jude A. Rivers, Hsien-Hsin S. Lee
MICRO2
2010 Chameleon: Virtualizing idle acceleration cores of a heterogeneous multicore processor for caching and prefetching
abstract
Heterogeneous multicore processors have emerged as an energy- and area-efficient architectural solution to improving performance for domain-specific applications such as those with a plethora of data-level parallelism. These processors typically contain a large number of small, compute-centric cores for acceleration while keeping one or two high-performance ILP cores on the die to guarantee single-thread performance. Although a major portion of the transistors are occupied by the acceleration cores, these resources will sit idle when running unparallelized legacy codes or the sequential part of an application. To address this underutilization issue, in this article, we introduce Chameleon, a flexible heterogeneous multicore architecture to virtualize these resources for enhancing memory performance when running sequential programs. The Chameleon architecture can dynamically virtualize the idle acceleration cores into a last-level cache, a data prefetcher, or a hybrid between these two techniques. In addition, Chameleon can operate in an adaptive mode that dynamically configures the acceleration cores between the hybrid mode and the prefetch-only mode by monitoring the effectiveness of the Chameleon cache mode. In our evaluation with SPEC2006 benchmark suite, different levels of performance improvements were achieved in different modes for different applications. In the case of the adaptive mode, Chameleon improves the performance of SPECint06 and SPECfp06 by 31% and 15%, on average. When considering only memory-intensive applications, Chameleon improves the system performance by 50% and 26% for SPECint06 and SPECfp06, respectively.
Dong Hyuk Woo, Joshua B. Fryman, Allan D. Knies, Hsien-Hsin S. Lee
ACM Trans. Archit. Code Optim.1
2006 Reducing energy of virtual cache synonym lookup using bloom filters
abstract
Virtual caches are employed as L1 caches of both high performance and embedded processors to meet their short latency requirements. However, they also introduce the synonym problem where the same physical cache line can be present at multiple locations in the cache due to their distinct virtual addresses, leading to potential data consistency issues. To guarantee correctness, common hardware solutions either perform serial lookups for all possible synonym locations in the L1 consuming additional energy or employ a reverse map in the L2 cache that incurs a large area overhead. Such preventive mechanisms are nevertheless indispensable even though synonyms may not always be present during the execution.In this paper, we study the synonym issue using Windows applications workload and propose a technique based on Bloom filters to reduce synonym lookup energy. By tracking the address stream using Bloom filters, we can confidently exclude the addresses that were never observed to eliminate unnecessary synonym lookups, thereby saving energy in the L1 cache. Bloom filters have a very small area overhead making our implementation a feasible and attractive solution for synonym detection. Our results show that synonyms in these applications actually constitutes less than 0.1% of the total cache misses. By applying our technique, the dynamic energy consumed in L1 data cache can be reduced up to 32.5%. When taking leakage energy into account, the savings is up to 27.6%.
Dong Hyuk Woo, Mrinmoy Ghosh, Emre Ozer 0001, Stuart Biles, Hsien-Hsin S. Lee
CASES1