EDBT 2026 Demo / reviewers in the wild / expert
William J. Song
dblp:131/9759
· DBLP profile ↗
20ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-9170-5986ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache CompressionabstractThe growing sequence length of large language models poses significant challenges for keyvalue (KV) caches.Existing state-of-the-art cache eviction methods primarily analyze the inference behavior of attention heads in successful retrieval-reasoning cases, often overlooking diverse behaviors in failure cases, such as bias and distraction.This oversight limits the potential to leverage heterogeneous head behaviors for improved eviction performance.Inspired by the confusion matrix, we introduce an Attention Behavior Matrix to comprehensively analyze attention head behaviors in both success and failure scenarios.By maximizing the signal-to-noise ratio -strengthening valid reasoning pathways in success cases while inhibiting noise from bias and distraction in failure cases -we propose REtrieval-reAsoning and Logic-constructed (REAL) KV cache eviction, the first method to leverage multi-behavior analysis.Comprehensive evaluations show that REAL achieves remarkable performance across various models and benchmarks; notably, on LongBench v2, it achieves comparable accuracy to the strongest baseline, HeadKV-R2, while requiring 32x less space (Figure 1).By offering a novel perspective on behavior analysis, we pave the way for a shift from success-only to comprehensive, failureaware methods in long-context modeling.Our code is available at https://github.com/ yonseicasl/REAL.Question: Which operating system comes installed on the cheaper device between black laptop and silver laptop?Needle: …… the Silver Laptop costs $1200 ……the Black Laptop…… costs $800…… comes installed with Linux…… The Red Tablet comes installed with Android…… 0 T h e 1 B l a c k 2 L a p t o p 3 c o m e s 4 i n s t a l l e d 5 w i t h 6 L i n u x 7 , 8 w h i c h 9 i s 1 0 t h e 1 1 c h e a p e r 1 2 d e v i c e 1 3 b e t w e e n 1 4 t h e 1 5 t w o 1 6 . 1 7 T h e 1 8 B l a c k 1 9 L a p t o p 2 0 c o s t s 2 1 $ 2 2 8 0 0 2 3 , 2 4 w h e r e a s 2 5 t h e 2 6 S i l v e r 2 7 L a p t o p 2 8 c o s t s 2 9 $ 3 0 1 2 0 3 1 0 3 2 .3 3 T h e r e f o r e 3 4 , 3 5 t h e 3 6 B l a c k 3 7 L a p t o p 3 8 i s 3 9 t h e 4 0 c h e a p e r 4 1 d e v i c e 4 2 .4 3 < | e o t _ i d | > Generation Step (Generated Token) 10 20 30 40 50 Attention Ratio (%) +6.5% +9.1% +0.2% +8. Xike Xie, William J. Song |
ACL (1) | 4 |
| 2026 | NPUWattch: ML-Based Power, Area, and Timing Modeling for Neural AcceleratorsabstractPre-silicon modeling tools for characterizing power, area, and timing (PAT) have enabled numerous architectural studies, but traditional analytical and table-based models begin to exhibit limitations in their applicability as architectural design complexity increases and process technology scales below 5 nm with the emergence of advanced transistors. Previous modeling techniques typically assume static scaling factors across different designs and technology nodes, derived from small circuit design benchmarks using old processes. Consequently, they do not reflect complex design variability and nonlinear projection to advanced technology nodes. Moreover, reference logic and SRAM implementations serving as the baseline for design and technology scaling are often created using different technologies and design rules, leading to significant estimation inaccuracies that distort the relative contributions of individual components. To address these challenges, this paper introduces NPUWattch, a machine learning-based PAT modeling framework for neural accelerators. It leverages neural network regression models to learn complex nonlinear relationships in technology and design scaling based on diverse post-layout logic and SRAM design datasets formulated using unified technology libraries. To this end, we developed technology libraries from 65 nm to 2 nm, constructed and validated diverse logic and SRAM datasets, and trained neural network models using an adaptive loss function to reinforce underrepresented regions of the design space. NPUWattch is validated against the post-layout results of numerous open-source neural accelerators, and evaluation results demonstrate that NPUWattch outperforms existing tools with an average estimation error of 2.7%, offering reliable and accurate PAT estimation. Minkwan Kim, Chanho Park 0004, Hanmok Park, Taigon Song, William J. Song |
HPCA | 7 |
| 2025 | RoTA: Rotational Torus Accelerator for Wear Leveling of Neural Processing ElementsabstractThis paper introduces a reliability-aware neural accelerator design with a wear-leveling solution that balances the utilization of processing elements (PEs). Neural accelerators deploy many PEs to exploit data-level parallelism, but their designs and operations have focused mostly on performance and energy efficiency metrics. Directional dataflows in PE arrays and dimensional misalignment with variable-sized neural layers cause the underutilization of PEs, which is biased to PE locations and gradually accumulated over time. Consequently, the accelerators experience severe usage imbalance between PEs. To resolve the problem, this paper proposes a rotational torus accelerator (RoTA) with an optimized wear-leveling scheme that shuffles PE utilization spaces to eliminate PE usage imbalance. Evaluation results show that RoTA improves lifetime reliability by 1.69x. Taesoo Lim, Jingu Park, Bogil Kim, William J. Song |
DATE | 5 |
| 2025 | Graphite: Hardware-Aware GNN Reshaping for Acceleration With GPU Tensor CoresabstractGraph neural networks (GNNs) have emerged as powerful tools for addressing non-euclidean problems. GNNs operate through two key execution phases: i) aggregation and ii) combination. In the aggregation phase, the feature data of neighboring graph nodes are gathered, which is expressed as sparse-dense matrix multiplication (SpMM) between an adjacency matrix and a feature embedding table. The combination phase takes the aggregated feature embedding as input to a neural network model with learnable weights. Typically, the adjacency matrix is extremely sparse due to inherent graph structures, making the aggregation phase a significant bottleneck in GNN computations. This paper introducesGraphite, a GNN acceleration framework to overcome the challenge of SpMM operations and enable graphics processing units (GPUs) to exploit massive thread-level parallelism more efficiently via existing dense acceleration units (i.e., tensor cores). To that end, Graphite employs three techniques for GNN acceleration. First,hardware-aware sparse graph reshaping (HAS)rearranges graph structures to replace sparse operations with dense computations, enabling hardware acceleration through GPU tensor cores. Additionally,balanced thread block scheduling (BTS)distributes sparse thread blocks evenly across streaming multiprocessors in GPUs, andzero-aware warp skipping (ZAWS)eliminates ineffective threads that operate on meaningless zeros. Experimental results show that Graphite achieves an average compression rate of 84.1% for adjacency matrices using HAS. Combined with BTS and ZAWS, Graphite delivers an average 1.55x speedup over the conventional SpMM-based GNN computation method. Taesoo Lim, William J. Song |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Nona: Accurate Power Prediction Model Using Neural NetworksabstractThis paper proposes a neural-network-based power model, Nona, that accurately predicts the power consumption of heterogeneous CPUs on a commercial mobile device. With aggressive on-device power management in action, it becomes increasingly challenging to make accurate power predictions for diverse applications. To overcome the limitations of the existing power models based on linear regression, Nona uses a lightweight neural network with a small number of performance monitoring counters (PMCs) chosen from a system analysis and a loss function designed for power prediction. Experiments on Google Pixel 6 show that Nona has a 3.4% average prediction error, improving on prior work by 2.6x. HoSun Choi, Chanho Park 0004, Euijun Kim, William J. Song |
DAC | 4 |
| 2024 | Genie Cache: Non-Blocking Miss Handling and Replacement in Page-Table-Based DRAM CacheabstractThis paper presents Genie Cache that enables non-blocking miss-handling and replacement in a page-table-based DRAM cache (DC). Various DC designs have been proposed to meet the growing bandwidth demand of emerging memory-bound applications. The related literature can be categorized into hardware-based (HW-based) and page-table-based (PT-based) schemes based on their tag storage methods. HW-based designs store DC metadata (e.g., tags) in on-package DRAM for scalability but use extra bandwidth and energy for metadata access. PT-based schemes store DC tags in page table entries (PTEs), enabling virtual-to-cache address translations using the existing memory management units (MMUs) without the DC bandwidth overhead. However, their miss-handling and eviction mechanisms relying on operating systems (OS) incur nontrivial latency overhead. To minimize the OS intervention, Genie Cache implements non-blocking miss handling and replacement using a hardware unit called DRAM cache management unit (DCMU) and a novel pre-write back mechanism. In Genie Cache, DC misses detected by MMUs are forwarded to the DCMU, which handles the misses by allocating page frames and updating PTEs without calling OS routines. When the PT-based DRAM cache runs low on free pages, an eviction routine is called to flush TLBs and evict a batch of cached pages to avoid frequent TLB shootdowns. Since writing back many dirty pages in a blocking manner causes substantial application stall cycles, Genie Cache proactively writes dirty pages back to off-package memory, allowing the eviction routine to simply evict cleaned pages. Experimental results show that Genie Cache achieves 51.3% speedup over the state-of-the-art PT-based design via non-blocking miss handling and replacement. Youngin Kim 0003, William J. Song |
MICRO | 2 |
| 2023 | NOMAD: Enabling Non-blocking OS-managed DRAM Cache via Tag-Data DecouplingabstractThis paper introduces a DRAM cache architecture that provides near-ideal access time and non-blocking miss handling. Previous DRAM cache (DC) designs are classified into two categories, HW-based and OS-managed schemes. Hardware-based designs implement non-blocking caches that can handle multiple DC misses using MSHRs, but they have drawbacks in metadata management since storing tags in on-package DRAM significantly increases the effective cycle time of DC accesses. In contrast, OS-managed schemes utilize PTEs for storing tags and caching them in TLBs, which can achieve ideal DC access time. However, they implement blocking caches that stall application threads on misses until cache fills are completed. To overcome the limitations of both HW-based and OS-managed schemes, this paper introduces a DRAM cache architecture named Non-blocking OS-managed DRAM cache (NOMAD). Unlike conventional caches that guarantee the presence of data on tag hits, NOMAD decouples tag and data management to enable non-blocking miss handling in an OS-managed DRAM cache. The front-end OS routines of NOMAD manage DC tags using PTEs and TLBs, and its back-end hardware handles data management in the DRAM cache. On a DC miss, the OS updates a tag, offloads a cache-fill command to the back-end, and immediately resumes an application thread without waiting for the cache fill to complete. Instead, the back-end hardware handles the cache fill without blocking the application thread. By decoupling tag and data management in NOMAD, a tag hit does not necessarily guarantee the presence of data in the DRAM cache. The back-end traces which DC lines are still in transfers and checks if the demanded part of a cache line has been transferred yet for every DC access. Notably, this back-end procedure does not require an OS intervention, thereby implementing a non-blocking DRAM cache. Experiment results show that NOMAD reduces application stall cycles by 76.1% and improves IPC by 16.7% over a state-of-the-art OS-managed scheme. Youngin Kim 0003, William J. Song |
HPCA | 3 |
| 2023 | SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUsabstractThis paper presents an address translation scheme in GPUs named SnakeByte that can dynamically manage variable-sized pages and maximize TLB reach by recursively merging contiguous pages. Memory virtualization has become an integral part of GPUs to enhance programmability and memory management efficiency. However, conventional memory virtualization methods using multi-level page tables and caching them in TLBs are insufficient to provide GPUs with enough address translation coverage for the massive volume of data. SnakeByte implements a hardware-based address translation mechanism that recursively merges contiguous pages into larger page groups and effectively extends TLB coverage. SnakeByte allows multiple equal-sized pages coalescing into a page table entry (PTE). It records the validity of pages to be merged using a bit vector, and few bits are annexed to indicate the size of merged pages. If all pages covered by the PTE are allocated with contiguity, the PTE is promoted to be further coalesced into a larger page group. The recursive coalescence of contiguous pages enables SnakeByte to handle variable-sized page groups with the exponentially increasing TLB reach. Associated with a contiguity-aware memory allocator, SnakeByte can consolidate vastly contiguous address spaces into a few TLB entries. Consequently, it significantly reduces TLB misses for large working sets in GPUs and achieves substantial performance improvements. Experiment results show that SnakeByte decreases the number of page table walks by 6.5x and enhances the GPU performance by 2.0x on average over the conventional paging scheme. Jiwon Lee 0001, Ju Min Lee, Yunho Oh, William J. Song, Won Woo Ro |
HPCA | 4 |
| 2023 | LAS: Locality-Aware Scheduling for GEMM-Accelerated Convolutions in GPUsabstractThis article presents a graphics processing unit (GPU) scheduling scheme that maximizes the exploitation of data locality in deep neural networks (DNNs). Convolution is one of the fundamental operations used in DNNs and accounts for more than 90% of the total execution time. To leverage massive thread-level parallelism (TLP) in a GPU, deeply nested convolution loops are lowered (or unrolled) into large matrix multiplication, which trades memory capacity and bandwidth for TLP augmentation. A large workspace matrix is split into tiles of general matrix multiplication (GEMM) and concurrently executed by many thread blocks. Notably, the workspace is filled with a number of duplicate data that originate from the same sources in the input feature map during the lowering process. However, conventional GPU scheduling is oblivious to data duplication patterns in the workspace, and thread blocks are assigned to streaming multiprocessors (SMs) irrespective of data similarity between GEMM tiles. Such scheduling misses a significant opportunity to exploit data locality manifested in the DNN convolution. This article proposes a GPU scheduling technique calledLocality-Aware Scheduling(LAS) that i) identifies which thread blocks share the largest amount of identical data based on the lowered patterns of a DNN convolution and ii) allocates such thread blocks showing the greatest data similarity to the same SM. In this way, small caches in SMs can efficiently utilize the data locality of the DNN convolution. Experimental results show that LAS with tensor cores achieves 20.1% performance improvements on average with 14.8% increases in L1 cache hit rates. William J. Song |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | NeuroSpector: Systematic Optimization of Dataflow Scheduling in DNN AcceleratorsabstractThis paper presents an optimization framework namedNeuroSpectorthat systematically analyzes the dataflow of deep neural network (DNN) accelerators and rapidly identifies optimal execution methods. The proposed methodology is demonstrated to work effectively with a variety of accelerator architectures and DNN workloads. It has been a baffling challenge to devise scheduling schemes for neural accelerators to maximize energy efficiency and performance. The challenge lies in that hardware specifications associated with multi-dimensional DNN data create an enormous number of possible scheduling options that can be exerted on accelerators. Related work suggested various techniques to solve the challenge encompassing brute-force search of massive solution spaces pruned by user constraints, solving the objective functions of system models, learning-based optimization, etc. However, each suggested technique was devised only for a specific accelerator model. Therefore, we find that they are not adaptively applicable to different accelerators and DNN workloads in that they produce hit-or-miss results with 100.1% greater energy and cycles on average compared to optimal scheduling schemes obtained from fully comprehensive brute-force searches. In contrast, NeuroSpector identifies efficient execution methods for various accelerators and workloads with only 1.5% differences on average to the optimal scheduling solutions. The optimization strategy of NeuroSpector is based on an observation that optimal executions are strongly correlated with minimizing data movements to the lower-level memory hierarchy of accelerators rather than maximizing the utilization of processing elements. Thus, NeuroSpector prioritizes optimizing lower-level components in the accelerator hierarchy, which is proven highly effective for various accelerators and DNN workloads. Chanho Park 0004, Bogil Kim, Sungmin Ryu, William J. Song |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | The Nebula Benchmark Suite: Implications of Lightweight Neural NetworksabstractThis article presents a benchmark suite namedNebulathat implements lightweight neural network benchmarks. Recent neural networks tend to form deeper and sizable networks to enhance accuracy and applicability. However, the massive volume of heavy networks makes them highly challenging to use in conventional research environments such as microarchitecture simulators. We notice that neural network computations are mainly comprised of matrix and vector calculations that repeat on multi-dimensional data encompassing batches, channels, layers, etc. This observation motivates us to develop a variable-sized neural network benchmark suite that provides users with options to select appropriate size of benchmarks for different research purposes or experiment conditions. Inspired by the implementations of well-known benchmarks such as PARSEC and SPLASH suites, Nebula offers various size options from large to small datasets for diverse types of neural networks. The Nebula benchmark suite is comprised of seven representative neural networks built on a C++ framework. The variable-sized benchmarks can be executed i) with acceleration libraries (e.g., BLAS, cuDNN) for faster and realistic application runs or ii) without the external libraries if execution environments do not support them, e.g., microarchitecture simulators. This article presents a methodology to develop the variable-sized neural network benchmarks, and their performance and characteristics are evaluated based on hardware measurements. The results demonstrate that the Nebula benchmarks reduce execution time as much as 25x while preserving similar architectural behaviors as the full-fledged neural networks. Bogil Kim, Chanho Park 0004, William J. Song |
IEEE Trans. Computers | 5 |
| 2020 | Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresabstractThis paper introduces a GPU architecture named Duplo that minimizes redundant memory accesses of convolutions in deep neural networks (DNNs). Convolution is one of the fundamental operations used in various classes of DNNs, and it takes the majority of execution time. Various approaches have been proposed to accelerate convolutions via general matrix multiplication (GEMM), Winograd convolution, fast Fourier transform (FFT), etc. Recent introduction of tensor cores in NVIDIA GPUs particularly targets on accelerating neural network computations. A tensor core in a streaming multiprocessor (SM) is a specialized unit dedicated to handling matrix-multiply-and-accumulate (MMA) operations. The underlying operations of tensor cores represent GEMM calculations, and lowering a convolution can effectively exploit the tensor cores by transforming deeply nested convolution loops into matrix multiplication. However, lowering the convolution has a critical drawback since it requires a larger memory space (or workspace) to compute the matrix multiplication, where the expanded workspace inevitably creates multiple duplicates of the same data stored at different memory addresses. The proposed Duplo architecture tackles this challenge by leveraging compile-time information and microarchitectural supports to detect and eliminate redundant memory accesses that repeatedly load the duplicates of data in the workspace matrix. Duplo identifies data duplication based on memory addresses and convolution information generated by a compiler. It uses a load history buffer (LHB) to trace the recent load history of workspace data and their presence in register file. Every load instruction of workspace data refers to the LHB to find if potentially the same copies of data exist in the register file. If data duplicates are found, Duplo simply renames registers and makes them point to the ones containing the same values instead of issuing memory requests to load the same data. Our experiment results show that Duplo improves the performance of DNNs by 29.4% on average and saves 34.1% of energy using tensor cores. Sungwoo Ahn, Yunho Oh, Bogil Kim, Won Woo Ro, William J. Song |
MICRO | 6 |
| 2018 | FineReg: Fine-Grained Register File Management for Augmenting GPU ThroughputabstractGraphics processing units (GPUs) include a large amount of hardware resources for parallel thread executions. However, the resources are not fully utilized during runtime, and observed throughput often falls far below the peak performance. A major cause is that GPUs cannot deploy enough number of warps at runtime. The limited size of register file constrains the number of cooperative thread arrays (CTAs) as one CTA takes up a few tens of kilobytes of registers. We observe that the actual working set size of a CTA is much smaller in general, and therefore there is room for additional CTAs to run. In this paper, we propose a novel GPU architecture called FineReg that improves overall throughput by increasing the number of concurrent CTAs. In particular, FineReg splits the monolithic register file into two regions, one for active CTAs and another for pending CTAs. Using FineReg, the GPU begins normal executions by allocating all registers required by active CTAs. If all warps of a CTA become stalled, FineReg moves the live registers (i.e., working set) of CTA to the pending-CTA region and launches an additional CTA by assigning registers to the newly activated CTA. If the registers of either active or pending-CTA region are used up, FineReg stops introducing additional CTAs and simply performs context switching between active and pending CTAs. Thus, FineReg increases the number of concurrent CTAs by reducing the effective size of per-CTA registers. Experiment results show that FineReg achieves 32.8% of performance improvement over a conventional GPU architecture. Yunho Oh, Myung Kuk Yoon, William J. Song, Won Woo Ro |
MICRO | 3 |
| 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processorsabstractHeterogeneous multicore processors have been suggested as alternative microarchitectural designs to enhance performance and energy efficiency. Using Amdahl's Law, heterogeneous models were primarily analyzed in performance and energy efficiency aspects to demonstrate its advantage over conventional homogeneous systems. In this paper, we further extend the study to understand the lifetime reliability consequences of heterogeneous multicore processors, as reliability becomes an increasingly important constraint. We present the lifetime reliability models of multicore processors based on Amdahl's Law, including compact thermal estimation that has strong correlation with device aging. Lifetime reliability is analyzed by varying i) core utilization (Amdahl's scaling factor), ii) processor composition (number of big and small cores), and iii) thread scheduling method. The study shows that the heterogeneous processor may have a serious reliability challenge. If the processor is comprised of only one big core and many small cores, stresses can be biased to the big core especially when workloads spend more time on sequential operations. Our study reveals that incorporating multiple big cores can mitigate reliability bottleneck in big cores and enhance processor lifetime, but adding too many big cores will have an adverse impact on lifetime reliability as well as performance. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
HPCA | 1 |
| 2016 | Measurement-Driven Methodology for Evaluating Processor Heterogeneity Options for Power-Performance EfficiencyabstractIt is generally perceived that heterogeneous multicore processors will provide better performance and power efficiency over conventional homogeneous cores. However, heterogeneity can also be achieved within a homogeneous core design, instantiated under different voltage-frequency settings or per-core simultaneous multi-treading (SMT) modes. In this paper, we pursue an architectural study motivated by the question, "Can we get by with a single, complex SMT-equipped core design that can operate at different voltage-frequency points? Or, is it mandatory to invest into two different core types, one complex and the other simple?" We propose a systematic, measurement-driven methodology to evaluate processor heterogeneity options. Our analysis particularly focuses on the domain of real-time constrained embedded processors. The study is based on a direct measurement of two real processors; one that uses simple in-order cores, and another that uses complex out-of-order cores. The effect of heterogeneous core composition (consisting of complex and simple cores in the same chip) is analytically projected from measurements gleaned from the two different systems. Our analysis yields new interesting insights. When dealing with two core types without SMT enabled, true core heterogeneity does not necessarily provide better performance or power efficiency under area and power constraints. If the complex-core homogeneous processor invokes SMT, it outperforms true heterogeneity by offering 28% better power efficiency, assuming that simple cores in the heterogeneous system operate only in single-threaded mode without SMT capability. If the small cores employ SMT, true heterogeneity yields 32% better power efficiency than the homogeneous processor with SMT. William J. Song, Alper Buyuktosunoglu, Chen-Yong Cher, Pradip Bose |
ISLPED | 1 |
| 2014 | Energy Introspector: A parallel, composable framework for integrated power-reliability-thermal modeling for multicore architecturesabstractSustaining processor performance growth is challenged by physical limitations due to increased power and heat dissipations. Power and thermal management techniques combined with inherent workload dynamics create the spatiotemporal variations of power, temperature, and degradation in processors. As industry moves to smaller feature sizes, the performance will become increasingly dominated by the physics. The challenge is in understanding how the physics is manifested at the microarchitecture level. This requires the modeling and simulation environment that can capture multiple, distinct physical phenomena and their concurrent impact on the microarchitecture. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
ISPASS | 1 |
| 2014 | Manifold: A parallel simulation framework for multicore systemsabstractThis paper presents Manifold, an open-source parallel simulation framework for multicore architectures. It consists of a parallel simulation kernel, a set of microarchitecture components, and an integrated library of power, thermal, reliability, and energy models. Using the components as building blocks, users can assemble multicore architecture simulation models and perform serial or parallel simulations to study the architectural and/or the physical characteristics of the models. Users can also create new components for Manifold or port existing models. Importantly, Manifold's component-based design provides the user with the ability to easily replace a component with another for efficient explorations of the design space. It also allows components to evolve independently and making it easy for simulators to incorporate new components as they become available. The distinguishing features of Manifold include i) transparent parallel execution, ii) integration of power, thermal, reliability, and energy models, iii) full system simulation, e.g., operating system and system binaries, and iv) component-based design. In this paper we provide a description of the software architecture of Manifold, and its main elements - a parallel multicore emulator front-end and a parallel component-based back-end timing model. We describe a few simulators that are built with Manifold components to illustrate its flexibility, and present test results of the scalability obtained on full-system simulation of coherent shared-memory multicore models with 16, 32, and 64 cores executing PARSEC and SPLASH-2 benchmarks. Jun Wang 0077, Jesse G. Beu, Rishiraj A. Bheda, Thomas M. Conte, Zhenjiang Dong, Chad D. Kersey, Mitchelle Rasquinha, George F. Riley, William J. Song, Sudhakar Yalamanchili |
ISPASS | 9 |
| 2014 | Power Modeling for GPU Architectures Using McPATabstractGraphics Processing Units (GPUs) are very popular for both graphics and general-purpose applications. Since GPUs operate many processing units and manage multiple levels of memory hierarchy, they consume a significant amount of power. Although several power models for CPUs are available, the power consumption of GPUs has not been studied much yet. In this article we develop a new power model for GPUs by utilizing McPAT, a CPU power tool. We generate initial power model data from McPAT with a detailed GPU configuration, and then adjust the models by comparing them with empirical data. We use the NVIDIA's Fermi architecture for building the power model, and our model estimates the GPU power consumption with an average error of 7.7% and 12.8% for the microbenchmarks and Merge benchmarks, respectively. Jieun Lim 0001, Nagesh B. Lakshminarayana, Hyesoon Kim, William J. Song, Sudhakar Yalamanchili, Wonyong Sung |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2014 | Control Principles and On-Chip Circuits for Active Cooling Using Integrated Superlattice-Based Thin-Film Thermoelectric DevicesabstractSuperlattice thin-film thermoelectric coolers (TECs) are emerging as a promising technology for hot spot mitigation in microprocessors. This paper studies the prospect of on-demand cooling with advanced TECs integrated at the back of the heat spreader inside a package (integrated TEC). Using thermal compact models of the chip and package with integrated TECs, the control principles for TEC-assisted transient cooling are presented. The control principles are implemented in a 130-nm CMOS process and cosimulated with the thermal system to show their feasibility and energy overheads. The simulation results show potential for extending the time for which a chip and package can sustain a high power load. Borislav Alexandrov, Owen Sullivan, William J. Song, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | SST: A Scalable Parallel Framework for Architecture-Level Performance, Power, Area and Thermal SimulationabstractIn this paper, we describe the integrated power, area and thermal modeling framework in the structural simulation toolkit (SST) for large-scale high performance computer simulation. It integrates various power and thermal modeling tools and computes run-time energy dissipation for core, network on chip, memory controller and shared cache. It also provides functionality to update the leakage power as temperature changes. We illustrate the utilization of the framework by applying it to explore interconnect options in manycore systems with consideration of temperature variation and leakage feedback. We compare power, energy-delay-area product (EDAP) and energy-delay product (EDP) of four manycore configurations-1 core, 2 cores, 4 cores and 8 cores per cluster. Results from simulation with or without consideration of temperature variation both show that the 4-core per cluster configuration has the best EDAP and EDP. Even so, considering that temperature variation increases total power dissipation, we demonstrate the importance of considering temperature variation in the design flow. With this power, area and thermal modeling capability, the SST can be used for hardware/software co-design of future exascale systems. Ming-yu Hsieh, Rolf Riesen, Kevin Thompson 0004, William J. Song, Arun Rodrigues |
Comput. J. | 4 |