VLDB 2026 Research / reviewers in the wild / expert
Paul Gratz
dblp:18/4038 · also Paul V. Gratz
· DBLP profile ↗
79ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0001-7120-7189ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 4 first-author · 20 since 2021Software engineering, systems software and programming languages · 15 · 9 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Machine Learning-Driven Early Performance Prediction Framework for Accelerated Microarchitecture SimulationabstractRapid and accurate performance estimation is critical in evaluating novel microarchitectures, as it enables efficient exploration of architectural trade-offs. Unfortunately, traditional simulation techniques, while precise in predicting performance and power, incur tremendous slowdowns versus real machines. Despite prior works having explored machine learning–based performance prediction, the area remains far from sufficiently studied with existing approaches typically requiring large comprehensive datasets, frequent retraining, and heavy memory footprints with limited accuracy. Here, we introduce a new, fast and accurate, early-stage preview framework that uses partial simulation data, and leverages a smaller, faster tree-based machine learning (ML) model to forecast performance metrics such as IPC and Power. By training on a diverse set of configurations, our framework dynamically captures relationships between microarchitectural parameters in large OoO cores versus overall performance and other metrics. Collecting data from as few as 10 sample points taken during warmup, representing only 25 million instructions, our models achieve mean absolute percentage errors of 3-4%, preserving a majority of the model’s predictive accuracy while achieving a 25× speedup (96% reduction in simulation time). By comparison, linear regression techniques from the same point in simulation show an error of 50%. In cache DSE, we improve ranking accuracy by 25× compared to state-of-the-art prediction methods. Our results also show the proposed framework can accurately predict the performance of unseen (untrained) microarchitectural components including new prefetchers and branch predictors. Aiden Stickney, Osvaldo Castro, Aaron Chan, Paul Gratz, Jiang Hu 0001, Aakash Tyagi, Jered Dominguez-Trujillo, Galen M. Shipman, Kevin Sheridan |
DATE | 4 |
| 2026 | R-Max: Extending BéLáDy's MIN with Prefetching to Bound Realistic Cache Performance
Chia-Hang Lee, Maccoy Merrell, Gino Chacon, Daniel A. Jiménez, Paul Gratz |
ISCA | 6 |
| 2026 | Correct Wrong Path SimulationabstractModern OoO CPUs employ deep pipelines with high branch misprediction recovery penalties. Instructions speculatively executed along mispredicted paths can significantly alter microarchitectural state. During design space exploration, architects often rely on trace-driven simulators, which are significantly faster than execution-driven models but trade accuracy for speed. Despite this benefit, trace-driven simulation often fails to adequately model the effects of wrong-path execution because traces are typically collected only from the correct-path. While prior work can accurately model wrong-path effects on the instruction stream, it often makes unrealistic assumptions when modeling the impact on the data stream. In this work, we examine the effects of wrong-path execution and present an infrastructure for enabling its modeling in a tracedriven simulator. Our analysis shows that wrong-path execution extensively affects structures on both the instruction and data sides, yielding performance variations ranging from $-3.6 \%$ to $85.7 \%$ compared to a baseline that ignores these effects. To benefit the research community and enhance the accuracy of simulators, we provide our traces and tracing utility. We aim for this to encourage industry to provide wrong-path traces generated by internal simulators, enabling fast academic research without exposing industry proprietary IP. Chrysanthos Pepi, Krishnam Tibrewala, Bhargav Reddy Godala, Sankara Prasad Ramesh, Alberto Ros 0001, Daniel A. Jiménez, Gilles Pokam, Paul Gratz |
ISPASS | 8 |
| 2025 | Skia: Exposing Shadow BranchesabstractModern processors implement a decoupled front-end, often using a form of Fetch Directed Instruction Prefetching (FDIP), to avoid front-end stalls. FDIP is driven by the Branch Prediction Unit (BPU), relying on the BPU's accuracy and branch target tracking structures to speculatively fetch instructions into the Instruction Cache (L1-I cache). As contemporary data center applications become more complex, their code footprints also grow, resulting in a high number of Branch Target Buffer (BTB) misses. These BTB missing branches typically have previously been decoded and placed in the BTB, but have since been evicted, leading to BTB misses now. FDIP can alleviate L1-I cache misses, but its reliance on the BPU's tracking structures means that when it encounters a BTB miss, the BPU may not identify the current instruction as a branch to FDIP. This can prevent FDIP from prefetching or cause it to speculate down the wrong path, further polluting the L1-I cache. Chrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino Chacon, Paul Gratz, Daniel A. Jiménez, Gilles Pokam, David I. August |
ASPLOS (2) | 5 |
| 2025 | Light-weight Cache Replacement for Instruction Heavy WorkloadsabstractThe last-level cache (LLC) is the last chance for memory accesses from the processor to avoid the costly latency of accessing the main memory.In recent years, an increasing number of instruction heavy workloads have put pressure on the last-level cache.We find that, for instruction heavy workloads, a simple replacement policy with minimal overhead provides at least the same benefit as a stateof-the-art, high-overhead replacement policy in the presence of aggressive prefetching.Our proposal is based on specifying insertion and promotion vectors (IPVs) as a generalization of re-reference interval prediction (RRIP) in such a way that the space of feasible policies may be searched exhaustively to find the best policy for the training set of workloads.The policies are formulated to deliver the best performance taking into account demand and prefetch accesses.We show that our technique, Prefetch Aware Coarse-grained Insertion and Promotion Vectors (PACIPV), improves performance over a state-of-the-art LLC replacement policy (Mockingjay) for instruction heavy workloads, and remains competitive for data heavy workloads with significantly less hardware overhead.We show that RRIP-based IPVs are very easy to implement but outperform far more complex replacement policies.PACIPV achieves a speedup of 3.3% over the baseline of LRU, outperforming SRRIP by 1.1% and the much more hardware intensive Mockingjay by 0.1%. Saba Mostofi, Setu Gupta, Ahmad Hassani, Krishnam Tibrewala, Elvira Teran, Paul Gratz, Daniel A. Jiménez |
ISCA | 6 |
| 2025 | Benchmarking 3D Gaussian Splatting RenderingabstractThe growing demand for 3D modeling, particularly in applications like augmented reality (AR) and virtual reality (VR), has underscored the need for more efficient techniques. Manual 3D modeling requires extensive effort and a specialized skill set. Although traditional Photogrammetry is faster than manual 3D modeling, it is still compute-intensive and time-consuming. Emerging methods like 3D Gaussian Splatting (3DGS) offer a faster and more cost-effective solution for model generation (training), at the cost of requiring a different rendering framework. For their widespread adoption in applications such as AR and VR, it is crucial to meet rendering performance requirements, such as power and frame rate. Existing works mainly focus on 3DGS training performance and lack comprehensive analysis and comparison with traditional graphics (TG) rendering techniques that use a mesh representation. Hence, a thorough benchmarking of 3DGS rendering is necessary to identify its limitations and potential areas for improvement. In this paper, we conduct a comprehensive performance study of 3DGS rendering using state-of-the-art frameworks on three different hardware platforms. We evaluate 3DGS performance against TG rendering in terms of frames-per-second (FPS), power, GPU memory footprint, GPU utilization, frametime breakdown, FPS to Watt, several GPU performance counters, and rendered image quality. We observe that 3DGS generates high-quality images with an average PSNR of 37. Our analysis reveals that 3DGS rendering requires a 3x improvement in the FPS to Watt to achieve the performance of TG rendering. Our evaluation shows that 3DGS occupies, on average, a 2x lesser GPU memory footprint compared to TG rendering. Our results indicate that around 10 % image quality can be a tradeoff for a 100 % improvement in frame rate. We see that the rasterization step, on average, consumes 64.76 % of frame time and is the main bottleneck to 3DGS. We identify that 3DGS has, on average, 25 % higher GPU utilization than TG rendering. However, its performance is limited by stalls due to branching and synchronization, pointing to possible improvements in 3DGS algorithm and making it GPU friendly. Saichand Samudrala, Sushant Kondguli, Paul Gratz |
ISPASS | 3 |
| 2024 | Flow Correlator: A Flow Table Cache Management StrategyabstractSwitching, routing, and security functions are the backbone of packet processing networks. Fast and efficient processing of packets requires maintaining the state information for many transient network connections. In particular, modern stateful firewalls, security monitoring devices, and Software-Defined Networking (SDN) dataplanes require maintaining state-ful flow tables. These flow tables often grow much larger than can fit on-chip, requiring caching to maintain performance.This paper focuses on improving caching efficiency, an important architectural component of the packet processing data planes. We present a novel predictive approach (Flow Correlator) to network flow table cache management by adapting the Hashed Perceptron binary classifier to improve the reliability and performance of the data plane caching. We also discovered an iterative approach to feature selection and ranking while adapting the Hashed Perceptron mechanism to network flow table cache management.Through extensive experimentation, we demonstrate improved caching efficiency of the proposed Flow Correlator mechanism. We also rigorously validate the performance and generic applicability of our technique across real-world datasets. Luke McHale, Paul Gratz, Alexander Sprintson |
ICCCN | 2 |
| 2024 | Aiding Microprocessor Performance Validation with Machine LearningabstractMicroprocessor validation is a complex task that consumes substantial engineering time. Degradation of the system performance that does not affect its functional correctness, is particularly difficult to address given the lack of a golden reference for performance. This work introduces an automated methodology based on machine learning to assist in localizing performance faults, aiming to speed up the validation process. Our results show that, for the injected performance issues, whose average IPC impact is$> 1{\%}$, our technique is able to help localize the exact microarchitectural unit where the degradation occurs$\sim$75% of the time while achieving a top-3 unit accuracy (out of 11 possible locations) of$> 97{\%}$. The proposed setup requires a few seconds to perform a localization inference, leading to a reduced validation time. Erick Carvajal Barboza, Mahesh Ketkar, Paul Gratz, Jiang Hu 0001 |
ISPASS | 3 |
| 2024 | Coherence Attacks and Countermeasures in Interposer-based Chiplet SystemsabstractIndustry is moving towards large-scale hardware systems that bundle processor cores, memories, accelerators, and so on. via 2.5D integration. These components are fabricated separately as chiplets and then integrated using an interposer as an interconnect carrier. This new design style is beneficial in terms of yield and economies of scale, as chiplets may come from various vendors and are relatively easy to integrate into one larger sophisticated system. However, the benefits of this approach come at the cost of new security challenges, especially when integrating chiplets that come from untrusted or not fully trusted, third- party vendors. In this work, we explore these challenges for modern interposer-based systems of cache-coherent, multi-core chiplets. First, we present basic coherence-oriented hardware Trojan attacks that pose a significant threat to chiplet-based designs and demonstrate how these basic attacks can be orchestrated to pose a significant threat to interposer-based systems. Second, we propose a novel scheme using an active interposer as a generic, secure-by-construction platform that forms a physical root of trust for modern 2.5D systems. The implementation of our scheme is confined to the interposer, resulting in little cost and leaving the chiplets and coherence system untouched. We show that our scheme prevents a range of coherence attacks with low overheads on system performance, ∼4%. Further, we demonstrate that our scheme scales efficiently as system size and memory capacities increase, resulting in reduced performance overheads. Gino Chacon, Johann Knechtel, Ozgur Sinanoglu, Paul Gratz, Vassos Soteriou |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | A Characterization of the Effects of Software Instruction Prefetching on an Aggressive Front-endabstractGrowing application sizes continue to strain the memory system. As more complex applications are developed, and the instruction memory footprint increases, the cache hierarchy cannot contain relevant instructions causing the front end to become idle as it awaits fetched instructions. Hardware instruction prefetchers can alleviate this problem by learning instruction stream behavior and prefetching instructions into the cache before use. Fetch Directed Prefetching (FDP) is a ubiquitous form of hardware instruction prefetching that uses branch predictor to predict future instruction cache references. Modern processors generally implement aggressive, deep FDP to decouple the front-end from the rest of the machine. Instruction accesses, however, provide a small amount of information over long periods due to instruction stream variability and lack of information regarding the context of instruction accesses. Capturing the instruction stream’s context requires significant storage overhead to correlate instruction accesses and ensure timely accesses. Software prefetching techniques overcome this problem by profiling and statically analyzing an application’s behavior. Prior work demonstrates software instruction prefetching’s potential performance benefit but does not evaluate performance in the context of aggressive, decoupled front-ends. While software prefetching provides ~ 20% improvement in conservative front-ends, we find that it does not yield performance benefit when modeling a baseline with an aggressive FDP, in some cases hurting performance. Our analysis finds that using software instruction prefetching negatively impacts an aggressive front-end’s behavior. We investigate this finding and characterize the different states a front-end can be in and how introducing instructions into an application can change the front-end’s behavior resulting in destructive interference with the software prefetcher. Gino Chacon, Nathan Gober, Krishnendra Nathella, Paul Gratz, Daniel A. Jiménez |
ISPASS | 4 |
| 2023 | KVRangeDB: Range Queries for a Hash-based Key-Value DeviceabstractKey–value (KV) software has proven useful to a wide variety of applications including analytics, time-series databases, and distributed file systems. To satisfy the requirements of diverse workloads, KV stores have been carefully tailored to best match the performance characteristics of underlying solid-state block devices. Emerging KV storage device is a promising technology for both simplifying the KV software stack and improving the performance of persistent storage-based applications. However, while providing fast, predictable put and get operations, existing KV storage devices do not natively support range queries that are critical to all three types of applications described above. In this article, we present KVRangeDB, a software layer that enables processing range queries for existing hash-based KV solid-state disks (KVSSDs). As an effort to adapt to the performance characteristics of emerging KVSSDs, KVRangeDB implements log-structured merge tree key index that reduces compaction I/O, merges keys when possible, and provides separate caches for indexes and values. We evaluated the KVRangeDB under a set of representative workloads, and compared its performance with two existing database solutions: a Rocksdb variant ported to work with the KVSSD, and Wisckey, a key–value database that is carefully tuned for conventional block devices. On filesystem aging workloads, KVRangeDB outperforms Wisckey by 23.7× in terms of throughput and reduce CPU usage and external write amplifications by 14.3× and 9.8×, respectively. Qing Zheng, Jason Lee 0004, Bradley W. Settlemyer, Fei Wen 0003, A. L. Narasimha Reddy, Paul Gratz |
ACM Trans. Storage | 7 |
| 2022 | Stay in your Lane: A NoC with Low-overhead Multi-packet BypassingabstractNoCs are over-provisioned with large virtual channels to provide deadlock freedom and performance improvement. This use of virtual channels leads to considerable power and area overhead. In this paper, we introduce a novel flow control, called FastFlow, to enhance performance and avoid both protocol- and network-level deadlocks with an impressive reduction in number of virtual channels compared to the state-of-the-art NoCs. FastPass promotes a packet to traverse the network bufferlessly; the packet bypasses the routers to reach its destination. During the traversals, the packet is guaranteed to make forward progress every cycle. As a result, such a packet cannot be blocked by congestion nor deadlock. Promoting more packets to FastPass will provide higher throughput. To this end, FastPass allows multiple packets to be upgraded as FastPass packets simultaneously. To avoid any collision between these packets, FastPass provides multiple pre-defined non-overlapping lanes. Each lane is allowed to propagate only one FastPass packet. In a time-division multiplexed way, each router gets a chance to upgrade its packets to the FastPass packets and then transfer them via the pre-defined non-overlapping lanes. FastPass not only provides high throughput but also resolves both protocol-and network-level deadlocks. Compared to the state-of-the-art, FastPass presents a 1.8 × increase in throughput for synthetic traffic, 46% improvement in average packet latency for real applications, and 40% reduction in power and area consumption. Hossein Farrokhbakht, Paul Gratz, Tushar Krishna, Joshua San Miguel, Natalie D. Enright Jerger |
HPCA | 2 |
| 2022 | SLAP-CC: Set-Level Adaptive Prefetching for Compressed CachesabstractData prefetching and cache compression are well-studied techniques to reduce the impact of memory latency. Data prefetching predicts future memory accesses and prefills the cache with the corresponding memory blocks in advance of explicit demands. Cache compression tries to increase the effective capacity of the cache with minimum area overhead. As we show, however, naïvely integrating the two techniques does not yield additive gains. We show that the extra ways per set that cache compression provides generally produce fewer hits than the ways in the baseline uncompressed cache. Hence, prefetching more aggressively to those sets with compression-added ways and more conservatively to those sets with less compression provides substantial benefit. In this paper we present set-level adaptive prefetching for compressed caches (SLAP-CC), a compressibility aware prefetching technique that adapts prefetch aggressiveness to the workload compressibility to maximize prefetch coverage. SLAP-CC dynamically adjusts the prefetch confidence threshold on a per-set basis, based on the set’s compression. SLAP-CC achieves a geometric mean speedup of 18.0% over a baseline system with compressed cache and no prefetching. Further, SLAP-CC significantly outperforms existing state of the art L2 prefetchers, SPP [1] and Best Offset [2] when combined with cache compression by an average of 7% and up to 25% on some workloads. Laith M. AlBarakat, Paul Gratz, Daniel A. Jiménez |
ICCD | 2 |
| 2022 | Composite Instruction PrefetchingabstractPrefetching is a pivotal mechanism for effectively masking latencies due to the processor/memory performance gap. Instruction prefetchers prevent costly instruction fetch stalls by requesting blocks of instruction memory in advance of their use to keep the pipeline front-end busy. the rapidly increasing instruction footprints of modern workloads have amplified the importance of such research.We propose a framework to leverage the complementary prefetching behaviors of existing prefetching techniques to create composite prefetchers. We show that recently proposed instruction prefetching techniques leverage different mechanisms from one another and find that in many cases, different prefetchers are complementary to each other. Composite prefetching allows for higher performance at lower storage overheads by combining the coverage of different complex prefetchers. We demonstrate a framework for selecting and combining state-of-the-art complex prefetchers, in a "plug-and-play" fashion, to identify the best performing combinations at various hardware overheads. We show that for every storage capacity constraint analyzed, composite prefetching outperforms prior prefetching schemes with greater improvements shown at smaller capacity constraints. Gino Chacon, Elba Garza, Alexandra Jimborean, Alberto Ros 0001, Paul Gratz, Daniel A. Jiménez, Samira Mirbagher Ajorpaz |
ICCD | 5 |
| 2022 | Page Size Aware Cache PrefetchingabstractThe increase in working set sizes of contemporary applications outpaces the growth in cache sizes, resulting in frequent main memory accesses that deteriorate system performance due to the disparity between processor and memory speeds. Prefetching data blocks into the cache hierarchy ahead of demand accesses has proven successful at attenuating this bottleneck. However, spatial cache prefetchers operating in the physical address space leave significant performance on the table by limiting their pattern detection within 4KB physical page boundaries when modern systems use page sizes larger than 4KB to mitigate the address translation overheads. This paper exploits the high usage of large pages in modern systems to increase the effectiveness of spatial cache prefetching. We design and propose the Page-size Propagation Module (PPM), a $\mu$architectural scheme that propagates the page size information to the lower-level cache prefetchers, enabling safe prefetching beyond 4KB physical page boundaries when the accessed blocks reside in large pages, at the cost of augmenting the first-level caches’ Miss Status Holding Register (MSHR) entries with one additional bit. PPM is compatible with any cache prefetcher without implying design modifications. We capitalize on PPM’s benefits by designing a module that consists of two page size aware prefetchers that inherently use different page sizes to drive prefetching. The composite module uses adaptive logic to dynamically enable the most appropriate page size aware prefetcher. Finally, we show that the proposed designs are transparent to which cache prefetcher is used. We apply the proposed page size exploitation techniques to four state-of-the-art spatial cache prefetchers. Our evaluation shows that our proposals improve single-core geomean performance by up to 8.1% (2.1% at minimum) over the original implementation of the considered prefetchers, across 80 memory-intensive workloads. In multi-core contexts, we report geomean speedups up to 7.7% across different cache prefetchers and core configurations. Georgios Vavouliotis, Gino Chacon, Lluc Alvarez, Paul Gratz, Daniel A. Jiménez, Marc Casas |
MICRO | 4 |
| 2022 | Reducing Minor Page Fault Overheads through Enhanced Page WalkerabstractApplication virtual memory footprints are growing rapidly in all systems from servers down to smartphones. To address this growing demand, system integrators are incorporating ever larger amounts of main memory, warranting rethinking of memory management. In current systems, applications produce page fault exceptions whenever they access virtual memory regions that are not backed by a physical page. As application memory footprints grow, they induce more and more minor page faults. Handling of each minor page fault can take a few thousands of CPU cycles and blocks the application till the OS kernel finds a free physical frame. These page faults can be detrimental to the performance when their frequency of occurrence is high and spread across application runtime. Specifically, lazy allocation-induced minor page faults are increasingly impacting application performance. Our evaluation of several workloads indicates an overhead due to minor page faults as high as 29% of execution time. In this article, we propose to mitigate this problem through a hardware, software co-design approach. Specifically, we first propose to parallelize portions of the kernel page allocation to run ahead of fault time in a separate thread. Then we propose the Minor Fault Offload Engine (MFOE), a per-core hardware accelerator for minor fault handling. MFOE is equipped with a pre-allocated page frame table that it uses to service a page fault. On a page fault, MFOE quickly picks a pre-allocated page frame from this table, makes an entry for it in the TLB, and updates the page table entry to satisfy the page fault. The pre-allocation frame tables are periodically refreshed by a background kernel thread, which also updates the data structures in the kernel to account for the handled page faults. We evaluate this system in the gem5 architectural simulator with a modified Linux kernel running on top of simulated hardware containing the MFOE accelerator. Our results show that MFOE improves the average critical path fault handling latency by 33× and tail critical path latency by 51×. Among the evaluated applications, we observed an improvement of runtime by an average of 6.6%. Chandrahas Tirumalasetty, Chih-Chieh Chou, A. L. Narasimha Reddy, Paul Gratz, Ayman Abouelwafa |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | SIMD-Matcher: A SIMD-based Arbitrary Matching FrameworkabstractPacket classification methods rely upon matching packet content/header against pre-defined rules, which are generated by network applications and their configurations. With the rapid development of network technology and the fast-growing network applications, users seek more enhanced, secure, and diverse network services. Hence it becomes critical to improve the performance of arbitrary matching operations. This article presents SIMD-Matcher, an efficient Single Instruction Multiple Data (SIMD) and cache-friendly arbitrary matching framework. To further improve the arbitrary matching performance, SIMD-Matcher adopts a trie node with a fixed high fanout and a varying span for each node depending on the data distribution. The trie node layout leverages cache and modern processor features such as SIMD instructions. To support arbitrary matching, we first interpret arbitrary rules into three fields: value, mask, and priority. Second, to support insertion of randomly positioned wildcards to arbitrary rules, we propose the SIMD-Matcher extraction algorithm to process the wildcard bits. Third, we add an array of wildcard entries to the leaf entries, which store the wildcard rules and guarantee the correctness of matching results. Experiments show that SIMD-Matcher outperforms GenMatcher under large-scale ruleset and key set, in terms of search time, insert time, and memory cost. Specifically with 5M rules, our method achieves a 2.7X speedup on search time, and the insertion time takes \( ~\sim \!\! 7.3 \) seconds, gaining a 1.38X speedup; meanwhile, the memory cost reduction is up to 6.17X. Ping Wang 0043, Fei Wen 0003, Paul Gratz, Alexander Sprintson |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Software Hint-Driven Data Management for Hybrid Memory in Mobile SystemsabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. Recently emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of distinct memory classes render a new challenge for memory system design. Ideally, pages should be placed or migrated between the two types of memories according to the data objects’ access properties. Prior system software approaches exploit the program information from OS but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recent memory requests, due to the limited on-chip resources. In this work, we propose OpenMem: a hardware-software cooperative approach that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. First, we built a hardware-based memory manager unit (HMMU) that can learn the short-term access patterns by online profiling, and execute data migration efficiently. Then, we built a heap memory manager for the heterogeneous memory systems that allows the programmer to directly customize each data object’s allocation to a favorable memory device within the presumed object life cycle. With the programmer’s hints guiding the data placement at allocation time, data objects with similar properties will be congregated to reduce unnecessary page migrations. We implemented the whole system on the FPGA board with embedded ARM processors. In testing under a set of benchmark applications from SPEC 2017 and PARSEC, experimental results show that OpenMem reduces 44.6% energy consumption with only a 16% performance degradation compared to the all-DRAM memory system. The amount of writes to the NVM is reduced by 14% versus the HMMU-only, extending the NVM device lifetime. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | OpenMem: Hardware/Software Cooperative Management for Mobile Memory SystemabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of current mobile applications. NVM technologies have higher capacity density, minimal static power consumption, but longer access latency and limited write endurance compared to DRAM. The different characteristics of these two memory classes, however, pose new challenges for memory system design. Ideally, pages shall be placed or migrated between the two types of memories according to the data objects’ access properties. Prior works use the OS for placement and migration in these systems, but at the cost of high software latency incurred by related kernel processes. Hardware approaches can avoid these latencies, however, hardware’s vision is constrained to a short time window of recently memory request, due to the limited on-chip resources.In this work, we propose OpenMem: a hardware-software cooperative approach to address placement and migration within hybrid memory systems, that combines the execution time advantages of pure hardware approaches with the data object properties in a global scope. We emulate OpenMem on an FPGA board with embedded ARM CPU, and run a set of benchmark applications from SPEC 2017 and PARSEC. Experimental results show that OpenMem reduces energy consumption by 44.6% with only a 16% performance degradation compared to an all-DRAM memory system. Further, writes to the NVM are reduced by 14% versus a hardware-only approach, extending the NVM device lifetime. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
DAC | 3 |
| 2021 | CMRC: Comprehensive Microarchitectural Register Coalescing for GPGPUsabstractGraphics processing units (GPUs) deploy a large register file (RF) to achieve high compute throughput. This RF, however, consumes a large portion of the total dynamic power in the GPU. Additionally, the RF banks and operand collectors (OCs) are designed with limited number of ports causing access serialization and negatively impacting performance. In this work, we introduce CMRC, a coalescing-aware RF organization that takes advantage of frequent narrow-width data present in general purpose applications to increase performance and reduce energy for GPGPUs. CMRC is a low-cost comprehensive approach to register coalescing capable of combining narrow-width read and write accesses from same or different warp instructions into fewer accesses, reducing port contention and access pressure. On general purpose applications, CMRC reduces RF accesses by 31.8%, achieves a performance speedup of 16.5%, and reduces overall GPU energy by 32.2% on average, outperforming best of class prior work by ~1.8x without the requirement of compiler support. Ahmad M. Radaideh, Paul Gratz |
DATE | 2 |
| 2021 | An FPGA-based Hybrid Memory Emulation SystemabstractHybrid memory systems, comprised of emerging non-volatile memory (NVM) and DRAM, have been proposed to address the growing memory demand of applications. Emerging NVM technologies, such as phase-change memories (PCM), memristor, and 3D XPoint, have higher capacity density, minimal static power consumption and lower cost per GB. However, NVM has longer access latency and limited write endurance as opposed to DRAM. The different characteristics of two memory classes point towards the design of hybrid memory systems containing multiple classes of main memory.In the iterative and incremental development of new architectures, the timeliness of simulation completion is critical to project progression. Hence, a highly efficient simulation method is needed to evaluate the performance of different hybrid memory system designs. Design exploration for hybrid memory systems is challenging, because it requires emulation of the full system stack, including the OS, memory controller, and interconnect. Moreover, benchmark applications for memory performance tests typically have much larger working sets, thus taking an even longer simulation warm-up period.In this paper, we propose an FPGA-based hybrid memory system emulation platform. We target the mobile computing system, which is sensitive to energy consumption and is likely to adopt NVM for its power efficiency. The focus of our platform is on the design of hybrid memory system, so we leverage the on-board hard IP ARM processors to enhance simulation performance while improving the accuracy of results. Thus, users can implement their data placement/migration policies with the FPGA logic elements and evaluate new designs quickly and effectively. Results show that our emulation platform provides a speedup of 9280x in simulation time compared to the software counterpart gem5. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
FPL | 3 |
| 2021 | Automatic Microprocessor Performance Bug DetectionabstractProcessor design validation and debug is a difficult and complex task, which consumes the lion's share of the design process. Design bugs that affect processor performance rather than its functionality are especially difficult to catch, particularly in new microarchitectures. This is because, unlike functional bugs, the correct processor performance of new microarchitectures on complex, long-running benchmarks is typically not deterministically known. Thus, when performance benchmarking new microarchitectures, performance teams may assume that the design is correct when the performance of the new microarchitecture exceeds that of the previous generation, despite significant performance regressions existing in the design. In this work we present a two-stage, machine learning-based methodology that is able to detect the existence of performance bugs in microprocessors. Our results show that our best technique detects 91.5% of microprocessor core performance bugs whose average IPC impact across the studied applications is greater than 1% versus a bug-free design with zero false positives. When evaluated on memory system bugs, our technique achieves 100% detection with zero false positives. Moreover, the detection is automatic, requiring very little performance engineer time. Erick Carvajal Barboza, Sara Jacob, Mahesh Ketkar, Michael Kishinevsky, Paul Gratz, Jiang Hu 0001 |
HPCA | 5 |
| 2021 | Pitstop: Enabling a Virtual Network Free Network-on-ChipabstractMaintaining correctness is of paramount importance in the design of a computer system. Within a multiprocessor interconnection network, correctness is guaranteed by having deadlock-free communication at both the protocol and network levels. Modern network-on-chip (NoC) designs use multiple virtual networks to maintain protocol-level deadlock freedom, at the expense of high power and area overheads. Other techniques involve complex detection and recovery mechanisms, or use misrouting which incurs additional packet latency. Considering that the probability of deadlocks occurring is low, the additional resources needed to avoid/resolve deadlocks should also be low. To this end, we propose Pitstop, a low-cost technique that guarantees correctness by resolving both protocol and network-level deadlocks without the use of virtual networks, complex hardware, or misrouting. Pitstop transfers blocked packets to the network interface (NI) creating a bubble (empty buffer slot) which breaks deadlock. The blocked packet can make forward progress through NI to NI traversals using low complexity bypassing mechanisms. This scheme performs better due to higher utilization of virtual channels and works on arbitrary irregular topologies without any virtual networks. Compared to state-of-the-art solutions, Pitstop can improve performance up to 11% and reduce power and area up to 41% and 40%. Hossein Farrokhbakht, Henry Kao, Kamran Hasan, Paul Gratz, Tushar Krishna, Joshua San Miguel, Natalie D. Enright Jerger |
HPCA | 4 |
| 2021 | SEEC: stochastic escape express channelabstractAllocating a free buffer before moving to the next router is a fundamental tenet for packet movement in NoCs. Often, to solve head of line blocking and avoid deadlock, NoCs are provisioned with significant buffer resources in the form of virtual channels (VC) which consume area and power. We introduce stochastic escape express channels (SEEC) to enhance performance and avoid deadlock with dramatically fewer buffers than state-of-the-art NoCs. The network interfaces in SEEC periodically send special tokens called seekers to find packets destined for them and upgrade them to use a novel flow control called Free-Flow (FF). FF-packets traverse the network minimally from link to link, bypassing routers (bufferlessly) to the destination. As a result, FF-packets bypass regions of congestion in the NoC without needing more buffers. Furthermore, any deadlock that a FF-packet was originally involved in is guaranteed to break, without requiring turn restrictions or extra VCs. We also present an extension called multi-SEEC (mSEEC) that enables multiple simultaneous non-intersecting FF-packet traversals to enhance throughput further. We implement and evaluate SEEC and mSEEC on a mesh over a range of synthetic workloads and real applications and observe 34--40% reduction in average packet latency for real applications and 10--50% average improvement in throughput for synthetic traffic over the state-of-the-art at 1/6th the area/power budget. Mayank Parasar, Natalie D. Enright Jerger, Paul Gratz, Joshua San Miguel, Tushar Krishna |
SC | 3 |
| 2021 | KVRAID: high performance, write efficient, update friendly erasure coding scheme for KV-SSDsabstractKey-value (KV) stores have been widely deployed in a variety of scale-out enterprise applications such as online retail, big data analytics, social networks, etc. Key-Value SSDs (KVSSDs) provide a key-value interface directly from the device aiming at lowering software overhead and reducing I/O amplification for such applications. A. L. Narasimha Reddy, Paul Gratz, Rekha Pitchumani, Yang-Seok Ki |
SYSTOR | 3 |
| 2020 | A Generic FPGA Accelerator for Minimum Storage Regenerating CodesabstractErasure coding is widely used in storage systems to achieve fault tolerance while minimizing the storage overhead. Recently, Minimum Storage Regenerating (MSR) codes are emerging to minimize repair bandwidth while maintaining the storage efficiency. Traditionally, erasure coding is implemented in the storage software stacks, which hinders normal operations and blocks resources that could be serving other user needs due to poor cache performance and costs high CPU and memory utilizations. In this paper, we propose a generic FPGA accelerator for MSR codes encoding/decoding which maximizes the computation parallelism and minimizes the data movement between off-chip DRAM and the on-chip SRAM buffers. To demonstrate the efficiency of our proposed accelerator, we implemented the encoding/decoding algorithms for a specific MSR code called Zigzag code on Xilinx VCU1525 acceleration card. Our evaluation shows our proposed accelerator can achieve ~2.4-3.1x better throughput and ~4.2-5.7x better power efficiency compared to the state-of-art multi-core CPU implementation and ~2.8-3.3x better throughput and ~4.2-5.3x better power efficiency compared to a modern GPU accelerator. Joo Hwan Lee, Rekha Pitchumani, Yang-Seok Ki, A. L. Narasimha Reddy, Paul Gratz |
ASP-DAC | 6 |
| 2020 | Exploiting Zero Data to Reduce Register File and Execution Unit Dynamic Power Consumption in GPGPUsabstractTo achieve high compute performance, graphics processing units (GPUs) provide a large register file and a large number of execution units. However, these design components consume a large portion of the total dynamic power in the GPU, particularly for general purpose applications. In this paper, we present a low-cost gating scheme to reduce dynamic power consumption in the register file and execution units without impacting performance. The scheme proposed dynamically exploit frequent found data value of zeros within and across registers in order to gate off register file reads and writes as well as execution units. We find that on general purpose applications from Rodinia, our low-cost gating scheme can reduce register file reads and writes on average by 35% and 40%, respectively. The register file and execution unit dynamic power are reduced on average by 19% and 13%, respectively. The reduction in total GPU dynamic power achieved is ranging from 3% to 19% with 8% on average with no performance loss. Ahmad M. Radaideh, Paul Gratz |
DAC | 2 |
| 2020 | DRAIN: Deadlock Removal for Arbitrary Irregular NetworksabstractCorrectness is a first-order concern in the design of computer systems. For multiprocessors, a primary correctness concern is the deadlock-free operation of the network and its coherence protocol; furthermore, we must guarantee the continued correctness of the network in the face of increasing faults. Designing for deadlock freedom is expensive. Prior solutions either sacrifice performance or power efficiency to proactively avoid deadlocks or impose high hardware complexity to reactively resolve deadlocks as they occur. However, the precise confluence of events that lead to deadlocks is so rare that minimal resources and time should be spent to ensure deadlock freedom. To that end, we propose DRAIN, a subactive approach to remove potential deadlocks without needing to explicitly detect or avoid them. We simply let deadlocks happen and periodically drain (i.e., force the movement of) packets in the network that may be involved in a cyclic dependency. As deadlocks are a rare occurrence, draining can be performed infrequently and at low cost. Unlike prior solutions, DRAIN eliminates not only routing-level but also protocol-level deadlocks without the need for expensive virtual networks. DRAIN dramatically simplifies deadlock freedom for irregular topologies and networks that are prone to wear-related faults. Our evaluations show that on an average, DRAIN can save 26.73% packet latency compared to proactive deadlock-freedom schemes in the presence of faults while saving 77.6% power compared to reactive schemes. Mayank Parasar, Hossein Farrokhbakht, Natalie D. Enright Jerger, Paul Gratz, Tushar Krishna, Joshua San Miguel |
HPCA | 4 |
| 2020 | SB-Fetch: synchronization aware hardware prefetching for chip multiprocessorsabstractShared-memory, multi-threaded applications often require programmers to insert thread synchronization primitives (i.e. locks, barriers, and condition variables) in critical sections to synchronize data access between processes. Scaling performance requires balanced per-thread workloads with little time spent in critical sections. In practice, however, threads often waste time waiting to acquire locks/barriers, leading to thread imbalance and poor performance scaling. Moreover, critical sections often stall data prefetchers that mitigate the effects of waiting by ensuring data is preloaded in core caches when the critical section is done. Laith M. AlBarakat, Paul Gratz, Daniel A. Jiménez |
ICS | 2 |
| 2020 | Virtualize and share non-volatile memories in user space
Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt |
CCF Trans. High Perform. Comput. | 4 |
| 2020 | Hardware Memory Management for Future Mobile Hybrid Memory SystemsabstractThe current mobile applications have rapidly growing memory footprints, posing a great challenge for memory system design. Insufficient DRAM main memory will incur frequent data swaps between memory and storage, a process that hurts performance, consumes energy, and deteriorates the write endurance of typical flash storage devices. Alternately, a larger DRAM has higher leakage power and drains the battery faster. Furthermore, DRAM scaling trends make further growth of DRAM in the mobile space prohibitive due to cost. Emerging nonvolatile memory (NVM) has the potential to alleviate these issues due to its higher capacity per cost than DRAM and minimal static power. Recently, a wide spectrum of NVM technologies, including phase-change memories (PCMs), memristor, and 3-D XPoint has emerged. Despite the mentioned advantages, NVM has longer access latency compared to DRAM and NVM writes can incur higher latencies and wear costs. Therefore, the integration of these new memory technologies in the memory hierarchy requires a fundamental rearchitecting of traditional system designs. In this work, we propose a hardware-accelerated memory manager (HMMU) that addresses in a flat address space, with a small partition of the DRAM reserved for subpage block-level management. We design a set of data placement and data migration policies within this memory manager such that we may exploit the advantages of each memory technology. By augmenting the system with this HMMU, we reduce the overall memory latency while also reducing writes to the NVM. The experimental results show that our design achieves a 39% reduction in energy consumption with only a 12% performance degradation versus an all-DRAM baseline that is likely untenable in the future. Fei Wen 0003, Paul Gratz, A. L. Narasimha Reddy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | The Best of IEEE Computer Architecture Letters in 2018abstractThree papers are noted: "Amoeba: An Autonomous Backup and Recovery SSD for Ransomware Attack Defense", by Donghyun Min, Donggyu Park, Jinwoo Ahn, Ryan Walker, Junghee Lee, Sungyong Park, Youngjae Kim, Sogang (University and University of Texas at San Antonio, USA); "The Architectural Implications of Cloud Microservices", by Yu Gan and Christina Delimitrou, (Cornell University, USA); and "An Alternative Analytical Approach to Associative Processing", Soroosh Khoram, Yue Zha, and Jing Li, (University of Wisconsin-Madison, USA). Paul Gratz |
HPCA | 1 |
| 2019 | SpecLock: Speculative Lock ForwardingabstractAs core counts increase, lock acquisition and release become even more critical because they lie on the critical path of shared memory applications. In this paper, we show that many applications exhibit regular and repeating lock sharing patterns. Based on this observation, we introduce SpecLock, an efficient hardware mechanism which speculates on the lock acquisition pattern between cores. Upon the release of a lock, the cache line containing the lock is speculatively forwarded to the next consumer of the lock. This forwarding action is performed via a specialized prefetch request and does not require coherence protocol modification. Further, the lock is not speculatively acquired, only the cache line containing the lock variable is placed in the private cache of the predicted consumer. Speculative forwarding serves to hide the remote core's lock acquisition latency. SpecLock is distributed and all predictions are made locally at each core. We show that SpecLock captures 87% of predictable lock patterns correctly and improves performance by an average of 10% with 64 cores. SpecLock incurs a negligible overhead, with a 75% area reduction compared to past work. Compared to two state of the art methods, SpecLock provides a speedup of 8% and 4% respectively. Pooria M. Yaghini, George Michelogiannakis, Paul Gratz |
ICCD | 3 |
| 2019 | Perceptron-based prefetch filteringabstractHardware prefetching is an effective technique for hiding cache miss latencies in modern processor designs. Prefetcher performance can be characterized by two main metrics that are generally at odds with one another: coverage, the fraction of baseline cache misses which the prefetcher brings into the cache; and accuracy, the fraction of prefetches which are ultimately used. An overly aggressive prefetcher may improve coverage at the cost of reduced accuracy. Thus, performance may be harmed by this over-aggressiveness because many resources are wasted, including cache capacity and bandwidth. An ideal prefetcher would have both high coverage and accuracy. Eshan Bhatia, Gino Chacon, Seth H. Pugsley, Elvira Teran, Paul Gratz, Daniel A. Jiménez |
ISCA | 5 |
| 2019 | SWAP: Synchronized Weaving of Adjacent Packets for Network Deadlock ResolutionabstractAn interconnection network forms the communication backbone in both on-chip and off-chip systems. In networks, congestion causes packets to be blocked. Indefinite blocking can occur if cyclic dependencies exist, leading to deadlock. All modern networks devote resources to either avoid deadlock by eliminating cyclic dependences or to detect and recover from it. Mayank Parasar, Natalie D. Enright Jerger, Paul Gratz, Joshua San Miguel, Tushar Krishna |
MICRO | 3 |
| 2019 | vNVML: An Efficient User Space Library for Virtualizing and Sharing Non-Volatile MemoriesabstractThe emerging non-volatile memory (NVM) has attractive characteristics such as DRAM-like, low-latency together with the non-volatility of storage devices. Recently, byte-addressable, memory bus-attached NVM has become available. This paper addresses the problem of combining a smaller, faster byte-addressable NVM with a larger, slower storage device, like SSD, to create the impression of a larger and faster byte-addressable NVM which can be shared across many applications. In this paper, we propose vNVML, a user space library for virtualizing and sharing NVM. vNVML provides for applications transaction like memory semantics that ensures write ordering and persistency guarantees across system failures. vNVML exploits DRAM for read caching, to enable improvements in performance and potentially to reduce the number of writes to NVM, extending the NVM lifetime. vNVML is implemented and evaluated with realistic workloads to show that our library allows applications to share NVM, both in a single O/S and when docker like containers are employed. The results from the evaluation show that vNVML incurs less than 10% overhead while providing the benefits of an expanded virtualized NVM space to the applications, allowing applications to safely share the virtual NVM. Chih-Chieh Chou, Jaemin Jung, A. L. Narasimha Reddy, Paul Gratz, Doug Voigt |
MSST | 4 |
| 2019 | GenMatcher: A Generic Clustering-Based Arbitrary Matching FrameworkabstractPacket classification methods rely upon packet content/header matching against rules. Thus, throughput of matching operations is critical in many networking applications. Further, with the advent of Software Defined Networking (SDN), efficient implementation of software approaches to matching are critical for the overall system performance. This article presents 1 GenMatcher, a generic, software-only, arbitrary matching framework for fast, efficient searches. The key idea of our approach is to represent arbitrary rules with efficient prefix-based tries. To support arbitrary wildcards, we rearrange bits within the rules such that wildcards accumulate to one side of the bitstring. Since many non-contiguous wildcards often remain, we use multiple prefix-based tries. The main challenge in this context is to generate efficient trie groupings and expansions to support all arbitrary rules. Finding an optimal mix of grouping and expansion is an NP-complete problem. Our contribution includes a novel, clustering-based grouping algorithm to group rules based upon their bit-level similarities. Our algorithm generates near-optimal trie groupings with low configuration times and provides significantly higher match throughput compared to prior techniques. Experiments with synthetic traffic show that our method can achieve a 58.9X speedup compared to the baseline on a single core processor under a given memory constraint. Ping Wang 0043, Luke McHale, Paul Gratz, Alexander Sprintson |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | Synchronized Progress in Interconnection Networks (SPIN): A New Theory for Deadlock FreedomabstractOne of the most fundamental design challenges in any interconnection network is that of routing deadlocks. A deadlock is a cyclic dependence between buffers that renders forward progress impossible. Deadlocks are a necessary evil and almost every on-chip/HPC network today avoids it either via routing restrictions across physical channels (Dally's Theory) or with at least one escape virtual channel (Duato's Theory). This ensures that a cyclic dependence between buffers is never created in the first place. Moreover, each solution is tied to a specific topology, requiring an updated policy if the topology were to change. Alternately, solutions have also been proposed to reserve certain resources (buffers) and allocate them only upon detection of a deadlock, thereby breaking the dependence chain and recovering from the deadlock. Unfortunately, all these approaches fundamentally lead to a loss in available bandwidth due to routing restrictions or buffer resource usage restrictions. In this work, we challenge the theoretical notion of viewing deadlocks as a lack of routing resource (buffers) problem that every solution to date is based on. We argue that a deadlock can in fact be considered as a lack of coordination between distributed entities. We prove that orchestrating a forward movement of every flit in the deadlocked ring at exactly the same time, which we call a spin, can guarantee forward progress and eventually lead to deadlock resolution with a bounded number of spins. We name this novel theory as SPIN (Synchronized Progress in Interconnection Networks). SPIN eliminates the need for virtual channels to achieve deadlock freedom thereby enabling fully adaptive routing with only one buffer per message class. We illustrate this capability by designing FAvORS, a novel truly one VC fully-adaptive routing algorithm. We also present a low-cost distributed implementation of SPIN and compare it against state-of-the-art deadlock avoidance/recovery schemes. SPIN provides up to 80% higher throughput, 52% lower area and 50% lower power for an on-chip 64-core mesh, and up to 83% higher throughput, 53% lower area and 55% lower power for an off-chip 1024-node dragon-fly. Aniruddh Ramrakhyani, Paul Gratz, Tushar Krishna |
ISCA | 2 |
| 2018 | SDPR: Improving Latency and Bandwidth in On-Chip Interconnect Through Simultaneous Dual-Path RoutingabstractNetworks-on-chips (NoCs) are gaining in popularity as replacement for shared medium interconnects in chip-multiprocessors (CMPs) and multiprocessor systems-on-chips, and their performance becoming essential to system performance. There have been emerging studies to achieve better power/energy efficiency without performance degradation on NoCs. However, there are still non-negligible latency issues caused by the mechanism of power efficient approaches. To alleviate the latency problem and to transfer data efficiently with the high utilization of interconnect resources, we propose an on-chip network architecture that improves latency and bandwidth. Increasing the data/link widths across the network may considerably resolve this problem but is a costly proposition both in terms of device area and of power. Alternatively, we propose a dual-path router architecture that efficiently exploits path diversity to attain low latency without significant hardware overhead. By: 1) doubling the number of injection and ejection ports; 2) splitting packets into two halves; 3) recomposing routing policy to support path diversity; and 4) provisioning the network hardware design, we can considerably enhance network resource utilization to achieve much higher performance in latency. The proposed simultaneous dual-path routing (SDPR) scheme outperformed the conventional dimension order routing (DOR) technique across synthetic workloads by 31%-40% in average latency and up to a 100% improvement in throughput performance running on a 49-core CMP. Our synthesizable model for the SDPR router and network provides accurate power and area reports. According to the synthesis reports, SDPR incurs insignificant overhead compared to the baseline XY DOR router. Yoon Seok Yang, Hrishikesh Deshpande, Gwan S. Choi, Paul Gratz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Kill the Program Counter: Reconstructing Program Behavior in the Processor Cache HierarchyabstractData prefetching and cache replacement algorithms have been intensively studied in the design of high performance microprocessors. Typically, the data prefetcher operates in the private caches and does not interact with the replacement policy in the shared Last-Level Cache (LLC). Similarly, most replacement policies do not consider demand and prefetch requests as different types of requests. In particular, program counter (PC)-based replacement policies cannot learn from prefetch requests since the data prefetcher does not generate a PC value. PC-based policies can also be negatively affected by compiler optimizations. In this paper, we propose a holistic cache management technique called Kill-the-PC (KPC) that overcomes the weaknesses of traditional prefetching and replacement policy algorithms. KPC cache management has three novel contributions. First, a prefetcher which approximates the future use distance of prefetch requests based on its prediction confidence. Second, a simple replacement policy provides similar or better performance than current state-of-the-art PC-based prediction using global hysteresis. Third, KPC integrates prefetching and replacement policy into a whole system which is greater than the sum of its parts. Information from the prefetcher is used to improve the performance of the replacement policy and vice-versa. Finally, KPC removes the need to propagate the PC through entire on-chip cache hierarchy while providing a holistic cache management approach with better performance than state-of-the-art PC-, and non-PC-based schemes. Our evaluation shows that KPC provides 8% better performance than the best combination of existing prefetcher and replacement policy for multi-core workloads. Jinchun Kim, Elvira Teran, Paul Gratz, Daniel A. Jiménez, Seth H. Pugsley, Chris Wilkerson |
ASPLOS | 3 |
| 2017 | Minimal exercise vector generation for reliability improvementabstractNegative Bias Temperature Instability (NBTI) is a prominent physical failure mechanism which severely degrades the performance of PMOS transistors whenever the voltage at the gate is negatively biased. It leads to catastrophic timing violations in critical circuits and a severe shortening of the overall operational lifetime of the entire system. To alleviate such damaging effects due to NBTI, we present PRITEXT, a novel technique which generates a minimal set of deterministic exercise vectors based on test generation techniques which inherently near-optimizes the bit patterns across each of the generated vectors; the end target being to exercise the critical paths of a device when dormant so as to achieve near-ideal NBTI stress reduction. We explore the design-space of our generated vectors and apply them to our test processor platform under differing sequences, where our evaluation under realistic benchmarks shows that PRITEXT leads to an average 4.99× and a maximum of 13.91× lifetime improvement using 9 generated vectors. In an attempt to reduce hardware overheads even further, we next propose a heuristic to further reduce the number of exercise vectors with minimum loss in lifetime improvement. P. Madhukar Reddy, Stavros Hadjitheophanous, Vassos Soteriou, Paul Gratz, Maria K. Michael |
IOLTS | 4 |
| 2016 | Path confidence based lookahead prefetchingabstractDesigning prefetchers to maximize system performance often requires a delicate balance between coverage and accuracy. Achieving both high coverage and accuracy is particularly challenging in workloads with complex address patterns, which may require large amounts of history to accurately predict future addresses. This paper describes the Signature Path Prefetcher (SPP), which offers effective solutions for three classic challenges in prefetcher design. First, SPP uses a compressed history based scheme that accurately predicts complex address patterns. Second, unlike other history based algorithms, which miss out on many prefetching opportunities when address patterns make a transition between physical pages, SPP tracks complex patterns across physical page boundaries and continues prefetching as soon as they move to new pages. Finally, SPP uses the confidence it has in its predictions to adaptively throttle itself on a per-prefetch stream basis. In our analysis, we find that SPP improves performance by 27.2% over a no-prefetching baseline, and outperforms the state-of-the-art Best Offset prefetcher by 6.4%. SPP does this with minimal overhead, operating strictly in the physical address space, and without requiring any additional processor core state, such as the PC. Jinchun Kim, Seth H. Pugsley, Paul Gratz, A. L. Narasimha Reddy, Chris Wilkerson, Zeshan Chishti |
MICRO | 3 |
| 2016 | Resource Sharing Centric Dynamic Voltage and Frequency Scaling for CMP Cores, Uncore, and MemoryabstractWith the breakdown of Dennard’s scaling over the past decade, performance growth of modern microprocessor design has largely relied on scaling core count in chip multiprocessors (CMPs). The challenge of chip power density, however, remains and demands new power management solutions. This work investigates a coordinated CMP systemwide Dynamic Voltage and Frequency Scaling (DVFS) policy centered around shared resource utilization. This approach represents a new angle on the problem, differing from the conventional core-workload-driven approaches. The key component of our work is per-core DVFS leveraging a technique similar to TCP Vegas congestion control from networking. This TCP Vegas–based DVFS can potentially identify the synergy between power reduction and performance improvement. Further, this work includes uncore (on-chip interconnect and shared last level cache) and main memory DVFS policies coordinated with the per-core DVFS policy. Full system simulations on PARSEC benchmarks show that our technique reduces total energy dissipation by over 47% across all benchmarks with less than 2.3% performance degradation. Our work also leads to 12% more energy savings compared to a prior work CMP DVFS policy. Jae-Yeon Won, Paul Gratz, Srinivas Shakkottai, Jiang Hu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | GCA: Global Congestion Awareness for Load Balance in Networks-on-ChipabstractAs modern CMPs scale to ever increasing core counts, Networks-on-Chip (NoCs) are emerging as an interconnection fabric, enabling communication between components. While NoCs provide high and scalable bandwidth, current routing algorithms, such as dimension-ordered routing, suffer from poor load balance, leading to reduced throughput and high latencies. Improving load balance, hence, is critical in future CMP designs where increased latency leads to wasted power and energy waiting for outstanding requests to resolve. Adaptive routing is a known technique to improve load balance, however, prior adaptive routing techniques either use local or regionally-aggregated information to form their routing decisions. This paper proposes a new, light-weight, adaptive routing algorithm for on-chip routers based on global link state and congestion information, Global Congestion Awareness (GCA). GCA uses a simple, low-complexity route calculation unit, to calculate paths to their destination without the myopia of local decisions, nor the aggregation of unrelated status information, found in prior designs. In particular GCA outperforms local adaptive routing by 26 percent, Regional Congestion Awareness (RCA) by 15 percent, and a recent competing adaptive routing algorithm, DAR, by 8 percent on average on realistic workloads. Mukund Ramakrishna, Vamsi Krishna Kodati, Paul Gratz, Alexander Sprintson |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Bandwidth-efficient on-chip interconnect designs for GPGPUsabstractModern computational workloads require abundant thread level parallelism (TLP), necessitating highly-parallel, many-core accelerators such as General Purpose Graphics Processing Units (GPGPUs). GPGPUs place a heavy demand on the on-chip interconnect between the many cores and a few memory controllers (MCs). Thus, traffic is highly asymmetric, impacting on-chip resource utilization and system performance. Here, we analyze the communication demands of typical GPGPU applications, and propose efficient Network-on-Chip (NoC) designs to meet those demands. We show that the proposed schemes improve performance by up to 64.7%. Compared to the best of class prior work, our VC monopolizing and partitioning schemes improve performance by 25%. Hyunjun Jang, Jinchun Kim, Paul Gratz, Ki Hwan Yum, Eun Jung Kim 0001 |
DAC | 3 |
| 2015 | A control-theoretic approach for energy efficient CPU-GPU subsystem in mobile platformsabstractThis paper presents a control-theoretic approach to optimize the energy consumption of integrated CPU and GPU subsystems for graphic applications. It achieves this via a dynamic management of the CPU and GPU frequencies. To this end, we first model the interaction between the GPU and CPU as a queuing system. Second, we formulate a Multi-Input-Multi-Output state-space closed loop control to ensure robustness and stability. We evaluated this control on an Intel Baytrail-based Android platform. Experimental evaluations show energy savings of 17.4% in the CPU-GPU subsystem with a low performance impact of 0.9%. David Kadjo, Raid Ayoub, Michael Kishinevsky, Paul Gratz |
DAC | 4 |
| 2015 | Energy-efficient implementations of GF (p) and GF(2m) elliptic curve cryptographyabstractWhile public-key cryptography is essential for secure communications, the energy cost of even the most efficient algorithms based on Elliptic Curve Cryptography (ECC) is prohibitive on many ultra-low energy devices such as sensor-network nodes and identification tags. Although an abundance of hardware acceleration techniques for ECC have been proposed in literature, little research has focused on understanding the energy benefits of these techniques. Therefore, we evaluate the energy cost of ECC on several different hardware/software configurations across a range of security levels. Our work comprehensively explores implementations of both GF(p) and GF(2m) ECC, demonstrating that GF(2m) provides a 1.31 to 2.11 factor improvement in energy efficiency over GF(p) on an extended RISC processor. We also show that including a 4KB instruction cache in our system can reduce the energy cost of ECC by as much as 30%. Furthermore, our GF(2m) coprocessor achieves a 2.8 to 3.61 factor improvement in energy efficiency compared to instruction set extensions and significantly outperforms prior work. Andrew D. Targhetta, Donald E. Owen, Francis L. Israel, Paul Gratz |
ICCD | 4 |
| 2015 | Clotho: Proactive wearout deceleration in Chip-Multiprocessor interconnectsabstractWith advancing process technology, Chip-Multiprocessors (CMPs) are experiencing ever worsening reliability due to prolonged operational stresses. The network-on-chip that interconnects the components of CMPs is especially vulnerable to such wearout-induced failure. To tackle this ominous threat we present Clotho, a novel, wearout-aware routing algorithm. Clotho continuously considers the stresses the on-chip interconnect experiences at runtime, along with temperature and fabrication process variation metrics, steering traffic away from locations that are most prone to Electromigration (EM)- and Hot-Carrier Injection (HCI)-induced wear. Under realistic workloads Clotho yields 66% and 8% average increases in mean time to failure for EM and HCI, respectively. Arseniy Vitkovskiy, Vassos Soteriou, Paul Gratz |
ICCD | 3 |
| 2015 | Having your cake and eating it too: Energy savings without performance loss through resource sharing driven power managementabstractTypically in computer systems, performance must be traded-off to achieve energy savings or, conversely, performance gains come with significant energy overhead. Here, we present a novel approach that can achieve synergistic energy-savings and performance gain in chip multiprocessors (CMPs). Our key observation is that per-core dynamic voltage/frequency scaling (DVFS) can be used as a client regulation mechanism for shared resources on-die. Based on this observation, we propose a new DVFS technique inspired by TCP Vegas, a congestion control protocol from the IP-networking domain. Full system simulations on PARSEC benchmarks show that our technique reduces total CMP energy dissipation by over 40% with a small performance improvement. Jae-Yeon Won, Paul Gratz, Srinivas Shakkottai, Jiang Hu 0001 |
ISLPED | 2 |
| 2015 | Wear-Aware Adaptive Routing for Networks-on-ChipsabstractChip-multiprocessors are facing worsening reliability due to prolonged operational stresses, with their tile-interconnecting Network-on-Chip (NoC) being especially vulnerable to wearout-induced failure. To tackle this ominous threat we present a novel wear-aware routing algorithm that continuously considers the stresses the NoC experiences at runtime, along with temperature and fabrication process variation metrics, steering traffic away from locations that are most prone to Electromigration (EM)- and Hot-Carrier Injection (HCI)-induced wear. Under realistic applications our wear-aware algorithm yields 66% and 8% average increases in mean-time-to-failure for EM and HCI, respectively. Arseni Vitkovski, Vassos Soteriou, Paul Gratz |
NOCS | 3 |
| 2015 | Use It or Lose It: Proactive, Deterministic Longevity in Future Chip MultiprocessorsabstractMoore's Law scaling continues to yield higher transistor density with each succeeding process generation, leading to today's many-core chip multiprocessors (CMPs) with tens or even hundreds of interconnected cores or tiles. Unfortunately, deep submicron CMOS process technology is marred by increasing susceptibility to wear. Prolonged operational stress gives rise to accelerated wearout and failure due to several physical failure mechanisms, including hot-carrier injection (HCI) and negative-bias temperature instability (NBTI). Each failure mechanism correlates with different usage-based stresses, all of which can eventually generate permanent faults. While the wearout of an individual core in many-core CMPs may not necessarily be catastrophic, a single fault in the interprocessor network-on-chip (NoC) fabric could render the entire chip useless, as it could lead to protocol-level deadlocks, or even partition away vital components such as the memory controller or other critical I/O. In this article, we study HCI- and NBTI-induced wear due to actual stresses caused by real workloads, applied onto the interconnect microarchitecture and develop a critical path model for NBTI-induced wearout. A key finding of this modeling is that, counter to prevailing wisdom, wearout in the CMP's on-chip interconnect is correlated with lack of load observed in the NoC routers rather than high load. We then develop a novel wearout-decelerating scheme in which routers under low load have their wear-sensitive components exercised without significantly impacting cycle time, pipeline depth, area, or power consumption of the overall router. A novel deterministic approach is proposed for the generation of appropriate exercise-mode data, ensuring design parameter targets are met. We subsequently show that the proposed design yields an ∼2,300× decrease in the rate of wear. Siva Bhanu Krishna Boga, Arseniy Vitkovskiy, Stavros Hadjitheophanous, Paul Gratz, Vassos Soteriou, Maria K. Michael |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2014 | ILP and TLP in shared memory applications: a limit studyabstractWith the breakdown of Dennard scaling, future processor designs will be at the mercy of power limits as Chip Multi-Processor (CMP) designs scale out to many-cores. It is critical, therefore, that future CMPs be optimally designed in terms of performance efficiency with respect to power. A characterization analysis of future workloads is imperative to ensure maximum returns of performance per Watt consumed. Hence, a detailed analysis of emerging workloads is necessary to understand their characteristics with respect to hardware in terms of power and performance tradeoffs. In this paper, we conduct a limit study simultaneously analyzing the two dominant forms of parallelism exploited by modern computer architectures: Instruction Level Parallelism (ILP) and Thread Level Parallelism (TLP). This study gives insights into the upper bounds of performance that future architectures can achieve. Furthermore it identifies the bottlenecks of emerging workloads. To the best of our knowledge, our work is the first study that combines the two forms of parallelism into one study with modern applications. We evaluate the PARSEC multithreaded benchmark suite using a specialized trace-driven simulator. We make several contributions describing the high-level behavior of next-generation applications. For example, we show these applications contain up to a factor of 929X more ILP than what is currently being extracted from real machines. We then show the effects of breaking the application into increasing numbers of threads (exploiting TLP), instruction window size, realistic branch prediction, realistic memory latency, and thread dependencies on exploitable ILP. Our examination shows that theses benchmarks differed vastly from one another. As a result, we expect no single, homogeneous, micro-architecture will work optimally for all, arguing for reconfigurable, heterogeneous designs. Ehsan Fatehi, Paul Gratz |
PACT | 2 |
| 2014 | Up by their bootstraps: Online learning in Artificial Neural Networks for CMP uncore power managementabstractWith increasing core counts in Chip Multi-Processor (CMP) designs, the size of the on-chip communication fabric and shared Last-Level Caches (LLC), which we term uncore here, is also growing, consuming as much as 30% of die area and a significant portion of chip power budget. In this work, we focus on improving the uncore energy-efficiency using dynamic voltage and frequency scaling. Previous approaches are mostly restricted to reactive techniques, which may respond poorly to abrupt workload and uncore utility changes. We find, however, there are predictable patterns in uncore utility which point towards the potential of a proactive approach to uncore power management. In this work, we utilize artificial intelligence principles to proactively leverage uncore utility pattern prediction via an Artificial Neural Network (ANN). ANNs, however, require training to produce accurate predictions. Architecting an efficient training mechanism without a priori knowledge of the workload is a major challenge. We propose a novel technique in which a simple Proportional Integral (PI) controller is used as a secondary classifier during ANN training, dynamically pulling the ANN up by its bootstraps to achieve accurate predictions. Both the ANN and the PI controller, then, work in tandem once the ANN training phase is complete. The advantage of using a PI controller to initially train the ANN is a dramatic acceleration of the ANN's initial learning phase. Thus, in a real system, this scenario allows quick power-control adaptation to rapid application phase changes and context switches during execution. We show that the proposed technique produces results comparable to those of pure offline training without a need for prerecorded training sets. Full system simulations using the PARSEC benchmark suite show that the bootstrapped ANN improves the energy-delay product of the uncore system by 27% versus existing state-of-the-art methodologies. Jae-Yeon Won, Paul Gratz, Jiang Hu 0001, Vassos Soteriou |
HPCA | 3 |
| 2014 | Stochastic Pre-classification for SDN Data Plane MatchingabstractThe Software Defined Networking (SDN) approach has numerous advantages, including the ability to program the network through simple abstractions, provide a centralized view of network state, and respond to changing network conditions. One of the main challenges in designing SDN enabled switches is efficient packet classification in the data plane. As the complexity of SDN applications increases, the data plane becomes more susceptible to Denial of Service (DoS) attacks, which can result in increased delays and packet loss. Accordingly, there is a strong need for network architectures that operate efficiently in the presence of malicious traffic. In particular, there is a need to protect authorized flows from DoS attacks. In this work we utilize a probabilistic data structure to pre-classify traffic with the aim of decoupling likely legitimate traffic from malicious traffic by leveraging the locality of packet flows. We validate our approach by examining a fundamental SDN application: software defined network firewall. For this application, our architecture dramatically reduces the impact of unknown/malicious flows on established/legitimate flows. We explore the effect of stochastic pre-classification in prioritizing data plane classification. We show how pre-classification can be used to increase the effective Quality of Service (QoS) for established flows and reduce the impact of adversarial traffic. Luke McHale, C. Jasson Casey, Paul Gratz, Alexander Sprintson |
ICNP | 3 |
| 2014 | The design space of ultra-low energy asymmetric cryptographyabstractThe energy cost of asymmetric cryptography, a vital component of modern secure communications, inhibits its wide spread adoption within the ultra-low energy regimes such as Implantable Medical Devices (IMDs), Wireless Sensor Networks (WSNs), and Radio Frequency Identification tags (RFIDs). Consequently, a gamut of hardware/software acceleration techniques exists to alleviate this energy burden. In this paper, we explore this design space, estimating the energy consumption for three levels of acceleration across the commercial security spectrum. First we examine an efficient baseline architecture centered around a pipelined RISC processor. We then include simple, yet beneficial instruction set extensions to our microarchitecture and evaluate the improvement in terms of energy per operation compared to baseline. Finally, we introduce a novel, dedicated accelerator to our microarchitecture and measure the energy per operation against the baseline and the ISA extensions. For ISA extensions, we show between 1.28 to 1.41 factor improvement in energy efficiency over baseline, while for full acceleration we demonstrate a 4.36 to 6.45 factor improvement. Andrew D. Targhetta, Donald E. Owen, Paul Gratz |
ISPASS | 3 |
| 2014 | B-Fetch: Branch Prediction Directed Prefetching for Chip-MultiprocessorsabstractFor decades, the primary tools in alleviating the "Memory Wall" have been large cache hierarchies and dataprefetchers. Both approaches, become more challenging in modern, Chip-multiprocessor (CMP) design. Increasing the last-level cache (LLC) size yields diminishing returns in terms of performance per Watt, given VLSI power scaling trends, this approach becomes hard to justify. These trends also impact hardware budgets for prefetchers. Moreover, in the context of CMPs running multiple concurrent processes, prefetching accuracy is critical to prevent cache pollution effects. These concerns point to the need for a light-weight prefetcher with high accuracy. Existing data prefetchers may generally be classified as low-overhead and low accuracy (Next-n, Stride, etc.) or high-overhead and high accuracy (STeMS, ISB). Wepropose B-Fetch: a data prefetcher driven by branch prediction and effective address value speculation. B-Fetch leverages control flow prediction to generate an expected future path of the executing application. It then speculatively computes the effective address of the load instructions along that path based upon a history of past register transformations. Detailed simulation using a cycle accurate simulator shows a geometric mean speedup of 23.4% for single-threaded workloads, improving to 28.6% for multi-application workloads over a baseline system without prefetching. We find that B-Fetch outperforms an existing "best-of-class" light-weight prefetcher under single-threaded and multi programmed workloads by 9% on average, with 65% less storage overhead. David Kadjo, Jinchun Kim, Prabal Sharma, Reena Panda, Paul Gratz, Daniel A. Jiménez |
MICRO | 5 |
| 2014 | STORM: A Simple Traffic-Optimized Router Microarchitecture for Networks-on-ChipabstractNetworks-on-Chip (NoCs) offer a scalable means of on-chip communication for future many-core chips. This work explores NoC router microarchitectures which leverage traffic pattern biases and imbalances to reduce latency and improve throughput. It introduces STORM, a new, low-latency, fair, highth-roughput NoC router design, customized for the traffic seen in a two-dimensional mesh network employing dimension-order routing. Compared to a baseline NoC router with equivalent buffer resources, STORM offers single cycle operation and reduced cycle time (17% less than the baseline on 45nm CMOS). This design yields a higher overall network saturation throughput (13% higher than the baseline) in an 8x8 2D mesh network for uniform random traffic. STORM also reduces packet latencies under realistic workloads by 36% on average. Shalimar Rasheed, Paul Gratz, Srinivas Shakkottai, Jiang Hu 0001 |
NOCS | 2 |
| 2014 | Spatial Locality Speculation to Reduce Energy in Chip-Multiprocessor Networks-on-ChipabstractAs processor chips become increasingly parallel, an efficient communication substrate is critical for meeting performance and energy targets. In this work, we target the root cause of network energy consumption through techniques that reduce link and router-level switching activity. We specifically focus on memory subsystem traffic, as it comprises the bulk of NoC load in a CMP. By transmitting only the flits that contain words predicted useful using a novel spatial locality predictor, our scheme seeks to reduce network activity. We aim to further lower NoC energy through microarchitectural mechanisms that inhibit datapath switching activity for unused words in individual flits. Using simulation-based performance studies and detailed energy models based on synthesized router designs and different link wire types, we show that 1) the prediction mechanism achieves very high accuracy, with an average rate of false-unused prediction of just 2.5 percent; 2) the combined NoC energy savings enabled by the predictor and microarchitectural support is 36 percent, on average, and up to 57 percent in the best case; and 3) there is no system performance penalty as a result of this technique. Boris Grot, Paul Gratz, Daniel A. Jiménez |
IEEE Trans. Computers | 3 |
| 2014 | LumiNOC: A Power-Efficient, High-Performance, Photonic Network-on-ChipabstractTo meet energy-efficient performance demands, the computing industry has moved to parallel computer architectures, such as chip multiprocessors (CMPs), internally interconnected via networks-on-chip (NoC) to meet growing communication needs. Achieving scaling performance as core counts increase to the hundreds in future CMPs, however, will require high performance, yet energy-efficient interconnects. Silicon nanophotonics is a promising replacement for electronic on-chip interconnect due to its high bandwidth and low latency, however, prior techniques have required high static power for the laser and ring thermal tuning. We propose a novel nano-photonic NoC (PNoC) architecture, LumiNOC, optimized for high performance and power-efficiency. This paper makes three primary contributions: a novel, nanophotonic architecture which partitions the network into subnets for better efficiency; a purely photonic, in-band, distributed arbitration scheme; and a channel sharing arrangement utilizing the same waveguides and wavelengths for arbitration as data transmission. In a 64-node NoC under synthetic traffic, LumiNOC enjoys 50% lower latency at low loads and ~40% higher throughput per Watt on synthetic traffic, versus other reported PNoCs. LumiNOC reduces latencies ~40% versus an electrical 2-D mesh NoCs on the PARSEC shared-memory, multithreaded benchmark suite. Mark Browning, Paul Gratz, Samuel Palermo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | WaveSync: Low-Latency Source-Synchronous Bypass Network-on-Chip ArchitectureabstractWaveSync is a network-on-chip architecture for a globally asynchronous locally-synchronous (GALS) design. The WaveSync design facilitates low-latency communication leveraging the source-synchronous clock sent along with the data to time components in the datapath of a downstream router, reducing the number of synchronizations needed. WaveSync accomplishes this by partitioning the router components at each node into different clock domains, each synchronized with one of the orthogonal incoming source-synchronous clocks in a GALS 2D mesh network. The data and clock subsequently propagate through each node/router synchronously until the destination is reached, regardless of the number of hops this may take. As long as the data travels in the path of clock propagation and no congestion is encountered, it will be propagated without latching as if in a long combinatorial path, with both the clock and the data accruing delay at the same rate. The result is that the need for synchronization between the mesochronous nodes and/or the asynchronous control associated with the typical GALS network is completely eliminated. To further reduce the latency overhead of synchronization, for those occasions when synchronization is still required (when a flit takes a turn or arrives at the destination), we propose a novel less-than-one-cycle synchronizer. The proposed WaveSync network outperforms conventional GALS networks by 87--90% in average latency, synthesized using a 45nm CMOS library. Yoon Seok Yang, Reeshav Kumar, Gwan S. Choi, Paul Gratz |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2013 | Dynamic voltage and frequency scaling for shared resources in multicore processor designsabstractAs the core count in processor chips grows, so do the on-die, shared resources such as on-chip communication fabric and shared cache, which are of paramount importance for chip performance and power. This paper presents a method for dynamic voltage/frequency scaling of networks-on-chip and last level caches in multicore processor designs, where the shared resources form a single voltage/frequency domain. Several new techniques for monitoring and control are developed, and validated through full system simulations on the PARSEC benchmarks. These techniques reduce energy-delay product by 56% compared to a state-of-the-art prior work. Zheng Xu 0006, Paul Gratz, Jiang Hu 0001, Michael Kishinevsky, Ümit Y. Ogras, Raid Ayoub |
DAC | 4 |
| 2013 | Stochastic Pre-Classification for Software Defined FirewallsabstractFirewalls are ubiquitous security functions and exist in almost all network connected devices whether protecting host stacks or providing transient packet filtering. Firewall performance, which is a key ingredient for network performance, can be greatly degraded by traffic crafted to exploit its filtering algorithms. These attacks can greatly reduce the Quality of Service (QoS) received by existing authorized flows in the firewall. This paper proposes a novel architecture that decouples this linkage between authorized flow QoS and adversarial traffic, marginalizing disruption caused by unauthorized flows, and ultimately improving overall performance of software defined firewalls. We show substantial improvements in throughput, packet loss, and latency over baseline software defined firewalls with varying ratios of attack traffic. All results are obtained using the cycle accurate architecture simulator gem5, and Internet packet traces obtained from 10 Gbps interfaces of core Internet routers. Pritha Ghoshal, C. Jasson Casey, Paul Gratz, Alexander Sprintson |
ICCCN | 3 |
| 2013 | Power gating with block migration in chip-multiprocessor last-level cachesabstractWe propose a novel technique to significantly reduce the leakage energy of last level caches while mitigating any significant performance impact. In general, cache blocks are not ordered by their temporal locality within the sets; hence, simply power gating off a partition of the cache, as done in previous studies, may lead to considerable performance degradation. We propose a solution that migrates the high temporal locality blocks to facilitate power gating, where blocks likely to be used in the future are migrated from the partition being shutdown to the live partition at a negligible performance impact and hardware overhead. Our detailed simulations show energy savings of 66% at low performance degradation of 2.16%. David Kadjo, Paul Gratz, Jiang Hu 0001, Raid Ayoub |
ICCD | 3 |
| 2013 | Use it or lose it: wear-out and lifetime in future chip multiprocessorsabstractMoore's Law scaling is continuing to yield even higher transistor density with each succeeding process generation, leading to today's multi-core Chip Multi-Processors (CMPs) with tens or even hundreds of interconnected cores or tiles. Unfortunately, deep sub-micron CMOS process technology is marred by increasing susceptibility to wearout. Prolonged operational stress gives rise to accelerated wearout and failure, due to several physical failure mechanisms, including Hot Carrier Injection (HCI) and Negative Bias Temperature Instability (NBTI). Each failure mechanism correlates with different usage-based stresses, all of which can eventually generate permanent faults. While the wearout of an individual core in many-core CMPs may not necessarily be catastrophic for the system, a single fault in the inter-processor Network-on-Chip (NoC) fabric could render the entire chip useless, as it could lead to protocol-level deadlocks, or even partition away vital components such as the memory controller or other critical I/O. In this paper, we develop critical path models for HCI- and NBTI-induced wear due to the actual stresses caused by real workloads, applied onto the interconnect microarchitecture. A key finding from this modeling being that, counter to prevailing wisdom, wearout in the CMP on-chip interconnect is correlated with lack of load observed in the NoC routers, rather than high load. We then develop a novel wearout-decelerating scheme in which routers under low load have their wearout-sensitive components exercised, without significantly impacting cycle time, pipeline depth, area or power consumption of the overall router. We subsequently show that the proposed design yields a 13.8x-65x increase in CMP lifetime. Arseniy Vitkovskiy, Paul Gratz, Vassos Soteriou |
MICRO | 3 |
| 2013 | GCA: Global congestion awareness for load balance in Networks-on-ChipabstractAs modern CMPs scale to ever increasing core counts, Networks-on-Chip (NoCs) are emerging as an interconnection fabric, enabling communication between components. While NoCs provide high and scalable bandwidth, current routing algorithms, such as dimension-ordered routing, suffer from poor load balance, leading to reduced throughput and high latencies. Improving load balance, hence, is critical in future CMP designs where increased latency leads to wasted power and energy waiting for outstanding requests to resolve. Adaptive routing is a known technique to improve load balance, however, prior adaptive routing techniques either use local or regionally aggregated information to form their routing decisions. This paper proposes a new, light-weight, adaptive routing algorithm for on-chip routers based on global link state and congestion information, Global Congestion Awareness (GCA). GCA uses a simple, low-complexity route calculation unit, to calculate paths to their destination without the myopia of local decisions, nor the aggregation of unrelated status information, found in prior designs. In particular GCA outperforms local adaptive routing by 26%, Regional Congestion Awareness (RCA) by 15%, and a recent competing adaptive routing algorithm, DAR, by 8% on average on realistic workloads. Mukund Ramakrishna, Paul Gratz, Alexander Sprintson |
NOCS | 2 |
| 2013 | ARI: Adaptive LLC-memory traffic managementabstractDecreasing the traffic from the CPU LLC to main memory is a very important issue in modern systems. Recent work focuses on cache misses, overlooking the impact of writebacks on the total memory traffic, energy consumption, IPC, and so forth. Policies that foster a balanced approach, between reducing write traffic to memory and improving miss rates, can increase overall performance and improve energy efficiency and memory system lifetime for NVM memory technology, such as phase-change memory (PCM). We propose Adaptive Replacement and Insertion (ARI), an adaptive approach to last-level CPU cache management, optimizing the two parameters (miss rate and writeback rate) simultaneously. Our specific focus is to reduce writebacks as much as possible while maintaining or improving the miss rate relative to conventional LRU replacement policy. ARI reduces LLC writebacks by 33%, on average, while also decreasing misses by 4.7%, on average. In a typical system, this boosts IPC by 4.9%, on average, while decreasing energy consumption by 8.9%. These results are achieved with minimal hardware overheads. Viacheslav V. Fedorov, Sheng Qiu, A. L. Narasimha Reddy, Paul Gratz |
ACM Trans. Archit. Code Optim. | 4 |
| 2013 | In-network monitoring and control policy for DVFS of CMP networks-on-chip and last level cachesabstractIn chip design today and for a foreseeable future, the last-level cache and on-chip interconnect is not only performance critical but also a substantial power consumer. This work focuses on employing dynamic voltage and frequency scaling (DVFS) policies for networks-on-chip (NoC) and shared, distributed last-level caches (LLC). In particular, we consider a practical system architecture where the distributed LLC and the NoC share a voltage/frequency domain that is separate from the core domain. This architecture enables the control of the relative speed between the cores and memory hierarchy without introducing synchronization delays within the NoC. DVFS for this architecture is more complex than individual link/core-based DVFS since it involves spatially distributed monitoring and control. We propose an average memory access time (AMAT)-based monitoring technique and integrate it with DVFS based on PID control theory. Simulations on PARSEC benchmarks yield a 27% energy savings with a negligible impact on system performance. Zheng Xu 0006, Paul Gratz, Jiang Hu 0001, Michael Kishinevsky, Ümit Y. Ogras |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | LumiNOC: a power-efficient, high-performance, photonic network-on-chip for future parallel architecturesabstractAchieving scaling performance as core counts increase to the hundreds in future chip-multi-processors (CMPs) requires high performing, yet energy-efficient interconnects. Silicon nanophotonics is a promising replacement for electronic on-chip interconnect due to its high bandwidth and low latency, however, prior techniques have required high static power for the laser and ring thermal tuning. We propose a novel nano-photonic NoC architecture, LumiNOC, optimized for high performance and power-efficiency. In a 64-node NoC under synthetic traffic, LumiNOC enjoys 50% lower latency at low loads and 40% higher throughput per Watt on synthetic traffic, versus other reported photonic NoCs. LumiNOC reduces latencies 40% versus an electrical 2D mesh NoCs on the PARSEC shared memory, multithreaded benchmark suite. Mark Browning, Paul Gratz, Samuel Palermo |
PACT | 3 |
| 2012 | WaveSync: A low-latency source synchronous bypass network-on-chip architectureabstractWaveSync is a low-latency focused, network-on-chip architecture for globally-asynchronous locally-synchronous (GALS) designs. WaveSync facilitates low-latency communication leveraging the source-synchronous clock sent with the data, to time components in the downstream routers data-path to reduce the number of synchronizations needed. WaveSync accomplishes this by partitioning the router components at each node into different clock-domains, each synchronized with one of the the orthogonal incoming source synchronous clocks in a GALS 2D mesh network. The data and clock subsequently propagate through each node/router, synchronously, until the destination is reached, regardless of the number of hops it may take. As long as the data travel in the path of clock propagation, and no congestion is encountered, it will be propagated without latching, as if in a long-combinatorial path, with both the clock and the data accruing delay at the same rate. The result is that the need for synchronization between the mesochronous nodes and/or the asynchronous control associated with typical GALS network is completely eliminated. The proposed WaveSync network outperforms conventional GALS networks by 87-90% in average nanosecond latency with 1.8-6.5 times more throughput across synthetic traffic patterns and SPLASH-2 benchmark suite. Yoon Seok Yang, Reeshav Kumar, Gwan S. Choi, Paul Gratz |
ICCD | 4 |
| 2012 | Exploiting path diversity for low-latency and high-bandwidth with the dual-path NoC routerabstractNetworks-on-Chips are gaining in popularity as replacement for shared medium interconnects in chip-multiprocessors (CMPs) and multiprocessor systems-on-chips (MPSoCs), and their performance becoming essential to system performance. We propose a dual-path router architecture that efficiently exploits path diversity to attain low latency and high throughput without significant hardware overhead. By 1) doubling the number of injection and ejection ports, 2) splitting packets into two halves, 3) recomposing routing policy to support path diversity, and 4) provisioning the network hardware design, we can significantly improve network resource utilization to achieve much higher throughput and lower latencies. Results show that the proposed dual-path router improves saturation bandwidth by 29% on uniform random synthetic traffic, while achieving a reduction in average packet latency of 31% and 17% for uniform random synthetic traffic and video benchmarks respectively. Yoon Seok Yang, Hrishikesh Deshpande, Gwan S. Choi, Paul Gratz |
ISCAS | 4 |
| 2012 | In-network Monitoring and Control Policy for DVFS of CMP Networks-on-Chip and Last Level CachesabstractIn chip design today and for a foreseeable future, on-chip communication is not only a performance bottleneck but also a substantial power consumer. This work focuses on employing dynamic voltage and frequency scaling (DVFS) policies for networks-on-chip (NoC) and shared, distributed last-level caches (LLC). In particular, we consider a practical system architecture where the distributed LLC and the NoC share a voltage/frequency domain which is separate from the core domain. This architecture enables controlling the relative speed between the cores and memory hierarchy without introducing synchronization delays within the NoC. DVFS for this architecture is more difficult than individual link/core-based DVFS since it involves spatially distributed monitoring and control. We propose an average memory access time (AMAT)-based monitoring technique and integrate it with DVFS based on PID control theory. Simulations on PARSEC benchmarks yield a 33% dynamic energy savings with a negligible impact on system performance. Zheng Xu 0006, Paul Gratz, Jiang Hu 0001, Michael Kishinevsky, Ümit Y. Ogras |
NOCS | 4 |
| 2011 | Reducing network-on-chip energy consumption through spatial locality speculationabstractAs processor chips become increasingly parallel, an efficient communication substrate is critical for meeting performance and energy targets. In this work, we target the root cause of network energy consumption through techniques that reduce link and router-level switching activity. We specifically focus on memory subsystem traffic, as it comprises the bulk of NoC load in a CMP. By transmitting only the flits that contain words predicted useful using a novel spatial locality predictor, our scheme seeks to reduce network activity. We aim to further lower NoC energy through microarchitectural mechanisms that inhibit datapath switching activity for unused words in individual flits. Using simulation-based performance studies and detailed energy models based on synthesized router designs and different link wire types, we show that (a) the prediction mechanism achieves very high accuracy, with an average misprediction rate of just 2.5%; (b) the combined NoC energy savings enabled by the predictor and microarchitectural support are 35% on average and up to 60% in the best case; and (c) the performance impact of these energy optimizations is negligible. Pritha Ghoshal, Boris Grot, Paul Gratz, Daniel A. Jiménez |
NOCS | 4 |
| 2011 | Asynchronous Bypass Channels for Multi-Synchronous NoCs: A Router Microarchitecture, Topology, and Routing AlgorithmabstractNetwork-on-chip (NoC) designs have emerged as a replacement for traditional shared-bus designs for on-chip communication. As with all current very large scale integration designs, however, reducing power consumption in NoCs is a critical challenge. One approach to reduce power consumption is to dynamically scale the voltage and frequency of each network node or groups of nodes (DVFS). Another approach is to replace the balanced clock tree with a globally-asynchronous, locally-synchronous (GALS) clocking scheme. In both DVFS and GALS designs, the chip as a whole is multi-synchronous. As the NoCs interconnecting those nodes must communicate across these clock domain boundaries, they tend to have high latencies as packets must be synchronized at the intermediate nodes. In this paper, we propose a novel router microarchitecture which offers superior performance with respect to typical synchronizing router designs for multi-synchronous networks. Our approach features asynchronous bypass channels which allow flit traversal of intermediate nodes within the network without the latching or synchronization overheads of typical designs. We also propose a new network topology and routing algorithm that leverage the advantages of the bypass channel offered by our router design. We present a detailed analysis of design decisions which affect the performance of the asynchronous bypass channel network. Our experiments show that our design improves the performance of a conventional synchronizing design with similar resources by up to 26% at low loads and increases saturation throughput by up to 50% for a uniform random traffic. Tushar N. K. Jain, Mukund Ramakrishna, Paul Gratz, Alexander Sprintson, Gwan S. Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Asynchronous Bypass Channels: Improving Performance for Multi-synchronous NoCsabstractNetworks-on-Chip (NoC) have emerged as a replacement for traditional shared-bus designs for on-chip communications. As with all current VLSI design, however, reducing power consumption in NoCs is a critical challenge. One approach to reduce power is to dynamically scale the voltage and frequency of each network node or groups of nodes (DVFS). Another approach to reduce power consumption is to replace the balanced clock tree with a globally-asynchronous, locally-synchronous (GALS) clocking scheme. NoCs implemented with either of these schemes, however, tend to have high latencies as packets must be synchronized at intermediate nodes between source and destination. In this paper, we propose a novel router microarchitecture which offers superior performance versus typical synchronizing router designs. Our approach features Asynchronous Bypass Channels (ABCs) at intermediate nodes thus avoiding synchronization delay. We also propose several new network topology and routing algorithm that leverage the advantages of the bypass channel offered by our router design. Our experiments show that our design improves the performance of a conventional synchronizing design with similar resources by up to 26% at low loads and increases saturation throughput by up to 50%. Tushar N. K. Jain, Paul Gratz, Alexander Sprintson, Gwan S. Choi |
NOCS | 2 |
| 2009 | An evaluation of the TRIPS computer systemabstractThe TRIPS system employs a new instruction set architecture (ISA) called Explicit Data Graph Execution (EDGE) that renegotiates the boundary between hardware and software to expose and exploit concurrency. EDGE ISAs use a block-atomic execution model in which blocks are composed of dataflow instructions. The goal of the TRIPS design is to mine concurrency for high performance while tolerating emerging technology scaling challenges, such as increasing wire delays and power consumption. This paper evaluates how well TRIPS meets this goal through a detailed ISA and performance analysis. We compare performance, using cycles counts, to commercial processors. On SPEC CPU2000, the Intel Core 2 outperforms compiled TRIPS code in most cases, although TRIPS matches a Pentium 4. On simple benchmarks, compiled TRIPS code outperforms the Core 2 by 10% and hand-optimized TRIPS code outperforms it by factor of 3. Compared to conventional ISAs, the block-atomic model provides a larger instruction window, increases concurrency at a cost of more instructions executed, and replaces register and memory accesses with more efficient direct instruction-to-instruction communication. Our analysis suggests ISA, microarchitecture, and compiler enhancements for addressing weaknesses in TRIPS and indicates that EDGE architectures have the potential to exploit greater concurrency in future technologies. Mark Gebhart, Bertrand A. Maher, Katherine E. Coons, Jeffrey R. Diamond, Paul Gratz, Mario Marino, Nitya Ranganathan, Behnam Robatmili, Aaron Smith, James H. Burrill, Stephen W. Keckler, Doug Burger, Kathryn S. McKinley |
ASPLOS | 5 |
| 2008 | Regional congestion awareness for load balance in networks-on-chipabstractInterconnection networks-on-chip (NOCs) are rapidly replacing other forms of interconnect in chip multiprocessors and system-on-chip designs. Existing interconnection networks use either oblivious or adaptive routing algorithms to determine the route taken by a packet to its destination. Despite somewhat higher implementation complexity, adaptive routing enjoys better fault tolerance characteristics, increases network throughput, and decreases latency compared to oblivious policies when faced with non-uniform or bursty traffic. However, adaptive routing can hurt performance by disturbing any inherent global load balance through greedy local decisions. To improve load balance in adapting routing, we propose Regional Congestion Awareness (RCA), a lightweight technique to improve global network balance. Instead of relying solely on local congestion information, RCA informs the routing policy of congestion in parts of the network beyond adjacent routers. Our experiments show that RCA matches or exceeds the performance of conventional adaptive routing across all workloads examined, with a 16% average and 71% maximum latency reduction on SPLASH-2 benchmarks running on a 49-core CMP. Compared to a baseline adaptive router, RCA incurs a negligible logic and modest wiring overhead. Paul Gratz, Boris Grot, Stephen W. Keckler |
HPCA | 1 |
| 2007 | Implementation and Evaluation of a Dynamically Routed Processor Operand NetworkabstractMicroarchitecturally integrated on-chip networks, or micronets, are candidates to replace busses for processor component interconnect in future processor designs. For micronets, tight coupling between processor microarchitecture and network architecture is one of the keys to improving processor performance. This paper presents the design, implementation and evaluation of the TRIPS operand network (OPN). The TRIPS OPN is a 5times5, dynamically routed, 2D mesh micronet that is integrated into the TRIPS microprocessor core. The TRIPS OPN is used for operand passing, register file I/O, and primary memory system I/O. We discuss in detail the OPN design, including the unique features that arise from its integration with the processor core, such as its connection to the execution unit's wakeup pipeline and its in flight mis-speculated traffic removal. We then evaluate the performance of the network under synthetic and realistic loads. Finally, we assess the processor performance implications of OPN design decisions with respect to the end-to-end latency of OPN packets and the OPN's bandwidth Paul Gratz, Karthikeyan Sankaralingam, Heather Hanson, Premkishore Shivakumar, Robert G. McDonald, Stephen W. Keckler, Doug Burger |
NOCS | 1 |
| 2006 | Implementation and Evaluation of On-Chip Network ArchitecturesabstractDriven by the need for higher bandwidth and complexity reduction, off-chip interconnect has evolved from proprietary busses to networked architectures. A similar evolution is occurring in on-chip interconnect. This paper presents the design, implementation and evaluation of one such on-chip network, the TRIPS OCN. The OCN is a wormhole routed, 4x10, 2D mesh network with four virtual channels. It provides a high bandwidth, low latency interconnect between the TRIPS processors, L2 cache banks and I/O units. We discuss the tradeoffs made in the design of the OCN, in particular why area and complexity were traded off against latency. We then evaluate the OCN using synthetic as well as realistic loads. We found that synthetic benchmarks do not provide sufficient indication of the behavior of realistic loads on this network. Finally, we examine the effect of link bandwidth and router FIFO depth on overall performance. Paul Gratz, Changkyu Kim, Robert G. McDonald, Stephen W. Keckler, Doug Burger |
ICCD | 1 |
| 2006 | Distributed Microarchitectural Protocols in the TRIPS Prototype ProcessorabstractGrowing on-chip wire delays will cause many future microarchitectures to be distributed, in which hardware resources within a single processor become nodes on one or more switched micronetworks. Since large processor cores will require multiple clock cycles to traverse, control must be distributed, not centralized. This paper describes the control protocols in the TRIPS processor, a distributed, tiled microarchitecture that supports dynamic execution. It details each of the five types of reused tiles that compose the processor, the control and data networks that connect them, and the distributed microarchitectural protocols that implement instruction fetch, execution, flush, and commit. We also describe the physical design issues that arose when implementing the microarchitecture in a 170M transistor, 130nm ASIC prototype chip composed of two 16-wide issue distributed processor cores and a distributed 1MB non-uniform (NUCA) on-chip memory system Karthikeyan Sankaralingam, Ramadass Nagarajan, Robert G. McDonald, Rajagopalan Desikan, Saurabh Drolia, Madhu Saravana Sibi Govindan, Paul Gratz, Divya Gulati, Heather Hanson, Changkyu Kim, Haiming Liu 0001, Nitya Ranganathan, Simha Sethumadhavan, Sadia Sharif, Premkishore Shivakumar, Stephen W. Keckler, Doug Burger |
MICRO | 7 |