VLDB 2026 Research / reviewers in the wild / expert
Mark Hempstead
dblp:53/1682
· DBLP profile ↗
38ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-9696-4741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CBUE: Conclusion Based Utility Evaluation for Differentially Private Categorical Data
Furkan Sarikaya, Shaohua Lu, Johes Bater, Mark Hempstead |
SP | 4 |
| 2023 | Boreas: A Cost-Effective Mitigation Method for Advanced Hotspots using Machine Learning and Hardware TelemetryabstractManaging advanced hotspots on modern microprocessors is a critical and worsening issue, affecting performance, product reliability, and device lifetime. Many thermal management techniques focus primarily on remaining below a critical temperature, and while they are successful to that end, they come with a significant performance cost. This cost stems from the large temperature guardbands that are needed to ensure device safety and correct IC operation. These guardbands must be sized to account for multiple factors including 1) control-loop latency, 2) thermal sensor delay, 3) thermal gradients inside timing paths that could result in timing violations, and 4) the instantaneous temperature delta between a temperature sensor and the true peak temperature on the IC. This work demonstrates the need for novel hotspot avoidance techniques that can react quickly and simultaneously account for each of these concerns in order to safely maximize performance. Recently introduced hotspot metrics–Hotspot-Severity and Maximum Local Temperature Difference (MLTD)–are used in order to allow model designers to have a single optimization target that accounts for each of these thermal concerns simultaneously. We present Boreas, a novel hotspot mitigation technique that uses a Machine Learning model implemented in an on-chip specialized hardware accelerator that leverages micro-architectural performance counters. Boreas outperforms existing thermal management techniques while remaining lightweight and well-suited for implementation in hardware. Even with a conservative thermal sensor delay, Boreas is able to predict severity with high precision, resulting in effective hotspot mitigation on unseen workloads. These machine learning models were, therefore, able to select a frequency that was 4.5% better than thermal only models on average, and up to 9.6% higher in the best case, while having the same reliability budget as the thermal models. Maziar Amiraski, David Werner, Alexander Hankin, Julien Sebot, Kaushik Vaidyanathan, Mark Hempstead |
ISPASS | 6 |
| 2022 | Hercules: Heterogeneity-Aware Inference Serving for At-Scale Personalized RecommendationabstractPersonalized recommendation is an important class of deep-learning applications that powers a large collection of internet services and consumes a considerable amount of datacenter resources. As the scale of production-grade recommendation systems continues to grow, optimizing their serving performance and efficiency in a heterogeneous datacenter is important and can translate into infrastructure capacity saving. In this paper, we propose Hercules, an optimized framework for personalized recommendation inference serving that targets diverse industry-representative models and cloud-scale heterogeneous systems. Hercules performs a two-stage optimization procedure — offline profiling and online serving. The first stage searches the large under-explored task scheduling space with a gradient-based search algorithm achieving up to 9.0× latency-bounded throughput improvement on individual servers; it also identifies the optimal heterogeneous server architecture for each recommendation workload. The second stage performs heterogeneity-aware cluster provisioning to optimize resource mapping and allocation in response to fluctuating diurnal loads. The proposed cluster scheduler in Hercules achieves 47.7% cluster capacity saving and reduces the provisioned power by 23.7% over a state-of-the-art greedy scheduler. Liu Ke 0001, Udit Gupta 0001, Mark Hempstead, Carole-Jean Wu, Hsien-Hsin S. Lee, Xuan Zhang 0001 |
HPCA | 3 |
| 2022 | NVMExplorer: A Framework for Cross-Stack Comparisons of Embedded Non-Volatile MemoriesabstractThe current computing landscape is dominated by data-intensive applications, making data movement one of the most prominent performance bottlenecks. With repeated off-chip memory access to DRAM driving up power, and SRAM technology scaling and leakage power limiting the efficiency of embedded memories, there is a need for new memory systems that can enable denser, more energy-efficient future on-chip storage. The actively expanding field of emerging, embeddable non-volatile memory (eNVM) technologies is providing many potential candidates to satisfy this need. However, eNVM cell technologies are in vastly different stages of development and introduce distinct trade-offs in terms of density, read, write, and reliability characteristics.We present NVMExplorer (http://nvmexplorer.seas.harvard.edu/): a cross-stack design space exploration framework to compare and evaluate future on-chip memory solutions with system constraints and application-level impacts in-the-loop. This work uses NVMExplorer to evaluate eNVM-based storage for a range of application and system contexts including machine learning on the edge, graph analytics, and general purpose cache. Additionally, NVMExplorer provides an interactive and easily navigable set of data visualizations, which allow users to quickly answer their specific questions regarding eNVMs, filter according to system and application constraints, and efficiently iterate and refine the design space. Lillian Pentecost, Alexander Hankin, Marco Donato, Mark Hempstead, Gu-Yeon Wei, David Brooks 0001 |
HPCA | 4 |
| 2022 | CASHT: Contention Analysis in Shared Hierarchies with TheftsabstractCache management policies should consider workloads’ contention behavior when managing a shared cache. Prior art makes estimates about shared cache behavior by adding extra logic or time to isolate per workload cache statistics. These approaches provide per-workload analysis but do not provide a holistic understanding of the utilization and effectiveness of caches under the ever-growing contention that comes standard with scaling cores. We present Contention Analysis in Shared Hierarchies using Thefts, or CASHT, 1 a framework for capturing cache contention information both offline and online. CASHT takes advantage of cache statistics made richer by observing a consequence of cache contention: inter-core evictions, or what we call THEFTS. We use thefts to complement more familiar cache statistics to train a learning model based on Gradient-boosting Trees (GBT) to predict the best ways to partition the last-level cache. GBT achieves 90+% accuracy with trained models as small as 100 B and at least 95% accuracy at 1 kB model size when predicting the best way to partition two workloads. CASHT employs a novel run-time framework for collecting thefts-based metrics despite partition intervention, and enables per-access sampling rather than set sampling that could add overhead but may not capture true workload behavior. Coupling CASHT and GBT for use as a dynamic policy results in a very lightweight and dynamic partitioning scheme that performs within a margin of error of Utility-based Cache Partitioning at a 1/8 the overhead. Cesar Gomes, Maziar Amiraski, Mark Hempstead |
ACM Trans. Archit. Code Optim. | 3 |
| 2021 | Understanding Capacity-Driven Scale-Out Neural Recommendation InferenceabstractDeep learning recommendation models have grown to the terabyte scale. Traditional serving schemes-that load entire models to a single server-are unable to support this scale. One approach to support these models is distributed serving, or distributed inference, which divides the memory requirements of a single large model across multiple servers. This work is a first-step for the systems community to develop novel model-serving solutions, given the huge system design space. Large-scale deep recommender systems are a novel workload and vital to study, as they consume up to 79% of all inference cycles in the data center. To that end, this work is the first to describe and characterize scale-out deep learning recommender inference using data-center serving infrastructure. This work specifically explores latency-bounded inference systems, compared to the throughput-oriented training systems of other recent works. We find that the latency and compute overheads of distributed inference are largely attributed to a model's static embedding table distribution and sparsity of inference request inputs. We evaluate three embedding table mapping strategies on three representative models and specify the challenging design trade-offs in terms of end-to-end latency, compute overhead, and resource efficiency. Overall, we observe a modest latency overhead with distributed inference-P99 latency is increased by only 1% in the best case configuration. The latency overheads are a result of the commodity infrastructure used and the sparsity of embedding tables. Encouragingly, we also show how distributed inference can account for efficiency improvements in data-center scale recommendation serving. Michael Lui, Yavuz Yetim, Özgür Özkan, Shin-Yeh Tsai, Carole-Jean Wu, Mark Hempstead |
ISPASS | 7 |
| 2021 | Thermal-Aware Overclocking for SmartphonesabstractHeat dissipation and battery life continue to be major challenges for smartphones. Smartphones seldom spend time at their highest performance points due to thermal concerns and frequently undergo thermal-throttling, where performance is limited while the device cools down. While overclocking and computational sprinting can be used to increase system performance, these techniques have not been evaluated on smartphones because they exacerbate both heat dissipation and battery life. In recent years, certain machine-learning workloads such as object detection and speech recognition have been moving away from the cloud and towards the edge. These workloads are short and user-facing making them excellent candidates for sprinting. To successfully overclock any workload however, any applied technique must ensure that it avoids forcing the system to throttle. In this paper, we describe and evaluate a system that can accurately predict the impact of workloads on the thermal state of a smartphone, enabling it to determine whether overclocking a specific workload will result in thermal-throttling. We show that careful application of overclocking using our system can decrease the latency of certain user-facing workloads by up to 18%. In this paper, we describe and evaluate a system that can accurately predict the impact of workloads on the thermal state of a smartphone, enabling it to determine whether overclocking a specific workload will result in thermal-throttling. We show that careful application of overclocking using our system can decrease the latency of certain user-facing workloads by up to 18%. Guru Prasad Srinivasa, David Werner, Mark Hempstead, Geoffrey Challen |
ISPASS | 3 |
| 2020 | Early-stage Automated Identification Tool for Shared AcceleratorsabstractThe use of application-specific accelerators to improve systems’ energy-efficiency and performance is becoming more prevalent. To overcome the tight area budget on embedded systems we propose an early detection tool that complements existing High-level Synthesis tools by identifying computationally similar synthesizable kernels that are used to build Shared Accelerators (SAs). SAs are specialized hardware accelerators that execute very different software kernels but share the common hardware functions between them. SAs can provide increased coverage if similarities between the dataflow and control flow of seemingly very different workloads are detected. Existing methods use either dynamic traces or analyze register transfer level (RTL) implementations to find these similarities which requires deep knowledge of RTL and the time-consuming RTL design process. Parnian Mokri, Mark Hempstead |
FCCM | 2 |
| 2020 | Early-stage Automated Identification of Similar Hardware Implementations with Abstract-Syntax-TreeabstractThe resource requirements of application-specific accelerators challenge embedded system designers who have a tight area budget but must cover a range of possible software kernels. We propose an early detection methodology (ReconfAST) to identify computationally similar synthesizable kernels to build Shared Accelerators (SAs). SAs are specialized hardware accelerators that execute very different software kernels but share the common hardware functions between them. SAs increase the fraction of workloads covered by specialized hardware by detecting similarities in dataflow and control flow between seemingly very different workloads. Existing methods use either dynamic traces or analyze register transfer level (RTL) implementations to find these similarities which require deep knowledge of RTL and time-consuming design process. Parnian Mokri, Maziar Amiraskari, Yuelin Liu, Mark Hempstead |
FPGA | 4 |
| 2020 | The Architectural Implications of Facebook's DNN-Based Personalized RecommendationabstractThe widespread application of deep learning has changed the landscape of computation in data centers. In particular, personalized recommendation for content ranking is now largely accomplished using deep neural networks. However, despite their importance and the amount of compute cycles they consume, relatively little research attention has been devoted to recommendation systems. To facilitate research and advance the understanding of these workloads, this paper presents a set of real-world, production-scale DNNs for personalized recommendation coupled with relevant performance metrics for evaluation. In addition to releasing a set of open-source workloads, we conduct in-depth analysis that underpins future system design and optimization for at-scale recommendation: Inference latency varies by 60% across three Intel server generations, batching and co-location of inference jobs can drastically improve latency-bounded throughput, and diversity across recommendation models leads to different optimization strategies. Udit Gupta 0001, Carole-Jean Wu, Xiaodong Wang 0020, Maxim Naumov, Brandon Reagen, David Brooks 0001, Bradford Cottel, Kim M. Hazelwood, Mark Hempstead, Bill Jia, Hsien-Hsin S. Lee, Andrey Malevich, Dheevatsa Mudigere, Mikhail Smelyanskiy, Liang Xiong, Xuan Zhang 0001 |
HPCA | 9 |
| 2020 | SnackNoC: Processing in the Communication LayerabstractIn this work, we propose and evaluate a Network-on-Chip (NoC) augmented with light-weight processing elements to provide a lean dataflow-style system. We show that contemporary NoC routers can frequently experience long periods of idle time, with less than 10% link utilization in HPC applications. By repurposing the temporal and spatial slack of the NoC, the proposed platform, SnackNoC, is able to compute linear algebra kernels efficiently within the communication layer with minimal additional resource costs. SnackNoC 'Snack' application kernels are programmed with a producer-consumer data model that uses the NoC slack to store and transmit intermediate data between processing elements. SnackNoC is demonstrated in a multi-program environment that continually executes linear algebra kernels on the NoC simultaneously with chip multiprocessor (CMP) applications on the processor cores. Linear algebra kernels are computed up to 14.2x faster on SnackNoC compared to an Intel Haswell EPx86 processing core. The cost of executing 'snack' kernels in parallel to the CMP applications is a minimal runtime impact of 0.01% to 0.83% due to higher link utilization, and an uncore area overhead of 1.1%. Karthik Sangaiah, Michael Lui, Ragh Kuttappa, Baris Taskin, Mark Hempstead |
HPCA | 5 |
| 2020 | Early-stage Automated Accelerator Identification Tool for Embedded Systems with Limited AreaabstractDesigners are turning toward hardware specialization through the use of application-specific accelerators to provide energy-efficiency and performance. We propose an early detection methodology to identify computationally similar and synthesize-able kernels that are used to build Shared Accelerators (SAs). SAs are specialized hardware accelerators that execute very different software kernels but share the common hardware functions between them. SAs can provide increased coverage if both data flow and control flow similarities between - seemingly very different- workloads are detected. Parnian Mokri, Mark Hempstead |
ICCAD | 2 |
| 2020 | RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingabstractPersonalized recommendation systems leverage deep learning models and account for the majority of data center AI cycles. Their performance is dominated by memory-bound sparse embedding operations with unique irregular memory access patterns that pose a fundamental challenge to accelerate. This paper proposes a lightweight, commodity DRAM compliant, near-memory processing solution to accelerate personalized recommendation inference. The in-depth characterization of production-grade recommendation models shows that embedding operations with high model-, operator and data-level parallelism lead to memory bandwidth saturation, limiting recommendation inference performance. We propose RecNMP which provides a scalable solution to improve system throughput, supporting a broad range of sparse embedding models. RecNMP is specifically tailored to production environments with heavy co-location of operators on a single server. Several hardware/software cooptimization techniques such as memory-side caching, tableaware packet scheduling, and hot entry profiling are studied, providing up to 9.8× memory latency speedup over a highly-optimized baseline. Overall, RecNMP offers 4.2× throughput improvement and 45.8% memory energy savings. Liu Ke 0001, Udit Gupta 0001, Benjamin Y. Cho, David Brooks 0001, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim M. Hazelwood, Bill Jia, Hsien-Hsin S. Lee, Meng Li 0004, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang 0020, Brandon Reagen, Carole-Jean Wu, Mark Hempstead, Xuan Zhang 0001 |
ISCA | 20 |
| 2020 | C^2AFE: Capacity Curve Annotation and Feature Extraction for Shared Cache AnalysisabstractRecent years see cache partitioning techniques employ commercial way partitioning to great effect. Techniques like partitioning build analysis on miss curves. However, the various forms of partitioning run into the issue of core scaling, requiring multi-tenant partitions and analysis of how workloads behave when sharing ever shrinking resources. We observe surprising variations in performance which can be disjoint with variations in cache misses that encourage such analysis. We present Capacity Curve Annotation and Feature Extraction, methodologies for annotation, feature extraction, and analysis of sensitivity studies based on performance curves. Cesar Gomes, Mark Hempstead |
ISPASS | 2 |
| 2019 | Quantifying Process Variations and Its Impacts on SmartphonesabstractProcess variation can cause the performance and energy consumption of smartphones of the same model to vary significantly. While process variation has been studied in detail, the effects on smartphone performance have not been quantified and evaluated. In this work we study the performance and energy differences of 5 recent SoC generations caused by underlying process variation. We make two important contributions. First, we present a methodology to construct a temperature-stabilized environment to perform repeatable power and performance measurements. Studying power-performance characteristics of smartphones is difficult. Running a benchmark back-to-back often produces significantly different results due to heat. Temperature, both device and ambient, play a significant role in determining performance and energy. Our methodology allows us to control for various factors and isolate the effects of the underlying process variation. We then apply our methodology to investigate performance and energy characteristics of several recent generations of smart-phone CPUs that result from process variation. Our results show that devices of the same model may exhibit differences of 10% and 12% difference in performance and energy over a fixed-duration workload. Guru Prasad Srinivasa, Scott Haseley, Geoffrey Challen, Mark Hempstead |
ISPASS | 4 |
| 2018 | Machine Learning on the Thermal Side-Channel: Analysis of Accelerator-Rich ArchitecturesabstractThe thermal profiles of integrated circuits (ICs) have been leveraged as a side-channel in multiple circuit and architectural scenarios. Applications range from identifying hardware Trojans to estimating the per-core power consumption of homogeneous multicore processors. Such scenarios leverage the correlation between the on-chip location of the consumed power with some target information of interest, such as correlating the extra power consumption at a specific circuit position with the presence of a hardware Trojan. While the spatial correlation between the power consumption and thermal profiles applies to all ICs, there is a fundamental difference in the context of modern SoCs. The difference stems from the presence of hardware accelerators, in which localized power consumption corresponds to the system performing the specific task that a given accelerator executes. The work described in the paper demonstrates the implications of correlating the thermal and power profiles of SoCs by presenting two working case studies that determine, at runtime, 1) the activity factor of each accelerator and 2) whether or not a system is infected by malware. This work relies on pre-processing thermal images in order to obtain a spatial profile of the estimated power density and uses a modified version of a previously developed technique that is tailored for use with accelerator-rich ICs. The resulting power estimates are fed into machine learning models that predict the core activity factor with mean average errors between 3% and 5% for the highest performing core. The statistical models used for malware detection result in an AuROC score of up to 1.0 and 0.9 when the malware offsets the activity factor of a single core by 2.5% and the 3-sigma width of the workload activity factor distribution is 2.5% and 5%, respectively. David Werner, Kyle Juretus, Ioannis Savidis, Mark Hempstead |
ICCD | 4 |
| 2018 | Towards Cross-Framework Workload Analysis via Flexible Event-Driven InterfacesabstractHardware/software co-design and software profiling rest on the ability to perform a range of workload analyses. State-of-the-art tools and methods used in such analyses utilize either custom solutions or complex frameworks. There are two problems with this approach: 1) duplicated development work when moving to new and unsupported frameworks or platforms, and 2) the additional burden of in-depth knowledge required to develop the analysis tools. This work presents a methodology to solve these inefficiencies by decoupling workload analysis from the underlying techniques used to observe the workload. The interface is designed to be cross-platform and presents workloads as a set of configurable events with scalable levels-of-detail. An implementation of the methodology, PRISM, is presented which leverages two popular dynamic binary instrumentation tools, Valgrind and DynamoRIO, and additionally Intel PT via Linux perf. The goals of the methodology are three-fold: modularity, flexibility, and productivity. Three analyses are conducted using PRISM to demonstrate these properties: 1) discrepancies are assessed between workloads generated with Valgrind, DynamoRIO, and perf, 2) scalability of a complex Valgrind trace generation tool is improved, and 3) prototyping of a new dynamic loop detection and data-dependence tool is demonstrated. The average overhead of PRISM compared to in-framework analysis is 33% in the worse case and under 1% during typical analysis. Michael Lui, Karthik Sangaiah, Mark Hempstead, Baris Taskin |
ISPASS | 3 |
| 2018 | SynchroTrace: Synchronization-Aware Architecture-Agnostic Traces for Lightweight Multicore Simulation of CMP and HPC WorkloadsabstractTrace-driven simulation of chip multiprocessor (CMP) systems offers many advantages over execution-driven simulation, such as reducing simulation time and complexity, allowing portability, and scalability. However, trace-based simulation approaches have difficulty capturing and accurately replaying multithreaded traces due to the inherent nondeterminism in the execution of multithreaded programs. In this work, we present SynchroTrace, a scalable, flexible, and accurate trace-based multithreaded simulation methodology. By recording synchronization events relevant to modern threading libraries (e.g., Pthreads and OpenMP) and dependencies in the traces, independent of the host architecture, the methodology is able to accurately model the nondeterminism of multithreaded programs for different hardware platforms and threading paradigms. Through capturing high-level instruction categories, the SynchroTrace average CPI trace Replay timing model offers fast and accurate simulation of many-core in-order CMPs. We perform two case studies to validate the SynchroTrace simulation flow against the gem5 full-system simulator: (1) a constraint-based design space exploration with traditional CMP benchmarks and (2) a thread-scalability study with HPC-representative applications. The results from these case studies show that (1) our trace-based approach with trace filtering has a peak speedup of up to 18.7× over simulation in gem5 full-system with an average of 9.6× speedup, (2) SynchroTrace maintains the thread-scaling accuracy of gem5 and can efficiently scale up to 64 threads, and (3) SynchroTrace can trace in one platform and model any platform in early stages of design. Karthik Sangaiah, Michael Lui, Radhika Jagtap, Stephan Diestelhorst, Siddharth Nilakantan, Ankit More, Baris Taskin, Mark Hempstead |
ACM Trans. Archit. Code Optim. | 8 |
| 2016 | Algorithms for CPU and DRAM DVFS under inefficiency constraintsabstractDynamic voltage and frequency scaling (DVFS) of both the core and DRAM provides opportunities to trade-off performance in order to save energy. Previous approaches to core and DRAM power management using DVFS used performance, specifically acceptable performance loss, as a constraint. We present energy management algorithms that coordinate core and DRAM frequency scaling under a specified energy budget. Approaches that work under performance constraints, as we will show, are not directly applicable to systems operating under energy constraints, as it is difficult to calculate the correct performance bounds in real-time to stay under an energy budget. Setting arbitrary energy budgets for a diverse set of applications can be harmful to application performance. We use the previously introduced concept of Inefficiency—the additional amount of energy above the minimum required energy that can be used to improve performance—to provide a dynamic energy constraint to our system. We introduce new power management algorithms that search the power and performance space to find the best performing point under this constraint. We demonstrate the efficacy of our algorithms using CPU DVFS and DRAM frequency scaling. We show that our algorithms have 24% lower tuning cost and save up to 5% energy with a little performance loss compared to a state-of-the-art performance constrained system. Rizwana Begum, Mark Hempstead, Guru Prasad Srinivasa, Geoffrey Challen |
ICCD | 2 |
| 2015 | Uncore RPD: Rapid Design Space Exploration of the Uncore via Regression ModelingabstractA regression-based design space exploration methodology is proposed that models the impacts of the memory hierarchy and the network-on-chip (NoC) on the overall chip multiprocessor (CMP) performance. Designers cannot explore all possible designs for a NoC without considering interactions with the rest of the uncore, in particular the cache configuration and memory hierarchy which determine the amount and pattern of the traffic on the NoC. The proposed regression model is able to capture the salient design points of the uncore for a comprehensive design space exploration by designing memory and NoC-specific regression models and leveraging recent advances in uncore simulation. To show the utility of our methodology, Uncore RPD, two case studies are presented: i) analyzing and refining regression models for an 8-core CMP and ii) performing a rapid design space exploration to find best performing designs of a NoC-based CMP given area-constraints for CMPs of up to 64 cores. Through these case studies, it is shown that i) simultaneous consideration of the memory and NoC parameters in the NoC design space exploration can refine uncore-based regression models, ii) sampling techniques must consider the dynamic design space of the uncore, and iii) overall, the proposed regression models reduce the amount of simulations required to characterize the NoC design space by up to four orders of magnitude. Karthik Sangaiah, Mark Hempstead, Baris Taskin |
ICCAD | 2 |
| 2015 | Power-agility metrics: Measuring dynamic characteristics of energy proportionalityabstractThere has been a recent call for energy proportional systems that consume power proportional to its usage. However, these systems are only theoretical. In order to achieve system-level energy proportionality, effective energy management algorithms are needed that can tune each individual component of the system dynamically in response to usage. We define the ability of a system to detect, select and transition to efficient power settings dynamically during the execution of applications as Power-Agility. However, there are no metrics available that can measure how well the device is achieving Power-Agility, guide designers on how to make better system components and management algorithms. We present two metrics, Selection Power Agility and Transition Power Agility that measure the ability of the systems to select the efficient power settings and transition to the selected settings respectively. We study the Power-Agility of various representative algorithms and modeled devices for SPEC2006 benchmarks. Our results show that the Power-Agility of systems across algorithms varies with applications. We also show that both the discrete power settings and the transition latencies of the devices have an impact on system's Power-Agility. Rizwana Begum, Mark Hempstead |
ICCD | 2 |
| 2015 | Combative cache efficacy techniques: Cache replacement in the context of independent prefetching in last level cacheabstractThe "memory wall" in CPU design refers to increasing divergence in performance growth between processors and memory. Prefetching and cache replacement were developed to overcome this, but there are few evaluations of the two techniques being used together. We contribute to this space by evaluating modern replacement policies in the context of a stride-N prefetcher. In exposing combative behavior between modern policies and the prefetcher, we hope to inform future design choices regarding such integration. Further, we present a means to either recoup or surpass performance gains seen in isolation for policies that exhibit conflict. Cesar Gomes, Mark Hempstead |
ICCD | 2 |
| 2015 | Synchrotrace: synchronization-aware architecture-agnostic traces for light-weight multicore simulationabstractTrace-driven simulation of chip multiprocessor (CMP) systems offers many advantages over execution-driven simulation, such as reducing simulation time and complexity, and allowing portability, and scalability. However, trace-based simulation approaches have encountered difficulty capturing and accurately replaying multi-threaded traces due to the inherent non-determinism in the execution of multi-threaded programs. In this work, we present SynchroTrace, a scalable, flexible, and accurate trace-based multi-threaded simulation methodology. The methodology captures synchronization- and dependency-aware, architectureagnostic, multi-threaded traces and uses a replay mechanism that plays back these traces correctly. By recording synchronization events and dependencies in the traces, independent of the host architecture, the methodology is able to accurately model the non-determinism of multi-threaded programs for different platforms. We validate the SynchroTrace simulation flow by successfully achieving the equivalent results of a constraint-based design space exploration with the Gem5 Full-System simulator. The results from simulating benchmarks from PARSEC 2.1 and Splash-2 show that our trace-based approach with trace filtering has a peak speedup of up to 18.4ξ over simulation in Gem5 Full-System with an average of about 7.5ξ speedup. We are also able to compress traces up to 74% of their original size with almost no impact on accuracy. Siddharth Nilakantan, Karthik Sangaiah, Ankit More, Giordano Salvador, Baris Taskin, Mark Hempstead |
ISPASS | 6 |
| 2014 | Static thread mapping for NoCs via binary instrumentation tracesabstractA novel methodology is proposed for thread mapping on a chip-multiprocessor (CMP) system with a network-on-chip (NoC). This novel mapping leverages multi-threaded traces produced by a binary instrumentation tool, which classifies the communication and computation events for each thread of a multi-threaded program application. Processing these binary instrumentation traces after profiling, a static thread mapping is computed to improve the NoC performance. Giordano Salvador, Siddharth Nilakantan, Baris Taskin, Mark Hempstead, Ankit More |
ICCD | 4 |
| 2013 | Characterizing the costs and benefits of hardware parallelism in accelerator coresabstractPower and utilization constraints are limiting the performance gains of traditional architectures. Designers are increasingly embracing specialization to improve performance in the era of dark-silicon. General purpose processors are beginning to resemble SOC's from the embedded domain, and now include many specialized accelerator cores to improve computation-throughput while reducing the energy-cost of computation. The design-space of accelerator cores is wide and varied. Designers are able to specify how much parallelism to expose in hardware by varying input width, pipeline depth, number of compute-lanes, etc. In this paper we study three accelerator cores: DES, FFT, and Jacobi Transform, exhibiting three different types of computation: streaming cryptographic, butterfly DSP, and stencil. We investigate methods to increase parallelism within the accelerator while remaining on the pareto-frontier, and examine the trade-offs faced by designers with respect to area, power, and throughput. We present models of these trade-offs and provide insight into the design of cores under real-world constraints. Steven J. Battle, Mark Hempstead |
ICCD | 2 |
| 2013 | Register allocation and VDD-gating algorithms for out-of-order architecturesabstractRegister Files (RF) in modern out-of-order microprocessors can account for up to 30% of total power consumed by the core. The complexity and size of the RF has increased due to the transition from ROB-based to MIPSR10K-style physical register renaming. Because physical registers are dynamically allocated, the RF is not fully occupied during every phase of the application. In this paper, we propose a comprehensive power management strategy of the RF through algorithms for register allocation and register-bank power-gating that are informed by both microarchitecture details and circuit costs. We investigate algorithms to control where to place registers in the RF, when to disable banks in the RF, and when to re-enable these banks. We include detailed circuit models to estimate the cost for banking and power-gating the RF. We are able to save up to 50% of the leakage energy vs. a baseline monolithic RF, and save 11% more leakage energy than fine-grained VDD-gating schemes. Steven J. Battle, Mark Hempstead |
ICCD | 2 |
| 2012 | Flexible register management using reference countingabstractConventional out-of-order processors that use a unified physical register file allocate and reclaim registers explicitly using a free list that operates as a circular queue. We describe and evaluate a more flexible register management scheme - reference counting. We implement reference counting using a bit-matrix with a column for every physical register and a row for every entity that can hold a physical register, e.g., an in-flight instruction. Columns are NOR'ed together to create a bitvector free list from which registers are allocated using priority encoders. We describe reference counting designs that support micro-architectural techniques including register file power gating, dynamic register move elimination, register file checkpointing, and latency tolerant execution. Performance and circuit simulation show that the energy cost of reference counting is low and is easily recouped by the savings of the techniques it enables. Steven J. Battle, Andrew D. Hilton, Mark Hempstead, Amir Roth |
HPCA | 3 |
| 2012 | The accelerator store: A shared memory framework for accelerator-based systemsabstractThis paper presents the many-accelerator architecture, a design approach combining the scalability of homogeneous multi-core architectures and system-on-chip's high performance and power-efficient hardware accelerators. In preparation for systems containing tens or hundreds of accelerators, we characterize a diverse pool of accelerators and find each contains significant amounts of SRAM memory (up to 90% of their area). We take advantage of this discovery and introduce the accelerator store, a scalable architectural component to minimize accelerator area by sharing its memories between accelerators. We evaluate the accelerator store for two applications and find significant system area reductions (30%) in exchange for small overheads (2% performance, 0%--8% energy). The paper also identifies new research directions enabled by the accelerator store and the many-accelerator architecture. Michael J. Lyons 0003, Mark Hempstead, Gu-Yeon Wei, David Brooks 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | Evaluation of an accelerator architecture for speckle reducing anisotropic diffusionabstractIncreasing chip power density has brought application specific accelerator architectures to the forefront as an energy and area efficient solution. While GPGPU systems take advantage of specialized hardware to perform computationally intensive tasks faster than chip multiprocessor (CMP) systems, accelerators are hardware units that are designed to execute a specific application efficiently. Real-time ultrasound imaging applications require the removal of multiplicative noise while maintaining a steady frame-rate, and are good candidates to explore accelerator-based systems. In this paper, we propose and evaluate the architecture of an accelerator designed to improve performance of SRAD image enhancing algorithm. We compare the projected performance of the SRAD accelerator to software implementations on a multi-core CPU and a CPU+GPU system. The proposed architecture achieves higher throughput by eliminating redundant fetches from memory and by storing intermediate data locally. The speedup of the GPU is found to be 3.2x over the CPU, while the accelerator achieved a speedup of 24x. The area efficiency of the GPU and accelerator is up to 1.6x and 370x better than the CPU, respectively. In comparison with the CPU, we find that the energy consumed for operation on a single frame is found to be 1.5x lesser on the GPU and upto 580x lesser on the accelerator. Siddharth Nilakantan, Srikanth Annangi, Nikhil Gulati, Karthik Sangaiah, Mark Hempstead |
CASES | 5 |
| 2011 | The Case for Power-Agile Computing
Geoffrey Challen, Mark Hempstead |
HotOS | 2 |
| 2009 | An accelerator-based wireless sensor network processor in 130nm CMOSabstractNetworks of ultra-low-power nodes capable of sensing, computation, and wireless communication have applications in medicine, science, industrial automation, and security. Over the past few years, deployments of wireless sensor networks (WSNs) have utilized nodes based on off-the-shelf general purpose microcontrollers. Reducing power consumption requires the development of System-on-chip (SoC) implementations that must provide both energy efficiency and adequate performance to meet the demands of the long deployment lifetimes and bursts of computation that characterize WSN applications. This work takes a holistic approach and, thus, studies all layers of the design space, from the applications and architecture, to process technology and circuits. This paper introduces the emerging application space of wireless sensor networks and describes the motivation and need for a custom system architecture. The proposed design fully embraces the accelerator-based computing paradigm, including acceleration for the network layer (routing) and application layer (data filtering). Moreover, the architecture can disable the accelerators via VDD-gating to minimize leakage current during the long idle times common to WSN applications. We have implemented a system architecture for wireless sensor network nodes in 130nm CMOS. It operates at 550 mV and 12.5 MHz. Our system uses 100x less power when idle than a traditional microcontroller, and 10-600x less energy when active. Mark Hempstead, Gu-Yeon Wei, David Brooks 0001 |
CASES | 1 |
| 2008 | System design considerations for sensor network applicationsabstractSystems research in the emerging space of wireless sensor networks has exploded. Researchers have deployed nodes composed of a wireless radio, MEMS sensors and low power computation for applications from medical sensing to volcanic monitoring. We must consider several requirements - including the need for inexpensive, long-lasting, highly reliable devices coupled with very low performance requirements - when designing devices for wireless sensor networks. An untethered, fully-integrated node that operates off of energy scavenged from the ambient environment is the ultimate goal. We take an application-driven approach to the design of a wireless sensor network node. Our approach addresses the event-driven nature that is characteristic of many sensor network workloads. We have completed a detailed architectural analysis of this space using a full-system simulator and RTL model. From this analysis, we chose to implement a design that best achieves the power goals and performance requirements of wireless sensor network applications. Mark Hempstead, Gu-Yeon Wei, David Brooks 0001 |
ISCAS | 1 |
| 2006 | Architecture and circuit techniques for low-throughput, energy-constrained systems across technology generationsabstractRising interest in the applications of wireless sensor networks has spurred research in the development of computing systems for low-throughput, energy-constrained applications. Unlike traditional performance oriented applications, sensor network nodes are primarily constrained by operation lifetime, which is limited by power consumption. Advanced CMOS process technologies provide ever increasing transistor density and improved performance characteristics. However, shrinking feature size and decreasing threshold voltages also lead to significant increases in leakage current, which is especially troublesome for applications with significant idle times. This work investigates tradeoffs between leakage and active power for low-throughput applications. We study these issues across a range of process technologies on a computing architecture that provides explicit support for fine-grain leakage-control techniques such as Vdd-gating and adaptive body bias. We present a methodology for selecting design parameters, including choice of process technology, that makes the optimal tradeoff between active power and leakage power for a given workload. Our results show that leakage power will dominate the selection of process technology, and architectures that support advanced leakage control techniques at the circuit level will be essential. We argue that without advanced low-power architectures future nano-scale process technologies will not be suited for sensor network applications. Mark Hempstead, Gu-Yeon Wei, David Brooks 0001 |
CASES | 1 |
| 2006 | A Realistic Power Consumption Model for Wireless Sensor Network DevicesabstractA realistic power consumption model of wireless communication subsystems typically used in many sensor network node devices is presented. Simple power consumption models for major components are individually identified, and the effective transmission range of a sensor node is modeled by the output power of the transmitting power amplifier, sensitivity of the receiving low noise amplifier, and RF environment. Using this basic model, conditions for minimum sensor network power consumption are derived for communication of sensor data from a source device to a destination node. Power consumption model parameters are extracted for two types of wireless sensor nodes that are widely used and commercially available. For typical hardware configurations and RF environments, it is shown that whenever single hop routing is possible it is almost always more power efficient than multi-hop routing. Further consideration of communication protocol overhead also shows that single hop routing will be more power efficient compared to multi-hop routing under realistic circumstances. This power consumption model can be used to guide design choices at many different layers of the design space including, topology design, node placement, energy efficient routing schemes, power management and the hardware design of future wireless sensor network devices Mark Hempstead, Woodward Yang |
SECON | 2 |
| 2005 | An Ultra Low Power System Architecture for Sensor Network ApplicationsabstractRecent years have seen a burgeoning interest in embedded wireless sensor networks with applications ranging from habitat monitoring to medical applications. Wireless sensor networks have several important attributes that require special attention to device design. These include the need for inexpensive, long-lasting, highly reliable devices coupled with very low performance requirements. Ultimately, the "holy grail" of this design space is a truly untethered device that operates off of energy scavenged from the ambient environment. In this paper, we describe an application-driven approach to the architectural design and implementation of a wireless sensor device that recognizes the event-driven nature of many sensor-network workloads. We have developed a full-system simulator for our sensor node design to verify and explore our architecture. Our simulation results suggest one to two orders of magnitude reduction in power dissipation over existing commodity-based systems for an important class of sensor network applications. We are currently in the implementation stage of design, and plan to tape out the first version of our system within the next year. Mark Hempstead, Nikhil Tripathi, Patrick Mauro, Gu-Yeon Wei, David Brooks 0001 |
ISCA | 1 |
| 2005 | Power and thermal effects of SRAM vs. Latch-Mux design styles and clock gating choicesabstractThis paper studies the impact on energy efficiency and thermal behavior of design style and clock-gating style in queue and array structures. These structures are major sources of power dissipation, and both design styles and various clock gating schemes can be found in modern, high-performance processors. Although some work in the circuits domain has explored these issues from a power perspective, thermal treatments are less common, and we are not aware of any work in the architecture domain.We study both SRAM and latch and multiplexer ("latch-mux") designs and their associated clock-gating options. Using circuit-level simulations of both design styles, we derive power-dissipation ratios which are then used in cycle-level power/performance/thermal simulations. We find that even though the "unconstrained" power of SRAM designs is always better than latch-mux designs, latch-mux designs dissipate less power in practice when a structure's average occupancy is low but access rate is high, especially when "stall gating" is used to minimize switching power. We also find that latch-mux designs with stall gating are especially promising from a thermal perspective, because they exhibit lower power density than SRAM designs. Overall, when combined with implementation and verification challenges for SRAMs, latch-mux designs with stall gating appear especially promising for designs with thermal constraints. This paper also shows the importance of considering the interaction between architectural and circuit-design choices when performing early-stage design exploration Yingmin Li, Mark Hempstead, Patrick Mauro, David Brooks 0001, Kevin Skadron |
ISLPED | 2 |
| 2004 | TinyBench: The Case For A Standardized Benchmark Suite for TinyOS Based Wireless Sensor Network DevicesabstractThe growing wireless sensor network research community lacks a standard method for evaluating hardware platforms. Traditional benchmark suites do not sufficiently address the needs of sensor network designers. This work provides motivation for a benchmark suite and details an approach for benchmarking TinyOS compatible hardware. To aid the development of future hardware architectures, we propose the creation of a standard single node benchmark suite, based on both real applications and "stressmarks." We present sample benchmark results and call for further work in this area. Mark Hempstead, Matt Welsh, David Brooks 0001 |
LCN | 1 |
| 2004 | Simulating the power consumption of large-scale sensor network applicationsabstractDeveloping sensor network applications demands a new set of tools to aid programmers. A number of simulation environments have been developed that provide varying degrees of scalability, realism, and detail for understanding the behavior of sensor networks. To date, however, none of these tools have addressed one of the most important aspects of sensor application design: that of power consumption. While simple approximations of overall power usage can be derived from estimates of node duty cycle and communication rates, these techniques often fail to capture the detailed, low-level energy requirements of the CPU, radio, sensors, and other peripherals. Victor Shnayder, Mark Hempstead, Bor-rong Chen, Geoffrey Challen, Matt Welsh |
SenSys | 2 |