EDBT 2026 Demo / reviewers in the wild / expert
David Meisner
dblp:18/389
· DBLP profile ↗
19ranked-venue papers
9as first author
1since 2021 · last 2024
0009-0000-6248-9671ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 7 first-authorSoftware engineering, systems software and programming languages · 11 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Energy-efficient computing · 36% Performance modeling and evaluation · 26% Cloud and datacenter computing · 25% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.8 | 4 | 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 CoScale: Coordinating CPU and Memory System DVFS in Server Systems · MICRO 2012 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.8 | 1 | 2024 | ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024 |
Performance modeling and evaluation › benchmarking
performance regression testing |
0.8 | 1 | 2024 | ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024 |
Performance modeling and evaluation
workload characterization |
0.3 | 3 | 2017 | The Mystery Machine: End-to-end Performance Analysis of Large-scale Internet Services · OSDI 2014 Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 |
Cloud and datacenter computing › datacenter architecture
warehouse-scale computer |
0.3 | 1 | 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using Adrenaline · ACM Trans. Comput. Syst. 2017 |
Energy-efficient computing › power management
memory power management |
0.3 | 2 | 2012 | CoScale: Coordinating CPU and Memory System DVFS in Server Systems · MICRO 2012 MemScale: active low-power modes for main memory · ASPLOS 2011 |
Energy-efficient computing
power management |
0.3 | 2 | 2012 | CoScale: Coordinating CPU and Memory System DVFS in Server Systems · MICRO 2012 Power management of online data-intensive services · ISCA 2011 |
Energy-efficient computing
datacenter power management |
0.3 | 3 | 2011 | The PowerNap Server Architecture · ACM Trans. Comput. Syst. 2011 Power routing: dynamic power provisioning in the data center · ASPLOS 2010 PowerNap: eliminating server idle power · ASPLOS 2009 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2024 | ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024 |
Cloud and datacenter computing
warehouse-scale computing |
0.2 | 1 | 2015 | Adrenaline: Pinpointing and reining in tail queries with quick voltage boosting · HPCA 2015 |
Performance modeling and evaluation › performance evaluation methodology
end-to-end performance analysis |
0.2 | 1 | 2014 | The Mystery Machine: End-to-end Performance Analysis of Large-scale Internet Services · OSDI 2014 |
Storage systems
key-value storage |
0.2 | 1 | 2013 | Thin servers with smart pipes: designing SoC accelerators for memcached · ISCA 2013 |
Storage systems › key-value storage
memcached |
0.2 | 1 | 2013 | Thin servers with smart pipes: designing SoC accelerators for memcached · ISCA 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2012 | DreamWeaver: architectural support for deep sleep · ASPLOS 2012 |
Energy-efficient computing › power management
low-power modes |
0.1 | 1 | 2012 | DreamWeaver: architectural support for deep sleep · ASPLOS 2012 |
Memory systems
DRAM |
0.1 | 1 | 2011 | MemScale: active low-power modes for main memory · ASPLOS 2011 |
Energy-efficient computing
energy proportionality |
0.1 | 1 | 2011 | Power management of online data-intensive services · ISCA 2011 |
Energy-efficient computing › power management › dynamic power management
idle power reduction |
0.1 | 1 | 2011 | The PowerNap Server Architecture · ACM Trans. Comput. Syst. 2011 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
online data-intensive services |
0.1 | 1 | 2011 | Power management of online data-intensive services · ISCA 2011 |
Energy-efficient computing › datacenter power management
power provisioning |
0.1 | 1 | 2011 | The PowerNap Server Architecture · ACM Trans. Comput. Syst. 2011 |
Energy-efficient computing › datacenter power management
server power management |
0.1 | 1 | 2011 | The PowerNap Server Architecture · ACM Trans. Comput. Syst. 2011 |
Electronic design automation › physical design › routing › VLSI routing
power/ground net routing |
0.1 | 1 | 2010 | Power routing: dynamic power provisioning in the data center · ASPLOS 2010 |
Cloud and datacenter computing
datacenter services |
0.1 | 1 | 2014 | The Mystery Machine: End-to-end Performance Analysis of Large-scale Internet Services · OSDI 2014 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
internet-scale services |
0.1 | 1 | 2014 | The Mystery Machine: End-to-end Performance Analysis of Large-scale Internet Services · OSDI 2014 |
Cloud and datacenter computing › datacenter architecture
datacenter server architecture |
0.0 | 1 | 2013 | Thin servers with smart pipes: designing SoC accelerators for memcached · ISCA 2013 |
Energy-efficient computing
datacenter energy efficiency |
0.0 | 1 | 2012 | DreamWeaver: architectural support for deep sleep · ASPLOS 2012 |
Energy-efficient computing › power management › device power management
processor power management |
0.0 | 1 | 2012 | CoScale: Coordinating CPU and Memory System DVFS in Server Systems · MICRO 2012 |
Energy-efficient computing
datacenter energy consumption |
0.0 | 1 | 2011 | The PowerNap Server Architecture · ACM Trans. Comput. Syst. 2011 |
Cloud and datacenter computing
datacenter workloads |
0.0 | 1 | 2011 | Power management of online data-intensive services · ISCA 2011 |
Cloud and datacenter computing
latency-critical applications |
0.0 | 1 | 2011 | Power management of online data-intensive services · ISCA 2011 |
Methods — techniques the papers use, named apart from their topics
DVFS · 0.5workload characterization · 0.3proactive boosting · 0.3per-query characteristics · 0.3statistical inference · 0.2quantile regression · 0.2voltage boosting · 0.2power analysis · 0.2performance counters · 0.1execution profiling · 0.1queueing theory · 0.1DFS · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
Mike Chow, Yang Wang 0009, Ayichew Hailu, Rohan Bopardikar, Jialiang Qu, David Meisner, Santosh Sonawane, Rodrigo Paim, Mack Ward, Ivor Huang, Matt McNally, Daniel Hodges, Zoltan Farkas, Caner Gocmen, Elvis Huang, Chunqiang Tang |
OSDI | 8 |
| 2017 | Reining in Long Tails in Warehouse-Scale Computers with Quick Voltage Boosting Using AdrenalineabstractReducing the long tail of the query latency distribution in modern warehouse scale computers is critical for improving performance and quality of service (QoS) of workloads such as Web Search and Memcached. Traditional turbo boost increases a processor’s voltage and frequency during a coarse-grained sliding window, boosting all queries that are processed during that window. However, the inability of such a technique to pinpoint tail queries for boosting limits its tail reduction benefit. In this work, we propose Adrenaline , an approach to leverage finer-granularity (tens of nanoseconds) voltage boosting to effectively rein in the tail latency with query-level precision. Two key insights underlie this work. First, emerging finer granularity voltage/frequency boosting is an enabling mechanism for intelligent allocation of the power budget to precisely boost only the queries that contribute to the tail latency; second, per-query characteristics can be used to design indicators for proactively pinpointing these queries, triggering boosting accordingly. Based on these insights, Adrenaline effectively pinpoints and boosts queries that are likely to increase the tail distribution and can reap more benefit from the voltage/frequency boost. By evaluating under various workload configurations, we demonstrate the effectiveness of our methodology. We achieve up to a 2.50 × tail latency improvement for Memcached and up to a 3.03 × for Web Search over coarse-grained dynamic voltage and frequency scaling (DVFS) given a fixed boosting power budget. When optimizing for energy reduction, Adrenaline achieves up to a 1.81 × improvement for Memcached and up to a 1.99 × for Web Search over coarse-grained DVFS. By using the carefully chosen boost thresholds, Adrenaline further improves the tail latency reduction to 4.82 × over coarse-grained DVFS. Chang-Hong Hsu, Michael Laurenzano, David Meisner, Thomas F. Wenisch, Ronald G. Dreslinski, Jason Mars, Lingjia Tang |
ACM Trans. Comput. Syst. | 4 |
| 2016 | Treadmill: Attributing the Source of Tail Latency through Precise Load Testing and Statistical InferenceabstractManaging tail latency of requests has become one of the primary challenges for large-scale Internet services. Data centers are quickly evolving and service operators frequently desire to make changes to the deployed software and production hardware configurations. Such changes demand a confident understanding of the impact on one's service, in particular its effect on tail latency (e.g., 95th-or 99th-percentile response latency of the service). Evaluating the impact on the tail is challenging because of its inherent variability. Existing tools and methodologies for measuring these effects suffer from a number of deficiencies including poor load tester design, statistically inaccurate aggregation, and improper attribution of effects. As shown in the paper, these pitfalls can often result in misleading conclusions. In this paper, we develop a methodology for statistically rigorous performance evaluation and performance factor attribution for server workloads. First, we find that careful design of the server load tester can ensure high quality performance evaluation, and empirically demonstrate the inaccuracy of load testers in previous work. Learning from the design flaws in prior work, we design and develop a modular load tester platform, Treadmill, that overcomes pitfalls of existing tools. Next, utilizing Treadmill, we construct measurement and analysis procedures that can properly attribute performance factors. We rely on statistically-sound performance evaluation and quantile regression, extending it to accommodate the idiosyncrasies of server systems. Finally, we use our augmented methodology to evaluate the impact of common server hardware features with Facebook production workloads on production hardware. We decompose the effects of these features on request tail latency and demonstrate that our evaluation methodology provides superior results, particularly in capturing complicated and counter-intuitive performance behaviors. By tuning the hardware features as suggested by the attribution, we reduce the 99th-percentile latency by 43% and its variance by 93%. David Meisner, Jason Mars, Lingjia Tang |
ISCA | 2 |
| 2015 | Adrenaline: Pinpointing and reining in tail queries with quick voltage boostingabstractReducing the long tail of the query latency distribution in modern warehouse scale computers is critical for improving performance and quality of service of workloads such as Web Search and Memcached. Traditional turbo boost increases a processor's voltage and frequency during a coarse-grain sliding window, boosting all queries that are processed during that window. However, the inability of such a technique to pinpoint tail queries for boosting limits its tail reduction benefit. In this work, we propose Adrenaline, an approach to leverage finer granularity, 10's of nanoseconds, voltage boosting to effectively rein in the tail latency with query-level precision. Two key insights underlie this work. First, emerging finer granularity voltage/frequency boosting is an enabling mechanism for intelligent allocation of the power budget to precisely boost only the queries that contribute to the tail latency; and second, per-query characteristics can be used to design indicators for proactively pinpointing these queries, triggering boosting accordingly. Based on these insights, Adrenaline effectively pinpoints and boosts queries that are likely to increase the tail distribution and can reap more benefit from the voltage/frequency boost. By evaluating under various workload configurations, we demonstrate the effectiveness of our methodology. We achieve up to a 2.50x tail latency improvement for Memcached and up to a 3.03x for Web Search over coarse-grained DVFS given a fixed boosting power budget. When optimizing for energy reduction, Adrenaline achieves up to a 1.81x improvement for Memcached and up to a 1.99x for Web Search over coarse-grained DVFS. Chang-Hong Hsu, Michael Laurenzano, David Meisner, Thomas F. Wenisch, Jason Mars, Lingjia Tang, Ronald G. Dreslinski |
HPCA | 4 |
| 2014 | The Mystery Machine: End-to-end Performance Analysis of Large-scale Internet Services
Michael Chow, David Meisner, Jason Flinn, Daniel Peek, Thomas F. Wenisch |
OSDI | 2 |
| 2013 | Thin servers with smart pipes: designing SoC accelerators for memcachedabstractDistributed in-memory key-value stores, such as memcached, are central to the scalability of modern internet services. Current deployments use commodity servers with high-end processors. However, given the cost-sensitivity of internet services and the recent proliferation of volume low-power System-on-Chip (SoC) designs, we see an opportunity for alternative architectures. We undertake a detailed characterization of memcached to reveal performance and power inefficiencies. Our study considers both high-performance and low-power CPUs and NICs across a variety of carefully-designed benchmarks that exercise the range of memcached behavior. We discover that, regardless of CPU microarchitecture, memcached execution is remarkably inefficient, saturating neither network links nor available memory bandwidth. Instead, we find performance is typically limited by the per-packet processing overheads in the NIC and OS kernel---long code paths limit CPU performance due to poor branch predictability and instruction fetch bottlenecks. Kevin T. Lim, David Meisner, Ali G. Saidi, Parthasarathy Ranganathan, Thomas F. Wenisch |
ISCA | 2 |
| 2012 | DreamWeaver: architectural support for deep sleepabstractNumerous data center services exhibit low average utilization leading to poor energy efficiency. Although CPU voltage and frequency scaling historically has been an effective means to scale down power with utilization, transistor scaling trends are limiting its effectiveness and the CPU is accounting for a shrinking fraction of system power. Recent research advocates the use of full-system idle low-power modes to combat energy losses, as such modes provide the deepest power savings with bounded response time impact. However, the trend towards increasing cores per die is undermining the effectiveness of these sleep modes, particularly for request-parallel data center applications, because the independent idle periods across individual cores are unlikely to align by happenstance. David Meisner, Thomas F. Wenisch |
ASPLOS | 1 |
| 2012 | MultiScale: memory system DVFS with multiple memory controllersabstractThe fraction of server energy consumed by the memory system has been increasing rapidly and is now on par with that consumed by processors. Recent work demonstrates that substantial memory energy can be saved with only a small, tightly-controlled performance degradation using memory Dynamic Frequency and Voltage Scaling (DVFS). Prior studies consider only servers with a single memory controller (MC); however, multicore server processors have begun to incorporate multiple MCs. We propose MultiScale, the first technique to coordinate DVFS across multiple MCs, memory channels, and memory devices. Under operating system control, MultiScale monitors application bandwidth requirements across MCs. It then uses a heuristic algorithm to select and apply a frequency combination that will minimize the overall system energy within user-specified per-application performance constraints. Our results demonstrate that MultiScale reduces system energy consumption significantly, compared to prior approaches, while respecting the user-specified performance constraints. Qingyuan Deng, David Meisner, Abhishek Bhattacharjee, Thomas F. Wenisch, Ricardo Bianchini |
ISLPED | 2 |
| 2012 | BigHouse: A simulation infrastructure for data center systemsabstractRecently, there has been an explosive growth in Internet services, greatly increasing the importance of data center systems. Applications served from “the cloud” are driving data center growth and quickly overtaking traditional workstations. Although there are a many tools for evaluating components of desktop and server architectures in detail, scalable modeling tools are noticeably missing. We describe BigHouse a simulation infrastructure for data center systems. Instead of simulating servers using detailed microarchitectural models, BigHouse raises the level of abstraction. Using a combination of queuing theory and stochastic modeling, BigHouse can simulate server systems in minutes rather than hours. BigHouse leverages statistical simulation techniques to limit simulation turnaround time to the minimum runtime needed for a desired accuracy. In this paper, we introduce BigHouse, describe its design, and present case studies for how it has already been applied to build and validate models of data center workloads and systems. Furthermore, we describe statistical techniques incorporated into BigHouse to accelerate and parallelize its simulations, and demonstrate its scalability to model large cluster systems while maintaining reasonable simulation time. David Meisner, Thomas F. Wenisch |
ISPASS | 1 |
| 2012 | CoScale: Coordinating CPU and Memory System DVFS in Server SystemsabstractRecent work has introduced memory system dynamic voltage and frequency scaling (DVFS), and has suggested that balanced scaling of both CPU and the memory system is the most promising approach for conserving energy in server systems. In this paper, we first demonstrate that CPU and memory system DVFS often conflict when performed independently by separate controllers. In response, we propose Co Scale, the first method for effectively coordinating these mechanisms under performance constraints. Co Scale relies on execution profiling of each core via (existing and new) performance counters, and models of core and memory performance and power consumption. Co Scale explores the set of possible frequency settings in such a way that it efficiently minimizes the full-system energy consumption within the performance bound. Our results demonstrate that, by effectively coordinating CPU and memory power management, Co Scale conserves a significant amount of system energy compared to existing approaches, while consistently remaining within the prescribed performance bounds. The results also show that Co Scale conserves almost as much system energy as an offline, idealized approach. Qingyuan Deng, David Meisner, Abhishek Bhattacharjee, Thomas F. Wenisch, Ricardo Bianchini |
MICRO | 2 |
| 2011 | MemScale: active low-power modes for main memoryabstractMain memory is responsible for a large and increasing fraction of the energy consumed by servers. Prior work has focused on exploiting DRAM low-power states to conserve energy. However, these states require entire DRAM ranks to be idled, which is difficult to achieve even in lightly loaded servers. In this paper, we propose to conserve memory energy while improving its energy-proportionality by creating active low-power modes for it. Specifically, we propose MemScale, a scheme wherein we apply dynamic voltage and frequency scaling (DVFS) to the memory controller and dynamic frequency scaling (DFS) to the memory channels and DRAM devices. MemScale is guided by an operating system policy that determines the DVFS/DFS mode of the memory subsystem based on the current need for memory bandwidth, the potential energy savings, and the performance degradation that applications are willing to withstand. Our results demonstrate that MemScale reduces energy consumption significantly compared to modern memory energy management approaches. We conclude that the potential benefits of the MemScale mechanisms and policy more than compensate for their small hardware cost. Qingyuan Deng, David Meisner, Luiz E. Ramos, Thomas F. Wenisch, Ricardo Bianchini |
ASPLOS | 2 |
| 2011 | Power management of online data-intensive servicesabstractMuch of the success of the Internet services model can be attributed to the popularity of a class of workloads that we call Online Data-Intensive (OLDI) services. These workloads perform significant computing over massive data sets per user request but, unlike their offline counterparts (such as MapReduce computations), they require responsiveness in the sub-second time scale at high request rates. Large search products, online advertising, and machine translation are examples of workloads in this class. Although the load in OLDI services can vary widely during the day, their energy consumption sees little variance due to the lack of energy proportionality of the underlying machinery. The scale and latency sensitivity of OLDI workloads also make them a challenging target for power management techniques. David Meisner, Christopher M. Sadler, Luiz André Barroso, Wolf-Dietrich Weber, Thomas F. Wenisch |
ISCA | 1 |
| 2011 | Does low-power design imply energy efficiency for data centers?
David Meisner, Thomas F. Wenisch |
ISLPED | 1 |
| 2011 | Towards a scalable data center-level evaluation methodologyabstractAs the popularity of Internet services continues to rise, the need to understand the design of the data center systems hosting these workloads becomes increasingly important. Unfortunately, research in this area has been stifled, primarily due to a lack of tools, workloads, and rigorous evaluation methodology. Traditional tools, such as architectural simulators, do not directly address data center-level issues and do not scale to simulate the thousands of machines needed for data center research. We introduce, Stochastic Queuing Simulation (SQS), our methodology for characterization and evaluation of data center systems. By leveraging techniques from stochastic modeling, queuing theory and statistical sampling, SQS uses discrete-event simulation to drive models that scale to tens of thousands of machines. Whereas detailed architectural simulations can last hours or days, SQS turnaround time is typically on the order of tens of minutes to an hour. Furthermore, computation can be distributed across cores and machines, achieving speedup using commodity clusters. David Meisner, Thomas F. Wenisch |
ISPASS | 1 |
| 2011 | The PowerNap Server ArchitectureabstractData center power consumption is growing to unprecedented levels: the EPA estimates U.S. data centers will consume 100 billion kilowatt hours annually by 2011. Much of this energy is wasted in idle systems: in typical deployments, server utilization is below 30%, but idle servers still consume 60% of their peak power draw. Typical idle periods---though frequent---last seconds or less, confounding simple energy-conservation approaches. In this article, we propose PowerNap, an energy-conservation approach where the entire system transitions rapidly between a high-performance active state and a near-zero-power idle state in response to instantaneous load. Rather than requiring fine-grained power-performance states and complex load-proportional operation from individual system components, PowerNap instead calls for minimizing idle power and transition time, which are simpler optimization goals. Based on the PowerNap concept, we develop requirements and outline mechanisms to eliminate idle power waste in enterprise blade servers. Because PowerNap operates in low-efficiency regions of current blade center power supplies, we introduce the Redundant Array for Inexpensive Load Sharing (RAILS), a power provisioning approach that provides high conversion efficiency across the entire range of PowerNap’s power demands. Using utilization traces collected from enterprise-scale commercial deployments, we demonstrate that, together, PowerNap and RAILS reduce average server power consumption by 74%. David Meisner, Brian T. Gold, Thomas F. Wenisch |
ACM Trans. Comput. Syst. | 1 |
| 2010 | Power routing: dynamic power provisioning in the data centerabstractData center power infrastructure incurs massive capital costs, which typically exceed energy costs over the life of the facility. To squeeze maximum value from the infrastructure, researchers have proposed over-subscribing power circuits, relying on the observation that peak loads are rare. To ensure availability, these proposals employ power capping, which throttles server performance during utilization spikes to enforce safe power budgets. However, because budgets must be enforced locally -- at each power distribution unit (PDU) -- local utilization spikes may force throttling even when power delivery capacity is available elsewhere. Moreover, the need to maintain reserve capacity for fault tolerance on power delivery paths magnifies the impact of utilization spikes. Steven Pelley, David Meisner, Pooya Zandevakili, Thomas F. Wenisch, Jack Underwood |
ASPLOS | 2 |
| 2010 | Peak power modeling for data center servers with switched-mode power suppliesabstractAccurately modeling server power consumption is critical in designing data center power provisioning infrastructure. However, to date, most research proposals have used average CPU utilization to infer the power consumption of clusters, typically averaging over tens of minutes per observation. We demonstrate that average CPU utilization is not sufficient to predict peak power consumption accurately. By characterizing the relationship between server utilization and power supply behavior, we can more accurately model the actual peak power consumption. Finally, we introduce a new operating system metric that can capture the needed information to design for peak power with low overhead. David Meisner, Thomas F. Wenisch |
ISLPED | 1 |
| 2009 | PowerNap: eliminating server idle powerabstractData center power consumption is growing to unprecedented levels: the EPA estimates U.S. data centers will consume 100 billion kilowatt hours annually by 2011. Much of this energy is wasted in idle systems: in typical deployments, server utilization is below 30%, but idle servers still consume 60% of their peak power draw. Typical idle periods though frequent--last seconds or less, confounding simple energy-conservation approaches. David Meisner, Brian T. Gold, Thomas F. Wenisch |
ASPLOS | 1 |
| 2007 | Hardware libraries: An architecture for economic acceleration in soft multi-core environmentsabstractIn single processor architectures, computationally- intensive functions are typically accelerated using hardware accelerators, which exploit the concurrency in the function code to achieve a significant speedup over software. The increased design constraints from power density and signal delay have shifted processor architectures in general towards multi-core designs. The migration to multi-core designs introduces the possibility of sharing hardware accelerators between cores. In this paper, we propose the concept of a hardware library, which is a pool of accelerated functions that are accessible by multiple cores. We find that sharing provides significant reductions in the area, logic usage and leakage power required for hardware acceleration. Contention for these units may exist in certain cases; however, the savings in terms of chip area are more appealing to many applications, particularly the embedded domain. We study the performance implications for our proposal using various multi-core arrangements, with actual implementations in FPGA fabrics. FPGAs are particularly appealing due to their cost effectiveness and the attained area savings enable designers to easily add functionality without significant chip revision. Our results show that is possible to save up to 37% of a chip's available logic and interconnect resources at a negligible impact (< 3%) to the performance. David Meisner, Sherief Reda |
ICCD | 1 |