EDBT 2026 Demo / reviewers in the wild / expert
Grigorios Magklis
dblp:65/3015
· DBLP profile ↗
16ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Processor architecture and microarchitecture · 33% Energy-efficient computing · 24% Parallel and multicore computing · 15% | |
| Software engineering, system software, and programming languages
2 papers |
Runtime systems and virtual machines · 82% Compilers and program optimization · 18% |
Topics — the 27 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
chip multiprocessor |
0.2 | 2 | 2010 | Thread-management techniques to maximize efficiency in multicore and simultaneous multithreaded microprocessors · ACM Trans. Archit. Code Optim. 2010 Understanding the Thermal Implications of Multi-Core Architectures · IEEE Trans. Parallel Distributed Syst. 2007 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.2 | 4 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 Dynamic frequency and voltage control for a multiple clock domain microarchitecture · MICRO 2002 |
Energy-efficient computing
thermal management |
0.1 | 2 | 2007 | Understanding the Thermal Implications of Multi-Core Architectures · IEEE Trans. Parallel Distributed Syst. 2007 Distributing the Frontend for Temperature Reduction · HPCA 2005 |
Memory systems
cache coherence |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Memory systems › cache coherence
conflict management |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Energy-efficient computing
power management |
0.1 | 1 | 2010 | Thread-management techniques to maximize efficiency in multicore and simultaneous multithreaded microprocessors · ACM Trans. Archit. Code Optim. 2010 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.1 | 1 | 2010 | Thread-management techniques to maximize efficiency in multicore and simultaneous multithreaded microprocessors · ACM Trans. Archit. Code Optim. 2010 |
Parallel and multicore computing › parallel programming runtimes
thread management |
0.1 | 1 | 2010 | Thread-management techniques to maximize efficiency in multicore and simultaneous multithreaded microprocessors · ACM Trans. Archit. Code Optim. 2010 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Energy-efficient computing › microprocessor power management
multiple clock domain processor |
0.1 | 2 | 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency Scaling · HPCA 2002 |
Energy-efficient computing › thermal management
dynamic thermal management |
0.1 | 1 | 2007 | Understanding the Thermal Implications of Multi-Core Architectures · IEEE Trans. Parallel Distributed Syst. 2007 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture › front-end
front-end design |
0.1 | 1 | 2005 | Distributing the Frontend for Temperature Reduction · HPCA 2005 |
Electronic design automation › design for manufacturability
area fill synthesis |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
Integrated circuit design › asynchronous circuit design
globally asynchronous locally synchronous design |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
High-performance computing
performance optimization |
0.0 | 1 | 2004 | Dynamically Trading Frequency for Complexity in a GALS Microprocessor · MICRO 2004 |
Integrated circuit design › clocking
multiple clock domain |
0.0 | 1 | 2002 | Dynamic frequency and voltage control for a multiple clock domain microarchitecture · MICRO 2002 |
Distributed systems
concurrency control |
0.0 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Storage systems › file systems
versioning |
0.0 | 1 | 2010 | A Dynamically Adaptable Hardware Transactional Memory · MICRO 2010 |
Parallel and multicore computing › parallel programming runtimes › thread management
thread migration |
0.0 | 1 | 2007 | Understanding the Thermal Implications of Multi-Core Architectures · IEEE Trans. Parallel Distributed Syst. 2007 |
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture |
0.0 | 1 | 2005 | Distributing the Frontend for Temperature Reduction · HPCA 2005 |
Processor architecture and microarchitecture › instruction fetch
trace cache |
0.0 | 1 | 2005 | Distributing the Frontend for Temperature Reduction · HPCA 2005 |
Compilers and program optimization
binary rewriting |
0.0 | 1 | 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain Microprocessor · ISCA 2003 |
Methods — techniques the papers use, named apart from their topics
speculative instruction-fusion optimization · 0.4cycle-accurate simulation · 0.4issue queue prioritization · 0.1frequency/voltage scaling · 0.1eager and lazy versioning · 0.1conflict prediction · 0.1simulation · 0.1microarchitectural simulation · 0.1phase detection · 0.0hardware-based control · 0.0profiling · 0.0binary rewriting · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Accelerating ML Recommendation with over a Thousand RISC-V/Tensor Processors on Esperanto's ET-SoC-1 ChipabstractThe ET-SoC-1 has over a thousand RISC-V processors on a single TSMC 7nm chip, including: • 1088 energy-efficient ET-Minion 64-bit RISC-V in-order cores each with a vector/tensor unit • 4 high-performance ET-Maxion 64-bit RISC-V out-of-order cores • >160 million bytes of on-chip SRAM • Interfaces for large external memory with low-power LPDDR4x DRAM and eMMC FLASH • PCIe x8 Gen4 and other common I/O interfaces • Innovative low-power architecture and circuit techniques allows entire chip to • Compute at peak rates of 100 to 200 TOPS • Operate using under 20 watts for ML recommendation workloads David R. Ditzel, Roger Espasa, Nivard Aymerich, Allen Baum, Tom Berg, Jim Burr, Eric Hao, Jayesh Iyer, Miquel Izquierdo, Shankar Jayaratnam, Darren Jones, Chris Klingner, Stephen Lee, Marc Lupon, Grigorios Magklis, Bojan Maric, Rajib Nath, Mike Neilly, J. Duane Northcutt, Bill Orner, Jose Renau, Gerard Reves, Xavier Reves, Tom Riordan, Pedro Sanchez, Sridhar Samudrala, Guillem Sole, Raymond Tang, Tommy Thorn, Sebastia Tortella, Daniel Yau |
HCS | 16 |
| 2014 | Speculative hardware/software co-designed floating-point multiply-add fusionabstractA Fused Multiply-Add (FMA) instruction is currently available in many general-purpose processors. It increases performance by reducing latency of dependent operations and increases precision by computing the result as an indivisible operation with no intermediate rounding. However, since the arithmetic behavior of a single-rounding FMA operation is different than independent FP multiply followed by FP add instructions, some algorithms require significant revalidation and rewriting efforts to work as expected when they are compiled to operate with FMA--a cost that developers may not be willing to pay. Because of that, abundant legacy applications are not able to utilize FMA instructions. In this paper we propose a novel HW/SW collaborative technique that is able to efficiently execute workloads with increased utilization of FMA, by adding the option to get the same numerical result as separate FP multiply and FP add pairs. In particular, we extended the host ISA of a HW/SW co-designed processor with a new Combined Multiply-Add (CMA) instruction that performs an FMA operation with an intermediate rounding. This new instruction is used by a transparent dynamic translation software layer that uses a speculative instruction-fusion optimization to transform FP multiply and FP add sequences into CMA instructions. The FMA unit has been slightly modified to support both single-rounding and double-rounding fused instructions without increasing their latency and to provide a conservative fall-back path in case of mispeculation. Evaluation on a cycle-accurate timing simulator showed that CMA improved SPECfp performance by 6.3% and reduced executed instructions by 4.7%. Marc Lupon, Enric Gibert, Grigorios Magklis, Sridhar Samudrala, Raúl Martínez, Kyriakos Stavrou, David R. Ditzel |
ASPLOS | 3 |
| 2011 | Thread shuffling: combining DVFS and thread migration toreduce energy consumptions for multi-core systems
Qiong Cai, José González 0002, Grigorios Magklis, Pedro Chaparro, Antonio González 0001 |
ISLPED | 3 |
| 2010 | A Dynamically Adaptable Hardware Transactional MemoryabstractMost Hardware Transactional Memory (HTM) implementations choose fixed version and conflict management policies at design time. While eager HTM systems store transactional state in-place in memory and resolve conflicts when they are produced, lazy HTM systems buffer the transactional state in specialized hardware and defer the resolution of conflicts until commit time. Each scheme has its strengths and weaknesses, but, unfortunately, both approaches are too inflexible in the way they manage data versioning and transactional contention. Thus, fixed HTM systems may result in a significant performance opportunity loss when they execute complex transactional applications. In this paper, we present DynTM (Dynamically Adaptable HTM), the first fully-flexible HTM system that permits the simultaneous execution of transactions using complementary version and conflict management strategies. In the heart of DynTM is a novel coherence protocol that allows tracking conflicts among eager and lazy transactions. Both the eager and the lazy execution modes of DynTM exhibit very high performance compared to modern HTM systems. For example, the DynTM lazy execution mode implements local commits to improve on previous proposals. In addition, lazy transactions share the majority of hardware support with eager transactions, reducing substantially the hardware cost compared to other lazy HTM systems. By utilizing a simple predictor to decide the best execution mode for each transaction at runtime, DynTM obtains an average speedup of 34% over HTM systems that employ fixed version and conflict management policies. Marc Lupon, Grigorios Magklis, Antonio González 0001 |
MICRO | 2 |
| 2010 | Thread-management techniques to maximize efficiency in multicore and simultaneous multithreaded microprocessorsabstractWe provide an analysis of thread-management techniques that increase performance or reduce energy in multicore and Simultaneous Multithreaded (SMT) cores. Thread delaying reduces energy consumption by running the core containing the critical thread at maximum frequency while scaling down the frequency and voltage of the cores containing noncritical threads. In this article, we provide an insightful breakdown of thread delaying on a simulated multi-core microprocessor. Thread balancing improves overall performance by giving higher priority to the critical thread in the issue queue of an SMT core. We provide a detailed breakdown of performance results for thread-balancing, identifying performance benefits and limitations. For those benchmarks where a performance benefit is not possible, we introduce a novel thread-balancing mechanism on an SMT core that can reduce energy consumption. We have performed a detailed study on an Intel microprocessor simulator running parallel applications. Thread delaying can reduce energy consumption by 4% to 44% with negligible performance loss. Thread balancing can increase performance by 20% or can reduce energy consumption by 23%. Ryan N. Rakvic, Qiong Cai, José González 0002, Grigorios Magklis, Pedro Chaparro, Antonio González 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2009 | FASTM: A Log-based Hardware Transactional Memory with Fast Abort RecoveryabstractVersion management, one of the key design dimensions of hardware transactional memory (HTM) systems, defines where and how transactional modifications are stored. Current HTM systems use either eager or lazy version management. Eager systems that keep new values in-place while they hold old values in a software log, suffer long delays when aborts are frequent because the pre-transactional state is recovered by software. Lazy systems that buffer new values in specialized hardware offer complex and inefficient solutions to handle hardware overflows, which are common in applications with coarse-grain transactions. In this paper, we present FASTM, an eager log-based HTM that takes advantage of the processorpsilas cache hierarchy to provide fast abort recovery. FASTM uses a novel coherence protocol to buffer the transactional modifications in the first level cache and to keep the non-speculative values in the higher levels of the memory hierarchy. This mechanism allows fast abort recovery of transactions that do not overflow the first level cache resources. Contrary to lazy HTM systems, committing transactions do not have to perform any actions in order to make their results visible to the rest of the system. FASTM keeps the pre-transactional state in a software-managed log as well, which permits the eviction of speculative values and enables transparent execution even in the case of cache overflow. This approach simplifies eviction policies without degrading performance, because it only falls back to a software abort recovery for transactions whose modified state has overflowed the cache. Simulation results show that FASTM achieves a speed-up of 43% compared to LogTM-SE, improving the scalability of applications with coarse-grain transactions and obtaining similar performance to an ideal eager HTM with zero-cost abort recovery. Marc Lupon, Grigorios Magklis, Antonio González 0001 |
PACT | 2 |
| 2008 | Meeting points: using thread criticality to adapt multicore hardware to parallel regionsabstractWe present a novel mechanism, called meeting point thread characterization, to dynamically detect critical threads in a parallel region. We define the critical thread the one with the longest completion time in the parallel region. Knowing the criticality of each thread has many potential applications. In this work, we propose two applications: thread delaying for multi-core systems and thread balancing for simultaneous multi-threaded (SMT) cores. Thread delaying saves energy consumptions by running the core containing the critical thread at maximum frequency while scaling down the frequency and voltage of the cores containing non-critical threads. Thread balancing improves overall performance by giving higher priority to the critical thread in the issue queue of an SMT core. Our experiments on a detailed microprocessor simulator with the Recognition, Mining, and Synthesis applications from Intel research laboratory reveal that thread delaying can achieve energy savings up to more than 40% with negligible performance loss. Thread balancing can improve performance from 1% to 20%. Qiong Cai, José González 0002, Ryan N. Rakvic, Grigorios Magklis, Pedro Chaparro, Antonio González 0001 |
PACT | 4 |
| 2008 | Thread fusionabstractThis work proposes Thread Fusion as an effective way of reducing power consumption when a Simultaneous Multi-Threaded (SMT) core is executing two threads from a homogeneous parallel application. Two dynamic instances of the same static instruction, each from a different thread are merged (fused) into a single instruction, consuming half of the resources of front-end pipeline stages. When the fused instruction is executed, it is cloned and it proceeds at full bandwidth. Our simulation results show average energy reduction of 10% with less than 1% impact on performance. José González 0002, Qiong Cai, Pedro Chaparro, Grigorios Magklis, Ryan N. Rakvic, Antonio González 0001 |
ISLPED | 4 |
| 2007 | Understanding the Thermal Implications of Multi-Core ArchitecturesabstractMulti-core architectures are becoming the main design paradigm for current and future processors. The main reason is that multi-core designs provide an effective way of overcoming ILP limitations by exploiting TLP. In addition, it is a power- and complexity-effective way of taking advantage of the huge number of transistors that can be integrated on a chip. On the other hand, today’s higher than ever power densities have made temperature one of the main limitations of microprocessor evolution. Thermal management in multi-core architectures is a fairly new area. Some works have addressed dynamic thermal management in bi/quad-core architectures. This work provides insight and explores different alternatives for thermal management in multi-core architectures with 16 cores. Schemes employing both energy reduction and activity migration are explored and improvements for thread migration schemes are proposed. Pedro Chaparro, José González 0002, Grigorios Magklis, Qiong Cai, Antonio González 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2006 | Independent front-end and back-end dynamic voltage scaling for a GALS microarchitectureabstractIn recent years, Globally Asynchronous Locally Synchronous (GALS) designs and dynamic voltage scaling (DVS) have emerged as some of the most popular approaches to address the ever increasing microprocessor energy consumption. In this work, we propose two on-line algorithms for adjusting dynamically, and independently, the voltage and frequency of the front-end and back-end domains of a novel two-domain microprocessor. We evaluate our mechanisms for both internal and external voltage regulators, and we present optimal dynamic voltage scaling results for the proposed microarchitecture. Our schemes achieve average improvement of 12% of the energy-delay2 metric, when using internal voltage regulators. Grigorios Magklis, Pedro Chaparro, José González 0002, Antonio González 0001 |
ISLPED | 1 |
| 2005 | Distributing the Frontend for Temperature ReductionabstractDue to increasing power densities, both on-chip average and peak temperatures are fast becoming a serious bottleneck in processor design. This is due to the cost of removing the heat generated, and the performance impact of dealing with thermal emergencies. So far microarchitectural techniques to control temperature have mainly focused on the processor backend (in particular the execution units), whereas the frontend has not received much attention. However, as the temperature of the backend remains controlled and the processor throughput increases, the heat dissipated by the frontend becomes more significant, and one of the major contributors to the total average temperature. This paper proposes and evaluates a distributed frontend for clustered microarchitectures that is able to reduce power density and temperature. First, a distributed mechanism for renaming and committing instructions is proposed. Second, a sub-banked trace cache with a bank hopping mechanism is presented. Finally, a method to improve the sub-banking is proposed based on a biased mapping function to distribute bank accesses to balance temperature. Pedro Chaparro, Grigorios Magklis, José González 0002, Antonio González 0001 |
HPCA | 2 |
| 2004 | Frontend Frequency-Voltage Adaptation for Optimal Energy-Delay^2abstractIn this paper, we present a clustered, multiple-clock domain (CMCD) microarchitecture that combines the benefits of both clustering and globally asynchronous locally synchronous (GALS) designs. We also present a mechanism for dynamically adapting the frequency and voltage of the frontend of the CMCD with the goal to optimize the energy-delay/sup 2/ product (ED2P). Our mechanism has minimal hardware cost, is entirely self-adjustable, does not depend on any thresholds, and achieves results close to optimal. We evaluate it on 16 SPEC 2000 applications and report 17.5% ED2P reduction on average (80% of the upper bound). Grigorios Magklis, José González 0002, Antonio González 0001 |
ICCD | 1 |
| 2004 | Dynamically Trading Frequency for Complexity in a GALS MicroprocessorabstractMicroprocessors are traditionally designed to provide "best overall" performance across a wide range of applications and operating environments. Several groups have proposed hardware techniques that save energy by "downsizing" hardware resources that are underutilized by the current application phase. Others have proposed a different energy-saving approach: dividing the processor into domains and dynamically changing the clock frequency and voltage within each domain during phases when the full domain frequency is not required. What has not been studied to date is how to exploit the adaptive nature of these approaches to improve performance rather than to save energy. In this paper, we describe an adaptive globally asynchronous, locally synchronous (GALS) microprocessor with a fixed global voltage and four independently clocked domains. Each domain is streamlined with modest hardware structures for very high clock frequency. Key structures can then be upsized on demand to exploit more distant parallelism, improve branch prediction, or increase cache capacity. Although doing so requires decreasing the associated domain frequency, other domain frequencies are unaffected. Our approach, therefore, is to maximize the throughput of each domain by finding the proper balance between the number of clock periods, and the clock frequency, for each application phase. To achieve this objective, we use novel hardware-based control techniques that accurately and efficiently capture the performance of all possible cache and queue configurations within a single interval, without having to resort to exhaustive online exploration or expensive offline profiling. Measuring across a broad suite of application benchmarks, we find that configuring our adaptive GALS processor just once per application yields 17.6% better performance, on average, than that of the "best overall" fully synchronous design. By adapting automatically to application phases, we can increase this advantage to more than 20%. Steven G. Dropsho, Greg Semeraro, David H. Albonesi, Grigorios Magklis, Michael L. Scott |
MICRO | 4 |
| 2003 | Profile-Based Dynamic Voltage and Frequency Scaling for a Multiple Clock Domain MicroprocessorabstractA Multiple Clock Domain (MCD) processor addresses the challenges of clock distribution and power dissipation by dividing a chip into several (coarse-grained) clock domains, allowing frequency and voltage to be reduced in domains that are not currently on the application’s critical path. Given a reconfiguration mechanism capable of choosing appropriate times and values for voltage/frequency scaling, an MCD processor has the potential to achieve significant energy savings with low performance degradation. Early work on MCD processors evaluated the potential for energy savings by manually inserting reconfiguration instructions into applications, or by employing an oracle driven by off-line analysis of (identical) prior program runs. Subsequent work developed a hardware-based on-line mechanism that averages 75–85% of the energy-delay improvement achieved via off-line analysis. In this paper we consider the automatic insertion of reconfiguration instructions into applications, using profiledriven binary rewriting. Profile-based reconfiguration introduces the need for “training runs” prior to production use of a given application, but avoids the hardware complexity of on-line reconfiguration. It also has the potential to yield significantly greater energy savings. Experimental results (training on small data sets and then running on larger, alternative data sets) indicate that the profile-driven approach is more stable than hardware-based reconfiguration, and yields virtually all of the energy-delay improvement achieved via off-line analysis. Grigorios Magklis, Michael L. Scott, Greg Semeraro, David H. Albonesi, Steven G. Dropsho |
ISCA | 1 |
| 2002 | Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency ScalingabstractAs clock frequency increases and feature size decreases, clock distribution and wire delays present a growing challenge to the designers of singly-clocked, globally synchronous systems. We describe an alternative approach, which we call a multiple clock domain (MCD) processor, in which the chip is divided into several clock domains, within which independent voltage and frequency scaling can be performed. Boundaries between domains are chosen to exploit existing queues, thereby minimizing inter-domain synchronization costs. We propose four clock domains, corresponding to the front end , integer units, floating point units, and load-store units. We evaluate this design using a simulation infrastructure based on SimpleScalar and Wattch. In an attempt to quantify potential energy savings independent of any particular on-line control strategy, we use off-line analysis of traces from a single-speed run of each of our benchmark applications to identify profitable reconfiguration points for a subsequent dynamic scaling run. Using applications from the MediaBench, Olden, and SPEC2000 benchmark suites, we obtain an average energy-delay product improvement of 20% with MCD compared to a modest 3% savings from voltage scaling a single clock and voltage system. Greg Semeraro, Grigorios Magklis, Rajeev Balasubramonian, David H. Albonesi, Sandhya Dwarkadas, Michael L. Scott |
HPCA | 2 |
| 2002 | Dynamic frequency and voltage control for a multiple clock domain microarchitectureabstractWe describe the design, analysis, and performance of an on-line algorithm to dynamically control the frequency/voltage of a Multiple Clock Domain (MCD) microarchitecture. The MCD microarchitecture allows the frequency/voltage of microprocessor regions to be adjusted independently and dynamically, allowing energy savings when the frequency of some regions can be reduced without significantly impacting performance. Our algorithm achieves on average a 19.0% reduction in Energy Per Instruction (EPI), a 3.2% increase in Cycles Per Instruction (CPI), a 16.7% improvement in Energy-Delay Product, and a Power Savings to Performance Degradation ratio of 4.6. Traditional frequency/voltage scaling techniques which apply reductions globally to a fully synchronous processor achieve a Power Savings to Performance Degradation ratio of only 2-3. Our Energy-Delay Product improvement is 85.5% of what has been achieved using an off-line algorithm. These results were achieved using a broad range of applications from the MediaBench, Olden, and Spec2000 benchmark suites using an algorithm we show to require minimal hardware resources. Greg Semeraro, David H. Albonesi, Steven G. Dropsho, Grigorios Magklis, Sandhya Dwarkadas, Michael L. Scott |
MICRO | 4 |