Charles Lefurgy

dblp:37/6803 · also Charles R. Lefurgy · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Energy-efficient computing · 62% Hardware reliability and fault tolerance · 14% Processor architecture and microarchitecture · 9%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%

Topics — the 25 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
power management
1.472019
Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015
Energy-efficient computing
datacenter power management
0.522019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012
Energy-efficient computing › datacenter power management
server power capping
0.422019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012
Processor architecture and microarchitecture › adaptive architecture
timing margin adaptation
0.412019
Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019
Hardware reliability and fault tolerance
timing guardband
0.322015
Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015
Active management of timing guardband to save energy in POWER7 · MICRO 2011
Integrated circuit design › variation-aware design
adaptive guardbanding
0.212015
Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015
Hardware reliability and fault tolerance
processor reliability
0.212015
Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015
Energy-efficient computing › power management › system-level power management
hierarchical power management
0.112012
SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012
Energy-efficient computing › power management › energy-efficient networking
network power management
0.112011
Power shifting in Thrifty Interconnection Network · HPCA 2011
Distributed systems › fault tolerance
high availability
0.112019
A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019
Embedded and real-time systems › real-time scheduling
multicore scheduling
0.112019
Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019
Cloud and datacenter computing › resource management
datacenter resource management
0.112007
Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007
Energy-efficient computing › power management
dynamic power management
0.112007
Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007
Energy-efficient computing
voltage and frequency scaling
0.112015
Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015
Cloud and datacenter computing › resource management
cloud resource management
0.012012
Accurate Fine-Grained Processor Power Proxies · MICRO 2012
Compilers and program optimization › code size reduction
code compression
0.021999
Evaluation of a High Performance Code Compression Method · MICRO 1999
Improving Code Density Using Compression Techniques · MICRO 1997
Embedded and real-time systems
embedded processor
0.021999
Evaluation of a High Performance Code Compression Method · MICRO 1999
Improving Code Density Using Compression Techniques · MICRO 1997
Hardware reliability and fault tolerance
aging
0.012011
Active management of timing guardband to save energy in POWER7 · MICRO 2011
Hardware reliability and fault tolerance
process variation
0.012011
Active management of timing guardband to save energy in POWER7 · MICRO 2011
Embedded and real-time systems › embedded processor
code compression
0.012000
Reducing Code Size with Run-Time Decompression · HPCA 2000
Memory systems › cache › CPU cache
instruction cache
0.012000
Reducing Code Size with Run-Time Decompression · HPCA 2000
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems
0.012007
Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007
Compilers and program optimization
code size reduction
0.012000
Reducing Code Size with Run-Time Decompression · HPCA 2000
Performance modeling and evaluation
benchmarking
0.011999
Evaluation of a High Performance Code Compression Method · MICRO 1999
Processor architecture and microarchitecture
instruction set architecture
0.011997
Improving Code Density Using Compression Techniques · MICRO 1997

Methods — techniques the papers use, named apart from their topics

priority-aware scheduling · 0.4power capping · 0.4critical path monitor · 0.4application scheduling and throttling · 0.4simulation · 0.3power model training · 0.1on-chip power sensing · 0.1control theory · 0.1CPU frequency scaling · 0.1workload trace analysis · 0.1reinforcement learning · 0.1multi-criteria reward · 0.1selective compression · 0.0cache miss profiling · 0.0post-compilation analysis · 0.0
YearPublicationVenuePosition
2019 A Scalable Priority-Aware Approach to Managing Data Center Server Power
abstract
Power management is a key component of modern data center design. Power managers must (1) ensure the costand energy-efficient utilization of the data center infrastructure, (2) maintain availability of the services provided by the center, and (3) address environmental concerns associated with the center's power consumption. While several power management techniques have been proposed and deployed in production data centers, there are still many challenges to comprehensive data center power management. This is particularly true in public cloud environments, where different jobs have different priority levels, and where high availability is critical. One example of the challenges facing public cloud data centers involves power capping. As power delivery must be highly reliable and tolerate wide variation in the load drawn by the data center components, the power infrastructure (e.g., power supplies, circuit breakers, UPS) has high redundancy and overprovisioning. During normal operation (i.e., typical server power demands, and no failures in the center), the power infrastructure is significantly underutilized. Power capping is a common solution to reduce this underutilization, by allowing more servers to be added safely (i.e., without power shortfalls) to the existing power infrastructure, and throttling power consumption in the infrequent cases where the demanded power exceeds the provisioned power capacity to avoid shortfalls. However, state-of-the-art power capping solutions are (1) not directly applicable to the redundant power infrastructure used in highly-available data centers; and (2) oblivious to differing workload priorities across the entire center when power consumption needs to be throttled, which can unnecessarily slow down high-priority work. To address this need, we develop CapMaestro, a new power management architecture with three key features for public cloud data centers. First, CapMaestro is designed to work with multiple power feeds (i.e., sources), and exploits server-level power capping to independently cap the load on each feed of a server. Second, CapMaestro uses a scalable, global priority-aware power capping approach, which accounts for power capacity at each level of the power distribution hierarchy. It exploits the underutilization of commonly-employed redundant power infrastructure at each level of the hierarchy to safely accommodate a much greater number of servers. Third, CapMaestro exploits stranded power (i.e., power budgets that are not utilized) in redundant power infrastructure to boost the performance of workloads in the data center. We add CapMaestro to a real cloud data center control plane, and demonstrate the effectiveness of all three key features. Using a large-scale data center simulation, we demonstrate that CapMaestro significantly and safely increases the number of servers for existing infrastructure. We also call out other key technical challenges the industry faces in data center power management.
Yang Li 0183, Charles Lefurgy, Karthick Rajamani, Malcolm Allen-Ware, Guillermo J. Silva, Daniel D. Heimsoth, Saugata Ghose, Onur Mutlu
HPCA2
2019 Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server
abstract
Active Timing Margin (ATM) is a technology that improves processor efficiency by reducing the pipeline timing margin with a control loop that adjusts voltage and frequency based on real-time chip environment monitoring. Although ATM has already been shown to yield substantial performance benefits, its full potential has yet to be unlocked. In this paper, we investigate how to maximize ATM's efficiency gain with a new means of exposing the inter-core speed variation: finetuning the ATM control loop. We conduct our analysis and evaluation on a production-grade POWER7+ system. On the POWER7+ server platform, we fine-tune the ATM control loop by programming its Critical Path Monitors, a key component of its ATM design that measures the cores' timing margins. With a robust stress-test procedure, we expose over 200 MHz of inherent inter-core speed differential by fine-tuning the percore ATM control loop. Exploiting this differential, we manage to double the ATM frequency gain over the static timing margin; this is not possible using conventional means, i.e. by setting fixedpoints for each core, because the corelevelmust account for chip-wide worst-case voltage variation. To manage the significant performance heterogeneity of fine-tuned systems, we propose application scheduling and throttling to manage the chip's process and voltage variation. Our proposal improves application performance by more than 10% over the static margin, almost doubling the 6% improvement of the default, unmanaged ATM system. Our technique is general enough that it can be adopted by any system that employs an active timing margin control loop.
Yazhou Zu, Daniel Richins, Charles Lefurgy, Vijay Janapa Reddi
HPCA3
2015 FreqLeak: A frequency step based method for efficient leakage power characterization in a system
abstract
Accurate estimation of leakage power at runtime requires post-silicon power measurements across a wide range of temperature and voltage conditions. Testing individual chips, especially at high-temperature corner conditions, is expensive in cost and time. We examine this problem in an industrial context and introduce FreqLeak, a frequency step based method for inexpensive and efficient leakage power characterization in a system. It enables a more thorough characterization than can be accomplished on a wafer prober alone due to time and equipment costs. Experimental evaluation on IBM POWER8 based systems demonstrates the efficiency of the proposed method, within an error of 5%. Further, we discuss the application of FreqLeak in system level power management.
Anand Haridass, Charles Lefurgy, Sreekanth Pai, Spandana Rachamalla, Francesco Campisano
ISLPED3
2015 Adaptive guardband scheduling to improve system-level efficiency of the POWER7+
abstract
The traditional guardbanding approach to ensure processor reliability is becoming obsolete because it always over-provisions voltage and wastes a lot of energy. As a next-generation alternative, adaptive guardbanding dynamically adjusts chip clock frequency and voltage based on timing margin measured at runtime. With adaptive guardbanding, voltage guardband is only provided when needed, thereby promising significant energy efficiency improvement.
Yazhou Zu, Charles Lefurgy, Jingwen Leng, Matthew Halpern, Michael S. Floyd, Vijay Janapa Reddi
MICRO2
2012 Accurate Fine-Grained Processor Power Proxies
abstract
There are not yet practical and accurate ways to directly measure core power in a microprocessor. This limits the granularity of measurement and control for computer power management. We overcome this limitation by presenting an accurate runtime per-core power proxy which closely estimates true core power. This enables new fine-grained microprocessor power management techniques at the core level. For example, cloud environments could manage and bill virtual machines for energy consumption associated with the core. The power model underlying our power proxy also enables energy-efficiency controllers to perform what-if analysis, instead of merely reacting to current conditions. We develop and validate a methodology for accurate power proxy training at both chip and core levels. Our implementation of power proxies uses on-chip logic in a high-performance multi-core processor and associated platform firmware. The power proxies account for full voltage and frequency ranges, as well as chip-to-chip process variations. For fixed clock frequency operation, a mean unsigned error of 1.8% for fine-grained 32ms samples across all workloads was achieved. For an interval of an entire workload, we achieve an average error of-0.2%. Similar results were achieved for voltage-scaling scenarios, too. We also present two sample applications of the power proxy: (1) per-core power billing for cloud computing services, and (2) simultaneous runtime energy saving comparisons among different power management policies without running each policy separately.
Wei Huang 0004, Charles Lefurgy, William Kuk, Alper Buyuktosunoglu, Michael S. Floyd, Karthick Rajamani, Malcolm Allen-Ware, Bishop Brock
MICRO2
2012 SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers
abstract
In today's data centers, precisely controlling server power consumption is an essential way to avoid system failures caused by power capacity overload or overheating due to increasingly high server density. While various power control strategies have been recently proposed, existing solutions are not scalable to control the power consumption of an entire large-scale data center, because these solutions are designed only for a single server or a rack enclosure. In a modern data center, however, power control needs to be enforced at three levels: rack enclosure, power distribution unit, and the entire data center, due to the physical and contractual power limits at each level. This paper presents SHIP, a highly scalable hierarchical power control architecture for large-scale data centers. SHIP is designed based on well-established control theory for analytical assurance of control accuracy and system stability. Empirical results on a physical testbed show that our control solution can provide precise power control, as well as power differentiations for optimized system performance and desired server priorities. In addition, our extensive simulation results based on a real trace file demonstrate the efficacy of our control solution in large-scale data centers composed of 5,415 servers.
Ming Chen 0002, Charles Lefurgy, Tom W. Keller
IEEE Trans. Parallel Distributed Syst.3
2011 Power shifting in Thrifty Interconnection Network
abstract
This paper presents two complementary techniques to manage the power consumption of large-scale systems with a packet-switched interconnection network. First, we propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically with little or no overhead by using inherent system events to timely trigger link activation or de-activation. Second, we propose Network Power Shifting (NPS) that dynamically shifts the power budget between the compute nodes and their corresponding network components. TIN activates and trains the links in the interconnection network, just-in-time before the network communication is about to happen, and thriftily puts them into a low-power mode when communication is finished, hence reducing unnecessary network power consumption. Furthermore, the compute nodes can absorb the extra power budget shifted from its attached network components and increase their processor frequency for higher performance with NPS. Our simulation results on a set of real-world workload traces show that TIN can achieve on average 60% network power reduction, with the support of only one low-power mode. When NPS is enabled, the two together can achieve 12% application performance improvement and 13% overall system energy reduction. Further performance improvement is possible if the compute nodes can speed up more and fully utilize the extra power budget reinvested from the thrifty network with more aggressive cooling support.
Jian Li 0059, Wei Huang 0004, Charles Lefurgy, Lixin Zhang 0002, Wolfgang E. Denzel, Richard R. Treumann, Kun Wang 0005
HPCA3
2011 Active management of timing guardband to save energy in POWER7
abstract
Microprocessor voltage levels include substantial margin to deal with process variation, system power supply variation, workload induced thermal and voltage variation, aging, random uncertainty, and test inaccuracy. This margin allows the microprocessor to operate correctly during worst-case conditions, but during typical conditions it is larger than necessary and wastes energy. We present a mechanism that reduces excess voltage margin by (1) introducing a critical path monitor (CPM) circuit that measures available timing margin in real-time, (2) coupling the CPM output to the clock generation circuit to adjust clock frequency within cycles in response to excess or inadequate timing margin, and (3) adjusting the processor voltage level periodically in firmware to achieve a specified average clock frequency target. We implemented this mechanism in a prototype IBM POWER7 server. During better-than-worst case conditions our guardband management mechanism reduces the average voltage setting 137-152 mV below nominal, resulting in average processor power reduction of 24% with no performance loss while running industry-standard benchmarks.
Charles Lefurgy, Alan J. Drake, Michael S. Floyd, Malcolm Allen-Ware, Bishop Brock, José A. Tierno, John B. Carter
MICRO1
2010 Adaptive energy management features of the POWER7TM processor
Michael S. Floyd, Bishop Brock, Malcolm Allen-Ware, Karthick Rajamani, Alan J. Drake, Charles Lefurgy, Lorena Pesantez
Hot Chips Symposium6
2009 SHIP: Scalable Hierarchical Power Control for Large-Scale Data Centers
abstract
In today's data centers, precisely controlling server power consumption is an essential way to avoid system failures caused by power capacity overload or overheating due to increasingly high server density. While various power control strategies have been recently proposed, existing solutions are not scalable to control the power consumption of an entire large-scale data center, because these solutions are designed only for a single server or a rack enclosure. In a modern data center, however, power control needs to be enforced at three levels: rack enclosure, power distribution unit, and the entire data center, due to the physical and contractual power limits at each level. This paper presents SHIP, a highly scalable hierarchical power control architecture for large-scale data centers. SHIP is designed based on well-established control theory for analytical assurance of control accuracy and system stability. Empirical results on a physical testbed show that our control solution can provide precise power control, as well as power differentiations for optimized system performance. In addition, our extensive simulation results based on a real trace file demonstrate the efficacy of our control solution in large-scale data centers composed of thousands of servers.
Ming Chen 0002, Charles Lefurgy, Tom W. Keller
PACT3
2009 Thrifty interconnection network for HPC systems
abstract
We propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically to save power with little or no overhead by using inherent system events to overlap the link activation or de-activation time. Our simulation results on a set of real world HPC workload traces show on average 35% network power reduction.
Jian Li 0059, Lixin Zhang 0002, Charles Lefurgy, Richard R. Treumann, Wolfgang E. Denzel
ICS3
2008 Power management solutions for computer systems and datacenters
abstract
The growing power and cooling requirements of high-density computing systems pose significant challenges for the design and operation of computers and their facilities. The rising operating expenses for datacenters demand the implementation of energy-efficient technologies and the best power management solutions. This tutorial addresses power management and cooling solutions from the individual computer system level to the datacenter. The audience will learn about the fundamental nature of the problems, approaches to developing solutions, available commercial solutions, and current research directions.
Karthick Rajamani, Charles Lefurgy, Soraya Ghiasi, Juan C. Rubio, Heather Hanson, Tom W. Keller
ISLPED2
2007 Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning
abstract
Electrical power management in large-scale IT systems such as commercial data- centers is an application area of rapidly growing interest from both an economic and ecological perspective, with billions of dollars and millions of metric tons of CO2 emissions at stake annually. Businesses want to save power without sac- rificing performance. This paper presents a reinforcement learning approach to simultaneous online management of both performance and power consumption. We apply RL in a realistic laboratory testbed using a Blade cluster and dynam- ically varying HTTP workload running on a commercial web applications mid- dleware platform. We embed a CPU frequency controller in the Blade servers’ firmware, and we train policies for this controller using a multi-criteria reward signal depending on both application performance and CPU power consumption. Our testbed scenario posed a number of challenges to successful use of RL, in- cluding multiple disparate reward functions, limited decision sampling rates, and pathologies arising when using multiple sensor readings as state variables. We describe innovative practical solutions to these challenges, and demonstrate clear performance improvements over both hand-designed policies as well as obvious “cookbook” RL implementations.
Gerald Tesauro, Rajarshi Das, Hoi Y. Chan, Jeffrey O. Kephart, David W. Levine, Freeman L. Rawson III, Charles Lefurgy
NIPS7
2005 Improving energy efficiency by making DRAM less randomly accessed
abstract
Existing techniques manage power for the main memory by passively monitoring the memory traffic, and based on which, predict when to power down and into which low-power state to transition. However, passively monitoring the memory traffic can be far from being effective as idle periods between consecutive memory accesses are often too short for existing power-management techniques to take full advantage of the deeper power-saving state implemented in modern DRAM architectures. In this paper, we propose a new technique that will actively reshape the memory traffic to coalesce short idle periods --- which were previously unusable for power management --- into longer ones, thus enabling existing techniques to effectively exploit idleness in the memory
Hai Huang 0002, Kang G. Shin, Charles Lefurgy, Tom W. Keller
ISLPED3
2004 Improving Server Performance on Transaction Processing Workloads by Enhanced Data Placement
abstract
Modern servers access large volumes of data while running commercial workloads. The data is typically spread among several storage devices (e.g. disks). Carefully placing the data across the storage devices can minimize costly remote accesses and improve performance. We propose the use of simulated annealing to arrive at an effective layout of data on disk. The proposed technique considers the configuration of the system and the cost of data movement. An initial layout globally optimized across all queries, shows speedups of up to 13% for a group of DSS queries and up to 6% for selected OLTP queries. This technique can be re-applied at run-time to further improve performance beyond the initial, globally optimized data layout. This scheme monitors architecture parameters to prevent optimizations of multiple operations to conflict with each other. Such a dynamic reorganization results in speedups of up to 23% for the DSS queries and up to 10% for the OLTP queries.
Juan C. Rubio, Charles Lefurgy, Lizy Kurian John
SBAC-PAD2
2003 On evaluating request-distribution schemes for saving energy in server clusters
abstract
Power-performance optimization is a relatively new problem area particularly in the context of server clusters. Power-aware request distribution is a method of scheduling service requests among servers in a cluster so that energy consumption is minimized, while maintaining a particular level of performance. Energy efficiency is obtained by powering-down some servers when the desired quality of service can be met with fewer servers. We have found that it is critical to take into account the system and workload factors during both the design and the evaluation of such request distribution schemes. We identify the key system and workload factors that impact such policies and their effectiveness in saving energy. We measure a web cluster running an industry-standard commercial web workload to demonstrate that understanding this system-workload context is critical to performing valid evaluations and even for improving the energy-saving schemes.
Karthick Rajamani, Charles Lefurgy
ISPASS2
2002 Critical power slope: understanding the runtime effects of frequency scaling
abstract
Energy efficiency is becoming an increasingly important feature for both mobile and high-performance server systems. Most processors designed today include power management features that provide processor operating points which can be used in power management algorithms. However, existing power management algorithms implicitly assume that lower performance points are more energy efficient than higher performance points. Our empirical observations indicate that for many systems, this assumption is not valid.We introduce a new concept called critical power slope to explain and capture the power-performance characteristics of systems with power management features. We evaluate three systems - a clock throttled Pentium laptop, a frequency scaled PowerPC platform, and a voltage scaled system to demonstrate the benefits of our approach. Our evaluation is based on empirical measurements of the first two systems, and publicly available data for the third. Using critical power slope, we explain why on the Pentium-based system, it is energy efficient to run only at the highest frequency, while on the PowerPC-based system, it is energy efficient to run at the lowest frequency point. We confirm our results by measuring the behavior of a web serving benchmark. Furthermore, we extend the critical power slope concept to understand the benefits of voltage scaling when combined with frequency scaling. We show that in some cases, it may be energy efficient not to reduce voltage below a certain point.
Akihiko Miyoshi, Charles Lefurgy, Eric Van Hensbergen, Ramakrishnan Rajamony, Ragunathan Rajkumar
ICS2
2000 Reducing Code Size with Run-Time Decompression
abstract
Compressed representations of programs can be used to improve the code density in embedded systems. Several hardware decompression architectures have been proposed recently. In this paper, we present a method of decompressing programs using software. It relies on using a software-managed instruction cache under control of the decompressor. This is achieved by employing a simple cache management instruction that allows explicit writing into a cache line. We also consider selective compression (determining which procedures in a program should be compressed) and show that selection based on cache miss profiles can substantially outperform the usual execution time based profiles for some benchmarks.
Charles Lefurgy, Eva Piccininni, Trevor N. Mudge
HPCA1
1999 Evaluation of a High Performance Code Compression Method
abstract
Compressing the instructions of an embedded program is important for cost-sensitive low-power control-oriented embedded computing. A number of compression schemes have been proposed to reduce program size. However, the increased instruction density has an accompanying performance cost because the instructions must be decompressed before execution. In this paper, we investigate the performance penalty of a hardware-managed code compression algorithm recently introduced in IBM's PowerPC 405. This scheme is the first to combine many previously proposed code compression techniques, making it an ideal candidate for study. We find that code compression with appropriate hardware optimizations does not have to incur much performance loss. Furthermore, our studies show this holds for architectures with a wide range of memory configurations and issue widths. Surprisingly, we find that a performance increase over native code is achievable in many situations.
Charles Lefurgy, Eva Piccininni, Trevor N. Mudge
MICRO1
1997 Improving Code Density Using Compression Techniques
abstract
Proposes a method for compressing programs in embedded processors where the instruction memory size dominates the cost. A post-compilation analyzer examines a program and replaces common sequences of instructions with a single instruction codeword. A microprocessor executes the compressed instruction sequences by fetching codewords from the instruction memory, expanding them back to the original sequence of instructions in the decode stage, and issuing them to the execution stages. We apply our technique to the PowerPC, ARM and i386 instruction sets and achieve an average size reduction of 39%, 34% and 26%, respectively, for SPEC CINT95 programs.
Charles Lefurgy, Peter L. Bird, I-Cheng K. Chen, Trevor N. Mudge
MICRO1