EDBT 2026 Demo / reviewers in the wild / expert
Charles Lefurgy
dblp:37/6803 · also Charles R. Lefurgy
· DBLP profile ↗
20ranked-venue papers
4as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Energy-efficient computing · 62% Hardware reliability and fault tolerance · 14% Processor architecture and microarchitecture · 9% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 25 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
power management |
1.4 | 7 | 2019 | Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019 A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015 |
Energy-efficient computing
datacenter power management |
0.5 | 2 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012 |
Energy-efficient computing › datacenter power management
server power capping |
0.4 | 2 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012 |
Processor architecture and microarchitecture › adaptive architecture
timing margin adaptation |
0.4 | 1 | 2019 | Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019 |
Hardware reliability and fault tolerance
timing guardband |
0.3 | 2 | 2015 | Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015 Active management of timing guardband to save energy in POWER7 · MICRO 2011 |
Integrated circuit design › variation-aware design
adaptive guardbanding |
0.2 | 1 | 2015 | Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015 |
Hardware reliability and fault tolerance
processor reliability |
0.2 | 1 | 2015 | Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015 |
Energy-efficient computing › power management › system-level power management
hierarchical power management |
0.1 | 1 | 2012 | SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data Centers · IEEE Trans. Parallel Distributed Syst. 2012 |
Energy-efficient computing › power management › energy-efficient networking
network power management |
0.1 | 1 | 2011 | Power shifting in Thrifty Interconnection Network · HPCA 2011 |
Distributed systems › fault tolerance
high availability |
0.1 | 1 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 |
Embedded and real-time systems › real-time scheduling
multicore scheduling |
0.1 | 1 | 2019 | Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER Server · HPCA 2019 |
Cloud and datacenter computing › resource management
datacenter resource management |
0.1 | 1 | 2007 | Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007 |
Energy-efficient computing › power management
dynamic power management |
0.1 | 1 | 2007 | Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007 |
Energy-efficient computing
voltage and frequency scaling |
0.1 | 1 | 2015 | Adaptive guardband scheduling to improve system-level efficiency of the POWER7+ · MICRO 2015 |
Cloud and datacenter computing › resource management
cloud resource management |
0.0 | 1 | 2012 | Accurate Fine-Grained Processor Power Proxies · MICRO 2012 |
Compilers and program optimization › code size reduction
code compression |
0.0 | 2 | 1999 | Evaluation of a High Performance Code Compression Method · MICRO 1999 Improving Code Density Using Compression Techniques · MICRO 1997 |
Embedded and real-time systems
embedded processor |
0.0 | 2 | 1999 | Evaluation of a High Performance Code Compression Method · MICRO 1999 Improving Code Density Using Compression Techniques · MICRO 1997 |
Hardware reliability and fault tolerance
aging |
0.0 | 1 | 2011 | Active management of timing guardband to save energy in POWER7 · MICRO 2011 |
Hardware reliability and fault tolerance
process variation |
0.0 | 1 | 2011 | Active management of timing guardband to save energy in POWER7 · MICRO 2011 |
Embedded and real-time systems › embedded processor
code compression |
0.0 | 1 | 2000 | Reducing Code Size with Run-Time Decompression · HPCA 2000 |
Memory systems › cache › CPU cache
instruction cache |
0.0 | 1 | 2000 | Reducing Code Size with Run-Time Decompression · HPCA 2000 |
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems |
0.0 | 1 | 2007 | Managing Power Consumption and Performance of Computing Systems Using Reinforcement Learning · NIPS 2007 |
Compilers and program optimization
code size reduction |
0.0 | 1 | 2000 | Reducing Code Size with Run-Time Decompression · HPCA 2000 |
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 1999 | Evaluation of a High Performance Code Compression Method · MICRO 1999 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1997 | Improving Code Density Using Compression Techniques · MICRO 1997 |
Methods — techniques the papers use, named apart from their topics
priority-aware scheduling · 0.4power capping · 0.4critical path monitor · 0.4application scheduling and throttling · 0.4simulation · 0.3power model training · 0.1on-chip power sensing · 0.1control theory · 0.1CPU frequency scaling · 0.1workload trace analysis · 0.1reinforcement learning · 0.1multi-criteria reward · 0.1selective compression · 0.0cache miss profiling · 0.0post-compilation analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server PowerabstractPower management is a key component of modern data center design. Power managers must (1) ensure the costand energy-efficient utilization of the data center infrastructure, (2) maintain availability of the services provided by the center, and (3) address environmental concerns associated with the center's power consumption. While several power management techniques have been proposed and deployed in production data centers, there are still many challenges to comprehensive data center power management. This is particularly true in public cloud environments, where different jobs have different priority levels, and where high availability is critical. One example of the challenges facing public cloud data centers involves power capping. As power delivery must be highly reliable and tolerate wide variation in the load drawn by the data center components, the power infrastructure (e.g., power supplies, circuit breakers, UPS) has high redundancy and overprovisioning. During normal operation (i.e., typical server power demands, and no failures in the center), the power infrastructure is significantly underutilized. Power capping is a common solution to reduce this underutilization, by allowing more servers to be added safely (i.e., without power shortfalls) to the existing power infrastructure, and throttling power consumption in the infrequent cases where the demanded power exceeds the provisioned power capacity to avoid shortfalls. However, state-of-the-art power capping solutions are (1) not directly applicable to the redundant power infrastructure used in highly-available data centers; and (2) oblivious to differing workload priorities across the entire center when power consumption needs to be throttled, which can unnecessarily slow down high-priority work. To address this need, we develop CapMaestro, a new power management architecture with three key features for public cloud data centers. First, CapMaestro is designed to work with multiple power feeds (i.e., sources), and exploits server-level power capping to independently cap the load on each feed of a server. Second, CapMaestro uses a scalable, global priority-aware power capping approach, which accounts for power capacity at each level of the power distribution hierarchy. It exploits the underutilization of commonly-employed redundant power infrastructure at each level of the hierarchy to safely accommodate a much greater number of servers. Third, CapMaestro exploits stranded power (i.e., power budgets that are not utilized) in redundant power infrastructure to boost the performance of workloads in the data center. We add CapMaestro to a real cloud data center control plane, and demonstrate the effectiveness of all three key features. Using a large-scale data center simulation, we demonstrate that CapMaestro significantly and safely increases the number of servers for existing infrastructure. We also call out other key technical challenges the industry faces in data center power management. Yang Li 0183, Charles Lefurgy, Karthick Rajamani, Malcolm Allen-Ware, Guillermo J. Silva, Daniel D. Heimsoth, Saugata Ghose, Onur Mutlu |
HPCA | 2 |
| 2019 | Fine-Tuning the Active Timing Margin (ATM) Control Loop for Maximizing Multi-core Efficiency on an IBM POWER ServerabstractActive Timing Margin (ATM) is a technology that improves processor efficiency by reducing the pipeline timing margin with a control loop that adjusts voltage and frequency based on real-time chip environment monitoring. Although ATM has already been shown to yield substantial performance benefits, its full potential has yet to be unlocked. In this paper, we investigate how to maximize ATM's efficiency gain with a new means of exposing the inter-core speed variation: finetuning the ATM control loop. We conduct our analysis and evaluation on a production-grade POWER7+ system. On the POWER7+ server platform, we fine-tune the ATM control loop by programming its Critical Path Monitors, a key component of its ATM design that measures the cores' timing margins. With a robust stress-test procedure, we expose over 200 MHz of inherent inter-core speed differential by fine-tuning the percore ATM control loop. Exploiting this differential, we manage to double the ATM frequency gain over the static timing margin; this is not possible using conventional means, i.e. by setting fixedpoints for each core, because the corelevelmust account for chip-wide worst-case voltage variation. To manage the significant performance heterogeneity of fine-tuned systems, we propose application scheduling and throttling to manage the chip's process and voltage variation. Our proposal improves application performance by more than 10% over the static margin, almost doubling the 6% improvement of the default, unmanaged ATM system. Our technique is general enough that it can be adopted by any system that employs an active timing margin control loop. Yazhou Zu, Daniel Richins, Charles Lefurgy, Vijay Janapa Reddi |
HPCA | 3 |
| 2015 | FreqLeak: A frequency step based method for efficient leakage power characterization in a systemabstractAccurate estimation of leakage power at runtime requires post-silicon power measurements across a wide range of temperature and voltage conditions. Testing individual chips, especially at high-temperature corner conditions, is expensive in cost and time. We examine this problem in an industrial context and introduce FreqLeak, a frequency step based method for inexpensive and efficient leakage power characterization in a system. It enables a more thorough characterization than can be accomplished on a wafer prober alone due to time and equipment costs. Experimental evaluation on IBM POWER8 based systems demonstrates the efficiency of the proposed method, within an error of 5%. Further, we discuss the application of FreqLeak in system level power management. Anand Haridass, Charles Lefurgy, Sreekanth Pai, Spandana Rachamalla, Francesco Campisano |
ISLPED | 3 |
| 2015 | Adaptive guardband scheduling to improve system-level efficiency of the POWER7+abstractThe traditional guardbanding approach to ensure processor reliability is becoming obsolete because it always over-provisions voltage and wastes a lot of energy. As a next-generation alternative, adaptive guardbanding dynamically adjusts chip clock frequency and voltage based on timing margin measured at runtime. With adaptive guardbanding, voltage guardband is only provided when needed, thereby promising significant energy efficiency improvement. Yazhou Zu, Charles Lefurgy, Jingwen Leng, Matthew Halpern, Michael S. Floyd, Vijay Janapa Reddi |
MICRO | 2 |
| 2012 | Accurate Fine-Grained Processor Power ProxiesabstractThere are not yet practical and accurate ways to directly measure core power in a microprocessor. This limits the granularity of measurement and control for computer power management. We overcome this limitation by presenting an accurate runtime per-core power proxy which closely estimates true core power. This enables new fine-grained microprocessor power management techniques at the core level. For example, cloud environments could manage and bill virtual machines for energy consumption associated with the core. The power model underlying our power proxy also enables energy-efficiency controllers to perform what-if analysis, instead of merely reacting to current conditions. We develop and validate a methodology for accurate power proxy training at both chip and core levels. Our implementation of power proxies uses on-chip logic in a high-performance multi-core processor and associated platform firmware. The power proxies account for full voltage and frequency ranges, as well as chip-to-chip process variations. For fixed clock frequency operation, a mean unsigned error of 1.8% for fine-grained 32ms samples across all workloads was achieved. For an interval of an entire workload, we achieve an average error of-0.2%. Similar results were achieved for voltage-scaling scenarios, too. We also present two sample applications of the power proxy: (1) per-core power billing for cloud computing services, and (2) simultaneous runtime energy saving comparisons among different power management policies without running each policy separately. Wei Huang 0004, Charles Lefurgy, William Kuk, Alper Buyuktosunoglu, Michael S. Floyd, Karthick Rajamani, Malcolm Allen-Ware, Bishop Brock |
MICRO | 2 |
| 2012 | SHIP: A Scalable Hierarchical Power Control Architecture for Large-Scale Data CentersabstractIn today's data centers, precisely controlling server power consumption is an essential way to avoid system failures caused by power capacity overload or overheating due to increasingly high server density. While various power control strategies have been recently proposed, existing solutions are not scalable to control the power consumption of an entire large-scale data center, because these solutions are designed only for a single server or a rack enclosure. In a modern data center, however, power control needs to be enforced at three levels: rack enclosure, power distribution unit, and the entire data center, due to the physical and contractual power limits at each level. This paper presents SHIP, a highly scalable hierarchical power control architecture for large-scale data centers. SHIP is designed based on well-established control theory for analytical assurance of control accuracy and system stability. Empirical results on a physical testbed show that our control solution can provide precise power control, as well as power differentiations for optimized system performance and desired server priorities. In addition, our extensive simulation results based on a real trace file demonstrate the efficacy of our control solution in large-scale data centers composed of 5,415 servers. Ming Chen 0002, Charles Lefurgy, Tom W. Keller |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | Power shifting in Thrifty Interconnection NetworkabstractThis paper presents two complementary techniques to manage the power consumption of large-scale systems with a packet-switched interconnection network. First, we propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically with little or no overhead by using inherent system events to timely trigger link activation or de-activation. Second, we propose Network Power Shifting (NPS) that dynamically shifts the power budget between the compute nodes and their corresponding network components. TIN activates and trains the links in the interconnection network, just-in-time before the network communication is about to happen, and thriftily puts them into a low-power mode when communication is finished, hence reducing unnecessary network power consumption. Furthermore, the compute nodes can absorb the extra power budget shifted from its attached network components and increase their processor frequency for higher performance with NPS. Our simulation results on a set of real-world workload traces show that TIN can achieve on average 60% network power reduction, with the support of only one low-power mode. When NPS is enabled, the two together can achieve 12% application performance improvement and 13% overall system energy reduction. Further performance improvement is possible if the compute nodes can speed up more and fully utilize the extra power budget reinvested from the thrifty network with more aggressive cooling support. Jian Li 0059, Wei Huang 0004, Charles Lefurgy, Lixin Zhang 0002, Wolfgang E. Denzel, Richard R. Treumann, Kun Wang 0005 |
HPCA | 3 |
| 2011 | Active management of timing guardband to save energy in POWER7abstractMicroprocessor voltage levels include substantial margin to deal with process variation, system power supply variation, workload induced thermal and voltage variation, aging, random uncertainty, and test inaccuracy. This margin allows the microprocessor to operate correctly during worst-case conditions, but during typical conditions it is larger than necessary and wastes energy. We present a mechanism that reduces excess voltage margin by (1) introducing a critical path monitor (CPM) circuit that measures available timing margin in real-time, (2) coupling the CPM output to the clock generation circuit to adjust clock frequency within cycles in response to excess or inadequate timing margin, and (3) adjusting the processor voltage level periodically in firmware to achieve a specified average clock frequency target. We implemented this mechanism in a prototype IBM POWER7 server. During better-than-worst case conditions our guardband management mechanism reduces the average voltage setting 137-152 mV below nominal, resulting in average processor power reduction of 24% with no performance loss while running industry-standard benchmarks. Charles Lefurgy, Alan J. Drake, Michael S. Floyd, Malcolm Allen-Ware, Bishop Brock, José A. Tierno, John B. Carter |
MICRO | 1 |
| 2010 | Adaptive energy management features of the POWER7TM processor
Michael S. Floyd, Bishop Brock, Malcolm Allen-Ware, Karthick Rajamani, Alan J. Drake, Charles Lefurgy, Lorena Pesantez |
Hot Chips Symposium | 6 |
| 2009 | SHIP: Scalable Hierarchical Power Control for Large-Scale Data CentersabstractIn today's data centers, precisely controlling server power consumption is an essential way to avoid system failures caused by power capacity overload or overheating due to increasingly high server density. While various power control strategies have been recently proposed, existing solutions are not scalable to control the power consumption of an entire large-scale data center, because these solutions are designed only for a single server or a rack enclosure. In a modern data center, however, power control needs to be enforced at three levels: rack enclosure, power distribution unit, and the entire data center, due to the physical and contractual power limits at each level. This paper presents SHIP, a highly scalable hierarchical power control architecture for large-scale data centers. SHIP is designed based on well-established control theory for analytical assurance of control accuracy and system stability. Empirical results on a physical testbed show that our control solution can provide precise power control, as well as power differentiations for optimized system performance. In addition, our extensive simulation results based on a real trace file demonstrate the efficacy of our control solution in large-scale data centers composed of thousands of servers. Ming Chen 0002, Charles Lefurgy, Tom W. Keller |
PACT | 3 |
| 2009 | Thrifty interconnection network for HPC systemsabstractWe propose Thrifty Interconnection Network (TIN), where the network links are activated and de-activated dynamically to save power with little or no overhead by using inherent system events to overlap the link activation or de-activation time. Our simulation results on a set of real world HPC workload traces show on average 35% network power reduction. Jian Li 0059, Lixin Zhang 0002, Charles Lefurgy, Richard R. Treumann, Wolfgang E. Denzel |
ICS | 3 |
| 2008 | Power management solutions for computer systems and datacentersabstractThe growing power and cooling requirements of high-density computing systems pose significant challenges for the design and operation of computers and their facilities. The rising operating expenses for datacenters demand the implementation of energy-efficient technologies and the best power management solutions. This tutorial addresses power management and cooling solutions from the individual computer system level to the datacenter. The audience will learn about the fundamental nature of the problems, approaches to developing solutions, available commercial solutions, and current research directions. Karthick Rajamani, Charles Lefurgy, Soraya Ghiasi, Juan C. Rubio, Heather Hanson, Tom W. Keller |
ISLPED | 2 |
| 2007 | Managing Power Consumption and Performance of Computing Systems Using Reinforcement LearningabstractElectrical power management in large-scale IT systems such as commercial data- centers is an application area of rapidly growing interest from both an economic and ecological perspective, with billions of dollars and millions of metric tons of CO2 emissions at stake annually. Businesses want to save power without sac- rificing performance. This paper presents a reinforcement learning approach to simultaneous online management of both performance and power consumption. We apply RL in a realistic laboratory testbed using a Blade cluster and dynam- ically varying HTTP workload running on a commercial web applications mid- dleware platform. We embed a CPU frequency controller in the Blade servers’ firmware, and we train policies for this controller using a multi-criteria reward signal depending on both application performance and CPU power consumption. Our testbed scenario posed a number of challenges to successful use of RL, in- cluding multiple disparate reward functions, limited decision sampling rates, and pathologies arising when using multiple sensor readings as state variables. We describe innovative practical solutions to these challenges, and demonstrate clear performance improvements over both hand-designed policies as well as obvious “cookbook” RL implementations. Gerald Tesauro, Rajarshi Das, Hoi Y. Chan, Jeffrey O. Kephart, David W. Levine, Freeman L. Rawson III, Charles Lefurgy |
NIPS | 7 |
| 2005 | Improving energy efficiency by making DRAM less randomly accessedabstractExisting techniques manage power for the main memory by passively monitoring the memory traffic, and based on which, predict when to power down and into which low-power state to transition. However, passively monitoring the memory traffic can be far from being effective as idle periods between consecutive memory accesses are often too short for existing power-management techniques to take full advantage of the deeper power-saving state implemented in modern DRAM architectures. In this paper, we propose a new technique that will actively reshape the memory traffic to coalesce short idle periods --- which were previously unusable for power management --- into longer ones, thus enabling existing techniques to effectively exploit idleness in the memory Hai Huang 0002, Kang G. Shin, Charles Lefurgy, Tom W. Keller |
ISLPED | 3 |
| 2004 | Improving Server Performance on Transaction Processing Workloads by Enhanced Data PlacementabstractModern servers access large volumes of data while running commercial workloads. The data is typically spread among several storage devices (e.g. disks). Carefully placing the data across the storage devices can minimize costly remote accesses and improve performance. We propose the use of simulated annealing to arrive at an effective layout of data on disk. The proposed technique considers the configuration of the system and the cost of data movement. An initial layout globally optimized across all queries, shows speedups of up to 13% for a group of DSS queries and up to 6% for selected OLTP queries. This technique can be re-applied at run-time to further improve performance beyond the initial, globally optimized data layout. This scheme monitors architecture parameters to prevent optimizations of multiple operations to conflict with each other. Such a dynamic reorganization results in speedups of up to 23% for the DSS queries and up to 10% for the OLTP queries. Juan C. Rubio, Charles Lefurgy, Lizy Kurian John |
SBAC-PAD | 2 |
| 2003 | On evaluating request-distribution schemes for saving energy in server clustersabstractPower-performance optimization is a relatively new problem area particularly in the context of server clusters. Power-aware request distribution is a method of scheduling service requests among servers in a cluster so that energy consumption is minimized, while maintaining a particular level of performance. Energy efficiency is obtained by powering-down some servers when the desired quality of service can be met with fewer servers. We have found that it is critical to take into account the system and workload factors during both the design and the evaluation of such request distribution schemes. We identify the key system and workload factors that impact such policies and their effectiveness in saving energy. We measure a web cluster running an industry-standard commercial web workload to demonstrate that understanding this system-workload context is critical to performing valid evaluations and even for improving the energy-saving schemes. Karthick Rajamani, Charles Lefurgy |
ISPASS | 2 |
| 2002 | Critical power slope: understanding the runtime effects of frequency scalingabstractEnergy efficiency is becoming an increasingly important feature for both mobile and high-performance server systems. Most processors designed today include power management features that provide processor operating points which can be used in power management algorithms. However, existing power management algorithms implicitly assume that lower performance points are more energy efficient than higher performance points. Our empirical observations indicate that for many systems, this assumption is not valid.We introduce a new concept called critical power slope to explain and capture the power-performance characteristics of systems with power management features. We evaluate three systems - a clock throttled Pentium laptop, a frequency scaled PowerPC platform, and a voltage scaled system to demonstrate the benefits of our approach. Our evaluation is based on empirical measurements of the first two systems, and publicly available data for the third. Using critical power slope, we explain why on the Pentium-based system, it is energy efficient to run only at the highest frequency, while on the PowerPC-based system, it is energy efficient to run at the lowest frequency point. We confirm our results by measuring the behavior of a web serving benchmark. Furthermore, we extend the critical power slope concept to understand the benefits of voltage scaling when combined with frequency scaling. We show that in some cases, it may be energy efficient not to reduce voltage below a certain point. Akihiko Miyoshi, Charles Lefurgy, Eric Van Hensbergen, Ramakrishnan Rajamony, Ragunathan Rajkumar |
ICS | 2 |
| 2000 | Reducing Code Size with Run-Time DecompressionabstractCompressed representations of programs can be used to improve the code density in embedded systems. Several hardware decompression architectures have been proposed recently. In this paper, we present a method of decompressing programs using software. It relies on using a software-managed instruction cache under control of the decompressor. This is achieved by employing a simple cache management instruction that allows explicit writing into a cache line. We also consider selective compression (determining which procedures in a program should be compressed) and show that selection based on cache miss profiles can substantially outperform the usual execution time based profiles for some benchmarks. Charles Lefurgy, Eva Piccininni, Trevor N. Mudge |
HPCA | 1 |
| 1999 | Evaluation of a High Performance Code Compression MethodabstractCompressing the instructions of an embedded program is important for cost-sensitive low-power control-oriented embedded computing. A number of compression schemes have been proposed to reduce program size. However, the increased instruction density has an accompanying performance cost because the instructions must be decompressed before execution. In this paper, we investigate the performance penalty of a hardware-managed code compression algorithm recently introduced in IBM's PowerPC 405. This scheme is the first to combine many previously proposed code compression techniques, making it an ideal candidate for study. We find that code compression with appropriate hardware optimizations does not have to incur much performance loss. Furthermore, our studies show this holds for architectures with a wide range of memory configurations and issue widths. Surprisingly, we find that a performance increase over native code is achievable in many situations. Charles Lefurgy, Eva Piccininni, Trevor N. Mudge |
MICRO | 1 |
| 1997 | Improving Code Density Using Compression TechniquesabstractProposes a method for compressing programs in embedded processors where the instruction memory size dominates the cost. A post-compilation analyzer examines a program and replaces common sequences of instructions with a single instruction codeword. A microprocessor executes the compressed instruction sequences by fetching codewords from the instruction memory, expanding them back to the original sequence of instructions in the decode stage, and issuing them to the execution stages. We apply our technique to the PowerPC, ARM and i386 instruction sets and achieve an average size reduction of 39%, 34% and 26%, respectively, for SPEC CINT95 programs. Charles Lefurgy, Peter L. Bird, I-Cheng K. Chen, Trevor N. Mudge |
MICRO | 1 |