Haris Lekatsas

dblp:75/5187 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 7 first-authorDatabases, data management, data science and information retrieval · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Embedded and real-time systems · 53% Processor architecture and microarchitecture · 16% Integrated circuit design · 13%
Software engineering, system software, and programming languages
3 papers
Operating systems · 76% Compilers and program optimization · 24%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Embedded and real-time systems › embedded processor
code compression
0.252003
CoCo: a hardware/software platform for rapid prototyping of code compression technologies · DAC 2003
Design of an one-cycle decompression hardware for performance increase in embedded systems · DAC 2002
A code decompression architecture for VLIW processors · MICRO 2001
Operating systems › resource management › memory management
memory compression
0.112006
High-performance operating system controlled memory compression · DAC 2006
Operating systems › resource management
memory management
0.112006
High-performance operating system controlled memory compression · DAC 2006
Embedded and real-time systems › embedded software › embedded operating systems
embedded memory management
0.112006
High-performance operating system controlled memory compression · DAC 2006
Integrated circuit design › parasitic capacitance
coupling capacitance
0.012001
A2BC: Adaptive Address Bus Coding for Low Power Deep Sub-Micron Designs · DAC 2001
Integrated circuit design › VLSI design
deep submicron design
0.012001
A2BC: Adaptive Address Bus Coding for Low Power Deep Sub-Micron Designs · DAC 2001
Embedded and real-time systems
embedded system design
0.012001
A code decompression architecture for VLIW processors · MICRO 2001
Processor architecture and microarchitecture
instruction set architecture
0.012001
A code decompression architecture for VLIW processors · MICRO 2001
Energy-efficient computing › low-power design
low-power interconnect
0.012001
A2BC: Adaptive Address Bus Coding for Low Power Deep Sub-Micron Designs · DAC 2001
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor
0.012001
A code decompression architecture for VLIW processors · MICRO 2001
Energy-efficient computing › low-power design
low-power embedded design
0.012000
Code compression for low power embedded system design · DAC 2000
Reconfigurable computing and FPGAs
FPGA prototyping
0.012003
CoCo: a hardware/software platform for rapid prototyping of code compression technologies · DAC 2003
Memory systems › cache › CPU cache
instruction cache
0.012000
Code compression for low power embedded system design · DAC 2000
Compilers and program optimization
code size reduction
0.011999
SAMC: a code compression algorithm for embedded processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999

Methods — techniques the papers use, named apart from their topics

adaptive memory management · 0.1parallel decompression · 0.1online compression · 0.1on-line compression · 0.1flexible instruction format compression · 0.1table-based decoding · 0.0markov model · 0.0arithmetic coding · 0.0physical bus modeling · 0.0bus encoding · 0.0
YearPublicationVenuePosition
2010 Online memory compression for embedded systems
abstract
Memory is a scarce resource during embedded system design. Increasing memory often increases packaging costs, cooling costs, size, and power consumption. This article presents CRAMES, a novel and efficient software-based RAM compression technique for embedded systems. The goal of CRAMES is to dramatically increase effective memory capacity without hardware or application design changes, while maintaining high performance and low energy consumption. To achieve this goal, CRAMES takes advantage of an operating system's virtual memory infrastructure by storing swapped-out pages in compressed format. It dynamically adjusts the size of the compressed RAM area, protecting applications capable of running without it from performance or energy consumption penalties. In addition to compressing working data sets, CRAMES also enables efficient in-RAM filesystem compression, thereby further increasing RAM capacity. CRAMES was implemented as a loadable module for the Linux kernel and evaluated on a battery-powered embedded system. Experimental results indicate that CRAMES is capable of doubling the amount of RAM available to applications running on the original system hardware. Execution time and energy consumption for a broad range of examples are rarely affected. When physical RAM is reduced to 62.5% of its original quantity, CRAMES enables the target embedded system to support the same applications with reasonable performance and energy consumption penalties (on average 9.5% and 10.5%), while without CRAMES those applications either may not execute or suffer from extreme performance degradation or instability. In addition to presenting a novel framework for dynamic data memory compression and in-RAM filesystem compression in embedded systems, this work identifies the software-based compression algorithms that are most appropriate for use in low-power embedded systems.
Lei Yang 0017, Robert P. Dick, Haris Lekatsas, Srimat T. Chakradhar
ACM Trans. Embed. Comput. Syst.3
2010 High-performance operating system controlled online memory compression
abstract
Online memory compression is a technology that increases the amount of memory available to applications by dynamically compressing and decompressing their working datasets on demand. It has proven extremely useful in embedded systems with tight physical RAM constraints. The technology can be used to increase functionality, reduce size, and reduce cost, without modifying applications or hardware. This article presents a new software-based online memory compression algorithm for embedded systems. In comparison with the best algorithms used in online memory compression, our new algorithm has a competitive compression ratio but is twice as fast. In addition, we describe several practical problems encountered in developing an online memory compression infrastructure and present solutions. We present a method of adaptively managing the uncompressed and compressed memory regions during application execution. This memory management scheme adapts to the predicted memory requirements of applications. It permits efficient compression for a wide range of applications. We have evaluated our techniques on a portable embedded device and have found that the memory available to applications can be increased by 2.5× with negligible performance and power consumption penalties, and with no changes to hardware or applications. Our techniques allow existing applications to execute with less physical memory. They also allow applications with larger working datasets to execute on unchanged embedded system hardware, thereby increasing functionality.
Lei Yang 0017, Robert P. Dick, Haris Lekatsas, Srimat T. Chakradhar
ACM Trans. Embed. Comput. Syst.3
2010 C-Pack: A High-Performance Microprocessor Cache Compression Algorithm
abstract
Microprocessor designers have been torn between tight constraints on the amount of on-chip cache memory and the high latency of off-chip memory, such as dynamic random access memory. Accessing off-chip memory generally takes an order of magnitude more time than accessing on-chip cache, and two orders of magnitude more time than executing an instruction. Computer systems and microarchitecture researchers have proposed using hardware data compression units within the memory hierarchies of microprocessors in order to improve performance, energy efficiency, and functionality. However, most past work, and all work on cache compression, has made unsubstantiated assumptions about the performance, power consumption, and area overheads of the proposed compression algorithms and hardware. It is not possible to determine whether compression at levels of the memory hierarchy closest to the processor is beneficial without understanding its costs. Furthermore, as we show in this paper, raw compression ratio is not always the most important metric. In this work, we present a lossless compression algorithm that has been designed for fast on-line data compression, and cache compression in particular. The algorithm has a number of novel features tailored for this application, including combining pairs of compressed lines into one cache line and allowing parallel compression of multiple words while using a single dictionary and without degradation in compression ratio. We reduced the proposed algorithm to a register transfer level hardware design, permitting performance, power consumption, and area estimation. Experiments comparing our work to previous work are described.
Xi Chen 0068, Lei Yang 0017, Robert P. Dick, Haris Lekatsas
IEEE Trans. Very Large Scale Integr. Syst.5
2008 Adaptive Filesystem Compression for Embedded Systems
abstract
Embedded system secondary storage size is often constrained, yet storage demands are growing as a result of increasing application complexity and storage of personal data and multimedia flies. Filesystem compression offers a solution. This paper formalizes the problem of automatic filesystem compression using multiple compression algorithms. The average latency of on-line file accesses is optimized under a constraint on filesystem capacity. Our solution is based on predictive control. Predicted latency implications are used to solve the file compression state selection problem using a multiple choice knapsack problem formulation. This approach is evaluated on filesystem traces and compared with other efficient heuristics. Our approach results in 34.1% reduction in file access latency compared to a straight-forward heuristic that decompresses frequently-accessed files and compresses least recently used files with more aggressive compression algorithms. It reduces file access latency by 67.7% compared to uniformly compressing files to the shallowest level required to meet storage capacity constraints.
Lan S. Bai, Haris Lekatsas, Robert P. Dick
DATE2
2008 Design and Implementation of a High-Performance Microprocessor Cache Compression Algorithm
abstract
Abstract Researchers have proposed using hardware data compression units within the memory hierarchies of microprocessors in order to improve performance, energy efficiency, and functionality. However, most past work, and in particular work on cache compression, has made unsubstantiated assumptions about the performance, power consumption, and area overheads of the required compression hardware. We present a lossless compression algorithm that has been designed for on-line memory hierarchy compression, and cache compression in particular. We reduced our algorithm to a register transfer level hardware implementation, permitting performance, power consumption, and area estimation. The results of experiments comparing our work to previous work are presented.
Xi Chen 0068, Lei Yang 0017, Haris Lekatsas, Robert P. Dick
DCC3
2007 Code Decompression Unit Design for VLIW Embedded Processors
abstract
Code size "bloating" in embedded very long instruction word (VLIW) processors is a major concern for embedded systems since memory is one of the most restricted resources. In this paper, we describe a code compression algorithm based on arithmetic coding, discuss how to design decompression architecture, and illustrate the tradeoffs between compression ratio and decompression overhead, by using different probability models. Experimental results for a VLIW embedded processor TMS320C6x show that compression ratios between 67% and 80% can be achieved, depending on the probability models used. A precache decompression unit design is implemented in TSMC 0.25 mum and a test chip is fabricated.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
IEEE Trans. Very Large Scale Integr. Syst.3
2006 High-performance operating system controlled memory compression
abstract
This article describes a new software-based on-line memory compression algorithm for embedded systems and presents a method of adaptively managing the uncompressed and compressed memory regions during application execution. The primary goal of this work is to save memory in disk-less embedded systems, resulting in greater functionality, smaller size, and lower overall cost, without modifying applications or hardware. In comparison with algorithms that are commonly used in on-line memory compression, our new algorithm has a comparable compression ratio but is twice as fast. The adaptive memory management scheme effectively responds to the predicted needs of applications and prevents on-line memory compression deadlock, permitting reliable and efficient compression for a wide range of applications. We have evaluated our technique on an embedded portable device and have found that the memory available to applications can be increased by 150%, allowing the execution of applications with larger working data sets, or allowing existing applications to run with less physical memory.
Lei Yang 0017, Haris Lekatsas, Robert P. Dick
DAC2
2006 Code Compression for Embedded VLIW Processors Using Variable-to-Fixed Coding
abstract
In embedded system design, memory is one of the most restricted resources, posing serious constraints on program size. Code compression has been used as a solution to reduce the code size for embedded systems. Lossless data compression techniques are used to compress instructions, which are then decompressed on-the-fly during execution. Previous work used fixed-to-variable coding algorithms that translate fixed-length bit sequences into variable-length bit sequences. In this paper, we present a class of code compression techniques called variable-to-fixed code compression (V2FCC), which uses variable-to-fixed coding schemes based on either Tunstall coding or arithmetic coding. Though the techniques are suitable for both reduced instruction set computer (RISC) and very long instruction word (VLIW) architectures, they favor VLIW architectures which require a high-bandwidth instruction prefetch mechanism to supply multiple operations per cycle, and fast decompression is critical to overcome the communication bottleneck between memory and CPU. Experimental results for a VLIW embedded processor TMS320C6x show that the compression ratios using memoryless V2FCC and Markov V2FCC are around 82.5% and 70%, respectively. Decompression unit designs for memoryless V2FCC and Markov V2FCC are implemented in TSMC 0.25-/spl mu/m technology.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Approximate arithmetic coding for bus transition reduction in low power designs
abstract
We present a method for reducing the power consumption of compressed-code systems by selectively inverting bits that are transmitted on the bus. By incorporating bus inversion into code compression/decompression, we reduce power consumption with no cost in hardware or power relative to code compression without inversion. Inverting has to be done carefully to ensure that the codes can still be decoded. As an additional challenge, compression will generally increase bit-toggling as it removes redundancies from the code transmitted. Therefore, we need to find the right balance between compression ratio and bit-toggling reduction. This paper presents a suitable algorithm that will combine approximate compression techniques with bit-toggling reduction and will explore the various tradeoffs. We take advantage of the approximations introduced to modify codes and reduce bit-toggling, while maintaining compression performance and decoding speed. An interesting result that is derived from our work is that high compression ratios do not necessarily result in the lowest power consumption. By using our method, bus-related power consumption has been reduced by as much as 35% compared to a system with no compression, and as much as 14% compared to a compressed-code system. Bit-toggling reduction does not impose any additional hardware costs other than the decompression engine. We also present a detailed analysis on how bus widths affect bit-toggling when transmitting compressed code, and we show experimental results on ARM, MIPS, and SPARC code. We finally compare our work with Bus Invert and show results that are superior except for the random data case where Bus Invert performs better.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.1
2004 Guest editorial: Special issue on embedded systems and security
abstract
No abstract available.
Dimitrios Serpanos, Haris Lekatsas
ACM Trans. Embed. Comput. Syst.2
2003 Multi-parametric improvements for embedded systems using code-placement and address bus coding
abstract
Code placement techniques for instruction code have shown to increase an SOCs performance mostly due to the increased cache hit ratios and as such those techniques can be a major optimization strategy for embedded systems. Little has been investigated on the interdependencies between code placement techniques and interconnect traffic (e.g. bus traffic) and optimization techniques combining both. In this paper we show as the first approach of its kind that a carefully designed known code placement strategy combined and adapted to a known interconnect encoding scheme does not only lead to a performance increase but it does also lead to a significant reduction of interconnect-related energy consumption. This becomes especially interesting since future SOC bus systems (or more general: "networks on a chip") are predicted to be a dominant energy consumer of an SOC. We show that a high-level optimization strategy like code placement and a lower-level optimization strategy like interconnect encoding are NOT orthogonal. Specifically, we report cache miss reduction ratios of 32% in average combined with bus related energy savings of 50.4% in average (with a maximum of up to 95.7%) by means of our combined optimization strategy. The results have been verified by means of diverse real-world SOC applications.
Sri Parameswaran, Jörg Henkel, Haris Lekatsas
ASP-DAC3
2003 CoCo: a hardware/software platform for rapid prototyping of code compression technologies
abstract
In recent years instruction code compression/decompression technologies have emerged as an efficient way to a) reduce the memory usage of an embedded system, b) to improve performance through effectively higher bandwidths and/or to c) reduce the overall power consumption of a system processing compressed code. We have presented efficient code compression/decompression techniques and architectures in the past. For the commercialization phase, we designed a novel hardware/software code compression/decompression platform (CoCo). It consists of a software platform that prepares, optimizes, compresses and compiles instruction code and a generic, parameterizable FPGA-based hardware architecture in form of a hardware platform that allows to rapidly evaluate prototypes of diverse compression/decompression technologies. We show the flexibility of CoCo, its ability to achieve code compression ratios (parameterizable) of up to 50% with a slight system performance gain and its ability to apply compression on real-world compiled code without any limitations where others have made implicit software-restrictive assumptions.
Haris Lekatsas, Jörg Henkel, Srimat T. Chakradhar, Venkata Jakkula, Murugan Sankaradass
DAC1
2003 Enhancing Signal Integrity through a Low-Overhead Encoding Scheme on Address Buses
abstract
Signal integrity is and will continue to be a major concern in deep sub-micron VLSI designs where the proximity of signal carrying lines leads to crosstalk, unpredictable signal delays and other parasitic side effects. Our scheme uses bus encoding that guarantees that at any time any two signal carrying lines will be separated by at least one grounded line and thus providing a high degree of signal integrity. This comes at a small overhead of only one additional bus line (the closest related work needs 14 additional lines for a 32-bit bus) and a small average performance decrease of 0.36%. By means of a large set of real-world applications, we compare our scheme to other state-of-the-art approaches and present comparisons in terms of degree of integrity, overhead (e.g. additional lines required) and a possible performance decrease.
Tiehan Lv, Jörg Henkel, Haris Lekatsas, Marilyn Wolf
DATE3
2003 Profile-Driven Selective Code Compression
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
DATE3
2003 Code Compression Using Variable-to-fixed Coding Based on Arithmetic Coding
abstract
Embedded computing systems are space and cost sensitive. Memory is one of the most restricted resources that post serious constraints on program size. Code compression, which is a special case of data compression where the input source is in machine instructions, has been proposed as a solution to this problem. Previous work in code compression has focused on either fixed-to-variable coding or dictionary-based algorithms. Code compression schemes that use variable-to-fixed (V2F) length coding were proposed, based on arithmetic coding. Experiments have shown that the compression ratio, using memoryless V2F coding for the TMS320C6x processor, have an average of 82.5% and decompression can be parallelized. A Markov-based V2F coding based on arithmetic coding has achieved an average compression ratio of 72% for TMS320C6x while decompression cannot be parallelized. Furthermore, the given experiments have shown that arithmetic coding based V2F coding has similar compression performance to Tunstall coding. Finally, a power reduction scheme for the instruction bus using the V2F coding scheme was presented.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
DCC3
2003 A dictionary-based en/decoding scheme for low-power data buses
abstract
As bus lengths on multihundred-million transistor systems-on-a-chip (SoC) grow, and as interwire capacitances of sub-0.10 /spl mu/m technologies advance, the resulting high-switching capacitances of buses (and interconnects in general) have a nonnegligible impact on the power consumption of a whole SoC. This trend has been recognized and recently addressed by various research groups. We address this problem by introducing our bus encoding technique, adaptive dictionary-encoding scheme "ADES" that minimizes the power consumption of data buses through a dictionary-based encoding technique. Based on exploration of data properties on buses, our technique saves on average more than 25% of bus energy compared to the nonencoded cases using a large set of real-world applications for both address and data buses. Furthermore, we compare our technique to the best-known data bus encoding techniques to date and we find that it exceeds all of them in terms of energy savings for the same set of applications.
Tiehan Lv, Jörg Henkel, Haris Lekatsas, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Design of an one-cycle decompression hardware for performance increase in embedded systems
abstract
Code compression is known as an effective technique to reduce instruction memory size on an embedded system. However, code compression can also be very effective in increasing processor-to-memory bandwidth and hence provide increased system performance. In this paper we describe our design and design methodology of the first running prototype of a one-cycle code decompression unit that decompresses compressed instructions on-the-fly. We describe in detail the architecture that enables decompression of multiple instructions in one cycle and we present the design methodologies and tools used. The stand-alone decompression unit does not require any modifications on the processor core. We observed up to 63% performance increase with 25% in average over a wide variety of applications running on the hardware prototype under various system configurations.
Haris Lekatsas, Jörg Henkel, Venkata Jakkula
DAC1
2002 An Adaptive Dictionary Encoding Scheme for SOC Data Buses
abstract
As bus lengths on multi-hundred-million transistor SOCs (systems-on-a-chip) grow and as inter-wire capacitances of sub-0.1 /spl mu/m technologies increase, the resulting high switching capacitances of buses (and interconnects in general) have a non-negligible impact on the power consumption of a whole SOC. In this paper, we address this problem by introducing our bus encoding technique 'ADES' that minimizes the power consumption of data buses through a dictionary-based encoding technique. We show that our technique saves between 18% and 40% of bus energy compared to the non-encoded cases using a large set of (freely-accessible) real-world applications. Furthermore, we compare our technique to the best-known data bus encoding techniques to date and show that it exceeds all of them in energy savings for the same set of applications. The additional hardware effort for our bus en/decoder is thereby very small.
Tiehan Lv, Marilyn Wolf, Jörg Henkel, Haris Lekatsas
DATE4
2001 A2BC: Adaptive Address Bus Coding for Low Power Deep Sub-Micron Designs
abstract
Due to larger buses (length, width) and deep sub-micron effects where coupling capacitances between bus lines are in the same order of magnitude as base capacitances, power consumption of interconnects starts to have a significant impact on a system's total power consumption. We present novel address bus encoding schemes that take coupling effects into consideration. The basis is a physical bus model that quantifies coupling capacitances. As a result, we report power/energy savings on the address buses of up to 56% compared to the best known ordinary power/energy efficient encoding schemes. Thereby, we exceed the only to-date approach that also takes coupling effects into consideration. Moreover, our encoding schemes do not assume any a priori knowledge that is particular to a specific application.
Jörg Henkel, Haris Lekatsas
DAC2
2001 Code Compression for VLIW Processors
Yuan Xie 0001, Haris Lekatsas, Marilyn Wolf
Data Compression Conference2
2001 A code decompression architecture for VLIW processors
abstract
In embedded system design, memory has been one of the most restricted resources. Reducing program size has been an important goal when designing an embedded system. Most of the previous work on code compression has targeted RISC architectures. Recently VLIW processors became very popular, particularly for signal processing. Decompression speed is especially important for VLIW architectures given that the length of the instruction word is long. Furthermore, modem VLIW architectures use flexible instruction formats, which require new code compression approaches. Previous work has assumed that instruction positions within the long instruction word correspond to specific functional units. In contrast, our code compression algorithm is capable of compressing flexible instruction formats, where any functional unit can be used for any position in the instruction word. We demonstrate our methods by applying it to the TMS320C6x architecture. We also compare two techniques for decompressing the VLIW instruction packet to reduce the decompression time. A fast parallel decompression architecture is described, which is implemented in TSMC 0.25 technology.
Yuan Xie 0001, Marilyn Wolf, Haris Lekatsas
MICRO3
2000 Code compression for low power embedded system design
abstract
We propose instruction code compression as an efficient method for reducing power on an embedded system. Our approach is the first one to measure and optimize the power consumption of a complete SOC (System--On--a--Chip) comprising a CPU, instruction cache, data cache, main memory, data buses and address bus through code compression. We compare the pre-cache architecture (decompressor between main memory and cache) to a novel post-cache architecture (decompressor between cache and CPU). Our simulations and synthesis results show that our methodology results in large energy savings between 22% and 82% compared to the same system without code compression. Furthermore, we demonstrate that power savings come with reduced chip area and the same or even improved performance.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
DAC1
2000 Arithmetic Coding for Low Power Embedded System Design
abstract
We present a novel algorithm that assigns codes to instructions during instruction code compression in order to minimize bus-related bit-toggling and thus reduce power consumption. The target application area is embedded systems, where power consumption is increasingly becoming a dominant design constraint. Our algorithm is based on a variant of quasi-arithmetic coding where coding allows for random access and fast table-based decoding. We take advantage of the approximations introduced to modify codes and reduce bit-toggling, while maintaining compression performance and decoding speed. We present the first work to explore the trade-offs between compression ratios and bus-related power consumption and show that high compression ratios do not necessarily result in the lowest power consumption. By using our method, bus-related power consumption has been reduced by as much as 35% without imposing any additional hardware costs.
Haris Lekatsas, Marilyn Wolf, Jörg Henkel
Data Compression Conference1
2000 A Decompression Architecture for Low Power Embedded Systems
abstract
We present an architecture for embedded systems that decompresses offline-compressed instructions during runtime. This is useful for compressed code systems where instructions are stored in a compressed format and decompressed on demand. The result is a significant reduction in power consumption, and in most cases a performance improvement. The stand-alone decompression engine is placed between the instruction cache and the CPU (post-cache architecture) as we have found this to be the most power-efficient architecture. This paper describes the design of this unit in detail and analyzes its power consumption and performance.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
ICCD1
1999 Random Access Decompression Using Binary Arithmetic Coding
abstract
We present an algorithm based on arithmetic coding that allows decompression to start at any point in the compressed file. This random access requirement poses some restrictions on the implementation of arithmetic coding and on the model used. Our main application area is executable code compression for computer systems where machine instructions are decompressed on-the-fly before execution. We focus on the decompression side of arithmetic coding and we propose a fast decoding scheme based on finite state machines. Furthermore, we present a method to decode multiple bits per cycle, while keeping the size of the decoder small.
Haris Lekatsas, Marilyn Wolf
Data Compression Conference1
1999 SAMC: a code compression algorithm for embedded processors
abstract
In this paper, we present a method for reducing the memory requirements of an embedded system by using code compression. We compress the instruction segment of the executable running on the embedded system, and we show how to design a run-time decompression unit to decompress code on the fly before execution. Our algorithm uses arithmetic coding in combination with a Markov model, which is adapted to the instruction set and the application. We provide experimental results on two architectures, Analog Devices' Share and ARM's ARM and Thumb instruction sets, and show that programs can often be reduced by more than 50%. Furthermore, we suggest a table-based design that allows multibit decoding to speed up decompression.
Haris Lekatsas, Marilyn Wolf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1998 Code Compression for Embedded Systems
abstract
Memory is one of the most restricted resources in many modern embedded systems. Code compression can provide substantial savings in terms of size. In a compressed code CPU, a cache miss triggers the decompression of a main memory block, before it gets transferred to the cache. Because the code must be decompressible starting from any point (or at least at cache block boundaries), most file-oriented compression techniques cannot be used. We propose two algorithms to compress code in a space-efficient and simple to decompress way, one which is independent of the instruction set and another which depends on the instruction set. We perform experiments on two instruction sets, a typical RISC (MIPS) and a typical CISC (x86) and compare our results to existing file-oriented compression algorithms.
Haris Lekatsas, Marilyn Wolf
DAC1