Ismail Kadayif

dblp:49/3168 · DBLP profile ↗
← Back
34ranked-venue papers
21as first author
1since 2021 · last 2022
0000-0002-2206-2726ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 19 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 6 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Memory systems · 57% Energy-efficient computing · 14% Processor architecture and microarchitecture · 10%
Software engineering, system software, and programming languages
6 papers
Compilers and program optimization · 92% Operating systems · 8%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.312018
Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › cache coherence
directory-based coherence
0.312018
Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › virtual memory management
page table management
0.312018
Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Processor architecture and microarchitecture
chip multiprocessor
0.142018
Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005
An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002
Energy-efficient computing
power management
0.132005
Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005
Generating physical addresses directly for saving instruction TLB energy · MICRO 2002
An energy saving strategy based on adaptive loop parallelization · DAC 2002
Processor architecture and microarchitecture › chip multiprocessor
tiled CMP
0.112018
Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Compilers and program optimization › code generation
address code generation
0.112007
Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Compilers and program optimization › compiler optimization
compiler-directed optimization
0.112007
Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Memory systems › memory management › virtual memory
address translation
0.112007
Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Hardware reliability and fault tolerance › memory reliability
cache reliability
0.112007
Modeling and improving data cache reliability · SIGMETRICS 2007
Parallel and multicore computing › loop transformation
loop parallelization
0.122002
An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002
An energy saving strategy based on adaptive loop parallelization · DAC 2002
Parallel and multicore computing
parallel programming models and scheduling
0.122002
An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002
An energy saving strategy based on adaptive loop parallelization · DAC 2002
Hardware reliability and fault tolerance
soft errors
0.112007
Modeling and improving data cache reliability · SIGMETRICS 2007
Memory systems › memory management
virtual memory
0.112007
Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Memory systems › memory management › virtual memory › address translation
TLB
0.122007
Generating physical addresses directly for saving instruction TLB energy · MICRO 2002
Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization
0.112005
Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005
Compilers and program optimization › memory optimization
data locality optimization
0.012004
Quasidynamic Layout Optimizations for Improving Data Locality · IEEE Trans. Parallel Distributed Syst. 2004
Compilers and program optimization
memory optimization
0.012004
A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Embedded and real-time systems › embedded software › embedded operating systems
embedded memory management
0.012004
A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Energy-efficient computing › power management
low-power modes
0.012004
Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004
Memory systems › DRAM › DRAM architecture
memory bank
0.012004
Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004
Energy-efficient computing › power management
memory power management
0.012004
Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004
Energy-efficient computing
memory system energy
0.012004
Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004
Embedded and real-time systems › embedded software › embedded operating systems › embedded memory management
scratchpad memory management
0.012004
A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Energy-efficient computing › energy-aware software
energy-aware compilation
0.012002
An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002
Memory systems
cache
0.022007
Modeling and improving data cache reliability · SIGMETRICS 2007
Quasidynamic Layout Optimizations for Improving Data Locality · IEEE Trans. Parallel Distributed Syst. 2004
Operating systems › resource management › memory management
dynamic memory allocation
0.012001
Dynamic Management of Scratch-Pad Memory Space · DAC 2001
Compilers and program optimization › compiler optimization › compiler-directed memory management
scratch-pad memory management
0.012001
Dynamic Management of Scratch-Pad Memory Space · DAC 2001
Hardware reliability and fault tolerance
reliability modeling
0.012007
Modeling and improving data cache reliability · SIGMETRICS 2007
Compilers and program optimization
loop transformation
0.012004
Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004

Methods — techniques the papers use, named apart from their topics

simulation · 0.3loop transformation · 0.2data transformation · 0.2translation registers · 0.1integer linear programming · 0.1compile-time analysis · 0.1code transformation · 0.1source-to-source translation · 0.1polyhedral compilation · 0.1soft error modeling · 0.1static layout optimizer · 0.0
YearPublicationVenuePosition
2022 Coherency Traffic Reduction in Manycore Systems
abstract
With the increasing number of cores in manycore accelerators and chip multiprocessors (CMPs), it gets more challenging to provide cache coherency efficiently. Although the snooping-based protocols are appropriate solutions to small-scale systems, they are inefficient for large systems because of the limited bandwidth. Therefore, large-scale manycores require directory-based solutions where a hardware structure called directory holds the information. This directory keeps track of all memory blocks and which cache stores a copy of these blocks. The directory sends messages only to caches that store relevant blocks and also coordinate simultaneous accesses to a cache block. As directory-based protocols scale to many cores, performance, network-on-chip (NoC) traffic, and bandwidth become major problems. In this paper, we present software mechanisms to improve the effectiveness of directory-based cache coherency in manycore and multicore systems with shared memory. In multithreaded applications, some of the data accesses do not disrupt cache coherency, but they still produce coherency messages among cores such as read-only (private) data. However, if data is accessed by at least two cores and at least one of them is a write operation, it is called shared data and requires cache coherency. In our proposed system, private data and shared data are determined at compile time, and cache coherency protocol only applies to shared data. We implement our approach in two stages. First, we use Andersen's static pointer analysis to analyze the program and mark its private instructions, i.e., instructions that load or store private data. Then, we use these analyses to decide if cache coherency protocol will be applied or not at runtime. Our simulation results on parallel benchmarks show that our approach reduces cycle count, dynamic random access memory (DRAM) accesses, and coherency traffic up to 13%.
Erdem Derebasoglu, Ismail Kadayif, Ozcan Ozturk 0001
DSD2
2018 Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs
abstract
As shown in some prior studies, a significant percentage of data blocks accessed in parallel codes are private, and not keeping track of those blocks can improve the effectiveness of directory structures in Chip multiprocessors (CMPs). In this paper, we have two major contributions. First, we showed that compared to the classification of cache blocks at page granularity, data block classification (DBC) at subpage level helps to detect considerably more private data blocks. Based on this idea, we propose two different approaches for enhancing the effectiveness of directory caches in tiled CMPs. In the first approach, which is called quasi-dynamic subpage level DBC (QDBC), a data block is assumed to be private from the beginning of the program execution and stays private as long as the corresponding subpage is accessed by only one core. Our second approach, which is called dynamic subpage level DBC, turns a data block into private again after all blocks within the corresponding subpage are evicted from private cache hierarchy. Memory block classification at subpage level, however, may increase the frequency of the operating system involvement in updating the maintenance bits in page table entries. To overcome this, we propose, as a second contribution, a distributed table called as on-chip page table (o-CPT), which stores recently accessed page translations in the system. Our simulation results show that, compared to page level data classification, QDBC and DBC approaches relying on the o-CPT can detect significantly more private data blocks and considerably improve system performance.
Mohammadreza Soltaniyeh, Ismail Kadayif, Ozcan Ozturk 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Hardware/software approaches for reducing the process variation impact on instruction fetches
abstract
As technology moves towards finer process geometries, it is becoming extremely difficult to control critical physical parameters such as channel length, gate oxide thickness, and dopant ion concentration. Variations in these parameters lead to dramatic variations in access latencies in Static Random Access Memory (SRAM) devices. This means that different lines of the same cache may have different access latencies. A simple solution to this problem is to adopt the worst-case latency paradigm. While this egalitarian cache management is simple, it may introduce significant performance overhead during instruction fetches when both address translation (instruction Translation Lookaside Buffer (TLB) access) and instruction cache access take place, making this solution infeasible for future high-performance processors. In this study, we first propose some hardware and software enhancements and then, based on those, investigate several techniques to mitigate the effect of process variation on the instruction fetch pipeline stage in modern processors. For address translation, we study an approach that performs the virtual-to-physical page translation once, then stores it in a special register, reusing it as long as the execution remains on the same instruction page. To handle varying access latencies across different instruction cache lines, we annotate the cache access latency of instructions within themselves to give the circuitry a hint about how long to wait for the next instruction to become available.
Ismail Kadayif, Mahir Turkcan, Seher Kiziltepe, Ozcan Ozturk 0001
ACM Trans. Design Autom. Electr. Syst.1
2007 Modeling and improving data cache reliability
abstract
Soft errors arising from energetic particle strikes pose a significant reliability concern for computing systems, especially for those running in noisy environments. Technology scaling and aggressive leakage control mechanisms make the problem caused by these transient errors even more severe. Therefore, it is very important to employ reliability enhancing mechanisms in processor/memory designs to protect them against soft errors. To do so, we first need to model soft errors, and then study cost/reliability tradeoffs among various reliability enhancing techniques based on the model so that system requirements could be met.
Ismail Kadayif, Mahmut T. Kandemir
SIGMETRICS1
2007 Reducing Data TLB Power via Compiler-Directed Address Generation
abstract
Address translation using the translation lookaside buffer (TLB) consumes as much as 16% of the chip power on some processors because of its high associativity and access frequency. While prior work has looked into optimizing this structure at the circuit and architectural levels, this paper takes a different approach to optimizing its power by reducing the number of data TLB (dTLB) lookups for data references. The main idea is to keep translations in a set of translation registers (TRs) and intelligently use them in software to directly generate the physical addresses without going through the dTLB. The software has to work within the confines of the TRs provided by the hardware and has to maximize the reuse of such translations to be effective. The authors propose strategies and code transformations for achieving this in array-based and pointer-based codes, looking to optimize data accesses. Results with a suite of Spec95 array-based and pointer-based codes show dTLB energy savings of up to 73% and 88%, respectively, compared to directly using the dTLB for all references. Despite the small increase in instructions executed with the mechanisms, the approach can, in fact, provide performance benefits in certain cache-addressing strategies
Ismail Kadayif, Partho Nath, Mahmut T. Kandemir, Anand Sivasubramaniam
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 Prefetching-aware cache line turnoff for saving leakage energy
abstract
While numerous prior studies focused on performance and energy optimizations for caches, their interactions have received much less attention. This paper studies this interaction and demonstrates how performance and energy optimizations can affect each other. More importantly, we propose three optimization schemes that turn off cache lines in a prefetching-sensitive manner. These schemes treat prefetched cache lines differently from the lines brought to the cache in a normal way (i.e., through a load operation) in turning off the cache lines. Our experiments with applications from the SPEC2000 suite indicate that the proposed approaches save significant leakage energy with very small degradation on performance.
Ismail Kadayif, Mahmut T. Kandemir, Feihui Li
ASP-DAC1
2005 Studying interactions between prefetching and cache line turnoff
abstract
While lots of prior studies focused on performance and energy optimizations for caches, their interactions have received much less attention. This is unfortunate since in general the performance-oriented techniques influence energy behavior of the cache, and the energy-oriented techniques usually increase program execution cycles. The overall energy and performance behavior of caches in embedded systems when multiple techniques co-exist remains an open research problem. This paper studies this interaction and illustrates how performance and energy optimizations affect each other. We also point out several potential optimizations that could be based on this study.
Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen
ASP-DAC1
2005 Compiling for memory emergency
abstract
There has been a continued growth in the sales of mobile and embedded devices, in spite of the economic recession in many parts of the world. Many of these devices operate under tight memory bounds. Past research dealt with this problem and proposed both hardware and software solutions oriented toward reducing memory space requirements of embedded and mobile applications. One of the common characteristics of most of these prior efforts is that they assume the memory space available to an embedded application is fixed for the entire execution. Unfortunately, this is not a valid assumption in many execution environments since a typical embedded/mobile platform can have multiple applications executing concurrently and sharing a common memory space. As a result, the amount of memory available to a particular application can vary during execution. A particularly interesting scenario is what we call "memory emergency", where the size of the memory available to an application suddenly drops. If the application is not written to cope with this emergency scenario, the result would normally be a premature termination due to insufficient memory. In this paper, we propose compiler-based solutions to this memory emergency problem. The proposed compiler support modifies a given application code assuming a memory emergency model and reduces memory space demand (when necessary) by recomputing data values, thereby performing a tradeoff between memory space reduction and performance overhead. Our goal is to be able to work with the reduced memory space but minimize the performance overhead it brings. We evaluate the proposed approaches using twelve array-based embedded benchmarks. Our experimental analysis shows that the proposed approaches are very successful in responding to many memory emergency scenarios.
Mahmut T. Kandemir, Guangyu Chen, Ismail Kadayif
LCTES3
2005 An integer linear programming-based tool for wireless sensor networks
Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin
J. Parallel Distributed Comput.1
2005 Data space-oriented tiling for enhancing locality
abstract
Improving locality of data references is becoming increasingly important due to increasing gap between processor cycle times and off-chip memory access latencies. Improving data locality not only improves effective memory access time but also reduces memory system energy consumption due to data references. An optimizing compiler can play an important role in enhancing data locality in array-intensive embedded media applications with regular data access patterns.This paper presents a compiler-based data space-oriented tiling approach (DST). In this strategy, the data space (e.g., an array of signals) is logically divided into chunks (called data tiles) and each data tile is processed in turn. In processing a data tile, our approach traverses the entire iteration space of all nests in the code and executes all iterations (potentially coming from different nests) that access the data tile being processed. In doing so, it also takes data dependences into account. Since a data space is common across all nests that access it, DST can potentially achieve better results than traditional iteration space (loop) tiling by exploiting internest data locality.We also present an example application of DST for improving the effectiveness of a scratch pad memory (SPM) for data accesses. SPMs are alternatives to conventional cache memories in embedded computing world. These small on-chip memories, like caches, provide fast and low-power access to data; but, they differ from conventional data caches in that their contents are managed by compiler instead of hardware. We have implemented DST in a source-to-source translator and quantified its benefits using a simulator. Our preliminary results with several array-intensive applications and varying input sizes show that our approach outperforms classical iteration space-oriented tiling as well as a data-oriented approach that considers each nest in isolation.
Ismail Kadayif, Mahmut T. Kandemir
ACM Trans. Embed. Comput. Syst.1
2005 Compiler-directed high-level energy estimation and optimization
abstract
The demand for high-performance architectures and powerful battery-operated mobile devices has accentuated the need for power optimization. While many power-oriented hardware optimization techniques have been proposed and incorporated in current systems, the increasingly critical power constraints have made it essential to look for software-level optimizations as well. The compiler can play a pivotal role in addressing the power constraints of a system as it wields a significant influence on the application's runtime behavior. This paper presents a novel Energy-Aware Compilation (EAC) framework that estimates and optimizes energy consumption of a given code, taking as input the architectural and technological parameters, energy models, and energy/performance/code size constraints. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time. As shown in this paper, EAC allows compiler writers and system designers to investigate power-performance tradeoffs of traditional compiler optimizations and to develop energy-conscious high-level code transformations.
Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
ACM Trans. Embed. Comput. Syst.1
2005 Optimizing instruction TLB energy using software and hardware techniques
abstract
Power consumption and power density for the Translation Look-aside Buffer (TLB) are important considerations not only in its design, but can have a consequence on cache design as well. After pointing out the importance of instruction TLB (iTLB) power optimization, this article embarks on a new philosophy for reducing the number of accesses to this structure. The overall idea is to keep a translation currently being used in a register and avoid going to the iTLB as far as possible---until there is a page change. We propose four different approaches for achieving this, and experimentally demonstrate that one of these schemes that uses a combination of compiler and hardware enhancements can reduce iTLB dynamic power by over 85% in most cases.The proposed approaches can work with different instruction-cache (iL1) lookup mechanisms and achieve significant iTLB power savings without compromising on performance. Their importance grows with higher iL1 miss rates and larger page sizes. They can work very well with large iTLB structures that can possibly consume more power and take longer to lookup, without the iTLB getting into the common case. Further, we also experimentally demonstrate that they can provide performance savings for virtually indexed, virtually tagged iL1 caches, and can even make physically indexed, physically tagged iL1 caches a possible choice for implementation.
Ismail Kadayif, Anand Sivasubramaniam, Mahmut T. Kandemir, Gokul B. Kandiraju, Guangyu Chen
ACM Trans. Design Autom. Electr. Syst.1
2005 Optimizing Array-Intensive Applications for On-Chip Multiprocessors
abstract
With energy consumption becoming one of the first-class optimization parameters in computer system design, compilation techniques that consider performance and energy simultaneously are expected to play a central role. In particular, compiling a given application code under performance and energy constraints is becoming an important problem. In this paper, we focus on an on-chip multiprocessor architecture and present a set of code optimization strategies. We first evaluate an adaptive loop parallelization strategy (i.e., a strategy that allows each loop nest to execute using a different number of processors if doing so is beneficial) and measure the potential energy savings when unused processors during execution of a nested loop are shut down (i.e., placed into a power-down or sleep state). Our results show that shutting down unused processors can lead to as much as 67 percent energy savings at the expense of up to 17 percent performance loss in a set of array-intensive applications. To eliminate this performance penalty, we also discuss and evaluate a processor preactivation strategy based on compile-time analysis of nested loops. Based on our experiments, we conclude that an adaptive loop parallelization strategy combined with idle processor shut down and preactivation can be very effective in reducing energy consumption without increasing execution time. We then generalize our strategy and present an application parallelization strategy based on integer linear programming (ILP). Given an array-intensive application, our optimization strategy determines the number of processors to be used in executing each loop nest based on the objective function and additional compilation constraints provided by the user/programmer. Our initial experience with this constraint-based optimization strategy shows that it is very successful in optimizing array-intensive applications on on-chip multiprocessors under multiple energy and performance constraints.
Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen, Ozcan Ozturk 0001, Mustafa Karaköy, Ugur Sezer
IEEE Trans. Parallel Distributed Syst.1
2004 Tuning In-Sensor Data Filtering to Reduce Energy Consumption in Wireless Sensor Networks
abstract
In recent years, research on wireless sensor networks has been undergoing a revolution, promising to have significant impact on a broad range of applications from military to health care to food safety. An important problem in many sensor network applications is to decide the amount of computation (or filtering) that needs to be done in the sensor nodes before the data are shifted to a central base station. Right amount of data filtering in the sensor nodes can lead to large savings in network-wide energy consumption. The main goal of this paper is to develop an automated strategy for data filtering in wireless sensor nodes. Assuming that one needs to reduce the overall energy consumption (as opposed to reducing just computation energy or communication energy), the proposed strategy attempts to strike a balance between computation energy consumption and communication energy consumption. Our experimental results clearly indicate that the proposed data filtering strategy generates substantial energy savings in practice.
Ismail Kadayif, Mahmut T. Kandemir
DATE1
2004 Exploiting Processor Workload Heterogeneity for Reducing Energy Consumption in Chip Multiprocessors
abstract
Advances in semiconductor technology are enabling designs with several hundred million transistors. Since building sophisticated single processor based systems is a complex process from design, verification, and software development perspectives, the use of chip multiprocessing is inevitable in future microprocessors. In fact, the abundance of explicit loop-level parallelism in many embedded applications helps us identify chip multiprocessing as one of the most promising directions in designing systems for embedded applications. Another architectural trend that we observe in embedded systems, namely, multi-voltage processors, is driven by the need of reducing energy consumption during program execution. Practical implementations such as Transmeta's Crusoe and Intel's XScale tune processor voltage/frequency depending on current execution load. Considering these two trends, chip multiprocessing and voltage/frequency scaling, this paper presents an optimization strategy for an architecture that makes use of both chip parallelism and voltage scaling. In our proposal, the compiler takes advantage of heterogeneity in parallel execution between the loads of different processors and assigns different voltages/frequencies to different processors if doing so reduces energy consumption without increasing overall execution cycles significantly. Our experiments with a set of applications show that this optimization can bring large energy benefits without much performance loss.
Ismail Kadayif, Mahmut T. Kandemir, Ibrahim Kolcu
DATE1
2004 Compiler-Guided Code Restructuring for Improving Instruction TLB Energy Behavior
Ismail Kadayif, Mahmut T. Kandemir, I. Demirkiran
Euro-Par1
2004 Compiler-directed physical address generation for reducing dTLB power
abstract
Address translation using the Translation Lookaside Buffer (TLB) consumes as much as 16% of the chip power on some processors because of its high associativity and access frequency. While prior work has looked into optimizing this structure at the circuit and architectural levels, this paper takes a different approach of optimizing its power by reducing the number of data TLB (dTLB) lookups for data references. The main idea is to keep translations in a set of translation registers, and intelligently use them in software to directly generate the physical addresses without going through the dTLB. The software has to work within the confines of the translation registers provided by the hardware, and has to maximize the reuse of such translations to be effective. We propose strategies and code transformations for achieving this in array-based and pointer-based codes, looking to optimize data accesses. Results with a suite of Spec95 array-based and pointer-based codes show dTLB energy savings of up to 73% and 88%, respectively, compared to directly using the dTLB for all references. Despite the small increase in instructions executed with our mechanisms, the approach can in fact provide performance benefits in certain cases.
Ismail Kadayif, Partho Nath, Mahmut T. Kandemir, Anand Sivasubramaniam
ISPASS1
2004 A compiler-based approach for dynamically managing scratch-pad memories in embedded systems
abstract
Optimizations aimed at improving the efficiency of on-chip memories in embedded systems are extremely important. Using a suitable combination of program transformations and memory design space exploration aimed at enhancing data locality enables significant reductions in effective memory access latencies. While numerous compiler optimizations have been proposed to improve cache performance, there are relatively few techniques that focus on software-managed on-chip memories. It is well-known that software-managed memories are important in real-time embedded environments with hard deadlines as they allow one to accurately predict the amount of time a given code segment will take. In this paper, we propose and evaluate a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework. Our framework includes an optimization suite that uses loop and data transformations, an on-chip memory partitioning step, and a code-rewriting phase that collectively transform an input code automatically to take advantage of the on-chip SPM. Compared with previous work, the proposed scheme is dynamic, and allows the contents of the SPM to change during the course of execution, depending on the changes in the data access pattern. Experimental results from our implementation using a source-to-source translator and a generic cost model indicate significant reductions in data transfer activity between the SPM and off-chip memory.
Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2004 Access Pattern Restructuring for Memory Energy
abstract
Improving memory energy consumption of programs that manipulate arrays is an important problem as these codes spend large amounts of energy in accessing off-chip memory. We propose a data-driven strategy to optimize the memory energy consumption in a banked memory system. Our compiler-based strategy modifies the original execution order of loop iterations in array-dominated applications to increase the length of the time period(s) in which memory banks are idle (i.e., not accessed by any loop iteration). To achieve this, it first classifies loop iterations according to their bank accesses patterns and then, with the help of a polyhedral tool, tries to bring the iterations with similar bank access patterns close together. Increasing the idle periods of memory banks brings two major benefits: first, it allows us to place more memory banks into low-power operating modes and, second, it enables us to use a more aggressive (i.e., more energy saving) operating mode (hence, saving more energy) for a given bank (instead of a less aggressive mode). The proposed strategy can reduce memory energy consumption in both sequential and parallel applications. Our strategy has been implemented in an experimental compiler using a polyhedral tool and evaluated using nine array-dominated applications on both a cacheless system and a system with cache memory. Our experimental results indicate that the proposed strategy is very successful in reducing the memory system energy and improves the memory energy by as much as 36.8 percent over a strategy that uses low-power modes without optimizing data access pattern. Our results also show that optimizations that target reducing off-chip memory energy can generate very different results from those that target at improving only cache locality.
Victor M. DeLaLuz, Ismail Kadayif, Mahmut T. Kandemir, Ugur Sezer
IEEE Trans. Parallel Distributed Syst.2
2004 Quasidynamic Layout Optimizations for Improving Data Locality
abstract
Compiler-directed locality optimization techniques are effective in reducing the number of cycles spent in off-chip memory accesses. Recently, methods have been developed that transform memory layouts of data structures at compile-time to improve spatial locality of nested loops beyond current control-centric (loop nest-based) optimizations. Most of these data-centric transformations use a single static (program-wide) memory layout for each array. A disadvantage of these static layout-based locality enhancement strategies is that they might fail to optimize codes that manipulate arrays, which demand different layouts in different parts of the code. We introduce a new approach, which extends current static layout optimization techniques by associating different memory layouts with the same array in different parts of the code. We call this strategy "quasidynamic layout optimization." In this strategy, the compiler determines memory layouts (for different parts of the code) at compile time, but layout conversions occur at runtime. We show that the possibility of dynamically changing memory layouts during the course of execution adds a new dimension to the data locality optimization problem. Our strategy employs a static layout optimizer module as a building block and, by repeatedly invoking it for different parts of the code, it checks whether runtime layout modifications bring additional benefits beyond static optimization. Our experiments indicate significant improvements in execution time over static layout-based locality enhancing techniques.
Ismail Kadayif, Mahmut T. Kandemir
IEEE Trans. Parallel Distributed Syst.1
2004 Compiler-directed scratch pad memory optimization for embedded multiprocessors
abstract
This paper presents a compiler strategy to optimize data accesses in regular array-intensive applications running on embedded multiprocessor environments. Specifically, we propose an optimization algorithm that targets at reducing extra off-chip memory accesses caused by interprocessor communication. This is achieved by increasing the application-wide reuse of data that resides in scratch-pad memories of processors. Our results obtained using four array-intensive image processing applications indicate that exploiting interprocessor data sharing can reduce energy-delay product significantly on a four-processor embedded system.
Mahmut T. Kandemir, Ismail Kadayif, Alok N. Choudhary, Ibrahim Kolcu
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Generalized Data Transformations for Enhancing Cache Behavior
Victor M. DeLaLuz, Mahmut T. Kandemir, Ismail Kadayif, Ugur Sezer
DATE3
2003 An Integrated Approach for Improving Cache Behavior
Gokhan Memik, Mahmut T. Kandemir, Alok N. Choudhary, Ismail Kadayif
DATE4
2003 Compiler-Directed Management of Instruction Accesses
abstract
We present a compiler-oriented strategy to reduce the memory system energy consumption due to instruction accesses and increase performance by exploiting scratch pad memories. Scratch pad memories (SPMs) are alternatives to conventional cache memories in embedded computing. These small on-chip memories, like caches, provide fast and low-power access to data and instructions; but, they differ from caches in that their contents are managed by software instead of hardware. Our compiler framework keeps the most frequently used instructions in SPM and dynamically changes the contents of the SPM as the (instruction) working set of the application changes.
Guilin Chen, Guangyu Chen, Ismail Kadayif, Wei Zhang 0002, Mahmut T. Kandemir, Ibrahim Kolcu, Ugur Sezer
DSD3
2003 CCC: Crossbar Connected Caches for Reducing Energy Consumption of On-Chip Multiprocessors
abstract
With shrinking feature size of silicon fabrication technology, architects are putting more and more logic into a single die. While one might opt to use these transistors for building complex single processor based architectures, recent trends indicate a shift towards on-chip multiprocessor systems since they are simpler to implement and can provide better performance. An important problem in on-chip multiprocessors is energy consumption. In particular, on-chip cache structures can be major energy consumers. In this work, we study energy behavior of different cache architectures, and propose a new architecture, where processors share a single, banked cache using crossbar interconnects. Our detailed cycle-accurate simulations show that this cache architecture brings energy benefits ranging from 9% to 26% (over an architecture where each processor has a private cache).
Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Ismail Kadayif
DSD5
2003 An Energy-Oriented Evaluation of Communication Optimizations for Microcensor Networks
Ismail Kadayif, Mahmut T. Kandemir, Alok N. Choudhary, Mustafa Karaköy
Euro-Par1
2002 Optimizing inter-nest data locality
abstract
By examining data reuse patterns of four array-intensive embedded applications, we found that these codes exhibit a significant amount of inter-nest reuse (i. e., the data reuse that occurs between different nests). While traditional compiler techniques that target array-intensive applications can exploit intra-nest data reuse, there has not been much success in the past in taking advantage of internest data reuse. In this paper, we present a compiler strategy that optimizes inter-nest reuse using loop (iteration space) transformations. Our approach captures the impact of execution of a nest on cache contents using an abstraction called footprint vector. Then, it transforms a given nest such that the new (transformed) access pattern reuses the data left in cache by the previous nest in the code. In optimizing inter-nest locality, our approach also tries to achieve good intra-nest locality. Our simulation results indicate large performance improvements. In particular, inter-nest loop optimization generates competitive results with intra-nest loop and data optimizations.
Mahmut T. Kandemir, Ismail Kadayif, Alok N. Choudhary, Joseph Zambreno
CASES2
2002 Influence of Loop Optimizations on Energy Consumption of Multi-bank Memory Systems
Mahmut T. Kandemir, Ibrahim Kolcu, Ismail Kadayif
CC3
2002 An energy saving strategy based on adaptive loop parallelization
abstract
In this paper, we evaluate an adaptive loop parallelization strategy (i.e., a strategy that allows each loop nest to execute using different number of processors if doing so is beneficial) and measure the potential energy savings when unused processors during execution of a nested loop in a multi-processor on-a-chip (MPoC) are shut down (i.e., placed into a power-down or sleep state). Our results show that shutting down unused processors can lead to as much as 67% energy savings with up to 17% performance loss in a set of array-intensive applications. We also discuss and evaluate a processor pre-activation strategy based on compile-time analysis of nested loops. Based on our experiments, we conclude that an adaptive loop parallelization strategy combined with idle processor shut-down and pre-activation can be very effective in reducing energy consumption without increasing execution time.
Ismail Kadayif, Mahmut T. Kandemir, Mustafa Karaköy
DAC1
2002 An integer linear programming based approach for parallelizing applications in On-chip multiprocessors
abstract
With energy consumption becoming one of the first-class optimization parameters in computer system design, compilation techniques that consider performance and energy simultaneously are expected to play a central role. In particular, compiling a given application code under performance and energy constraints is becoming an important problem. In this paper, we focus on an on-chip multiprocessor architecture and present a parallelization strategy based on integer linear programming. Given an array-intensive application, our optimization strategy determines the number of processors to be used in executing each nest based on the objective function and additional compilation constraints provided by the user. Our initial experience with this strategy shows that it is very successful in optimizing array-intensive applications on on chip multiprocessors under energy and performance constraints.
Ismail Kadayif, Mahmut T. Kandemir, Ugur Sezer
DAC1
2002 EAC: A Compiler Framework for High-Level Energy Estimation and Optimization
abstract
This paper presents a novel Energy-Aware Compilation (EAC) framework that can estimate and optimize energy consumption of a given code taking as input the architectural and technological parameters, energy models, and energy/performance constraints,. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time.
Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
DATE1
2002 Generating physical addresses directly for saving instruction TLB energy
abstract
Power consumption and power density for the Translation Lookaside Buffer (TLB) are important considerations not only in its design, but can have a consequence on cache design as well. This paper embarks on a new philosophy for reducing the number of accesses to the instruction TLB (iTLB) for power and performance optimizations. The overall idea is to keep a translation currently being used in a register and avoid going to the iTLB as far as possible - until there is a page change. We propose four different approaches for achieving this, and experimentally demonstrate that one of these schemes that uses a combination of compiler and hardware enhancements can reduce iTLB dynamic power by over 85% in most cases. These mechanisms can work with different instruction-cache (iLl) lookup mechanisms and achieve significant iTLB power savings without compromising on performance. Their importance grows with higher iLl miss rates and larger page sizes. They can work very well with large iTLB structures, that can possibly consume more power and take longer to lookup, without the iTLB getting into the common case. Further, we also experimentally demonstrate that they can provide performance savings for virtually-indexed, virtually-tagged iLl caches, and can even make physically-indexed, physically-tagged iLl caches a possible choice for implementation.
Ismail Kadayif, Anand Sivasubramaniam, Mahmut T. Kandemir, Gokul B. Kandiraju, Guangyu Chen
MICRO1
2001 Dynamic Management of Scratch-Pad Memory Space
abstract
Optimizations aimed at improving the efficiency of on-chip memories are extremely important. We propose a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework that uses both loop and data transformations. Experimental results obtained using a generic cost model indicate significant reductions in data transfer activity between SPM and off-chip memory.
Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh
DAC5
2001 vEC: virtual energy counters
abstract
Energy has become a critical issue in processor design, especially in embedded environments. Thus, there is a need for tools, which provide an accurate and fast estimation of energy. In this paper, we present the design and use of a tool, Virtual Energy Counters (vEC), for estimating the energy consumption of user programs. vEC is built on top of the Perfmon user library for the UltraSPARC platform, and provides a user interface, which can be used within user programs to estimate the energy consumption. The energy estimates are provided for those consumed in the data, instruction and extended caches, main memory, address bus, data bus, address pads, and data pads.
Ismail Kadayif, T. Chinoda, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam
PASTE1