VLDB 2026 Research / reviewers in the wild / expert
Ismail Kadayif
dblp:49/3168
· DBLP profile ↗
34ranked-venue papers
21as first author
1since 2021 · last 2022
0000-0002-2206-2726ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 19 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Memory systems · 57% Energy-efficient computing · 14% Processor architecture and microarchitecture · 10% | |
| Software engineering, system software, and programming languages
6 papers |
Compilers and program optimization · 92% Operating systems · 8% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.3 | 1 | 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems › cache coherence
directory-based coherence |
0.3 | 1 | 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems › virtual memory management
page table management |
0.3 | 1 | 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 4 | 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005 An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002 |
Energy-efficient computing
power management |
0.1 | 3 | 2005 | Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005 Generating physical addresses directly for saving instruction TLB energy · MICRO 2002 An energy saving strategy based on adaptive loop parallelization · DAC 2002 |
Processor architecture and microarchitecture › chip multiprocessor
tiled CMP |
0.1 | 1 | 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Compilers and program optimization › code generation
address code generation |
0.1 | 1 | 2007 | Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Compilers and program optimization › compiler optimization
compiler-directed optimization |
0.1 | 1 | 2007 | Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Memory systems › memory management › virtual memory
address translation |
0.1 | 1 | 2007 | Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Hardware reliability and fault tolerance › memory reliability
cache reliability |
0.1 | 1 | 2007 | Modeling and improving data cache reliability · SIGMETRICS 2007 |
Parallel and multicore computing › loop transformation
loop parallelization |
0.1 | 2 | 2002 | An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002 An energy saving strategy based on adaptive loop parallelization · DAC 2002 |
Parallel and multicore computing
parallel programming models and scheduling |
0.1 | 2 | 2002 | An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002 An energy saving strategy based on adaptive loop parallelization · DAC 2002 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2007 | Modeling and improving data cache reliability · SIGMETRICS 2007 |
Memory systems › memory management
virtual memory |
0.1 | 1 | 2007 | Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.1 | 2 | 2007 | Generating physical addresses directly for saving instruction TLB energy · MICRO 2002 Reducing Data TLB Power via Compiler-Directed Address Generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization |
0.1 | 1 | 2005 | Optimizing Array-Intensive Applications for On-Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2005 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 2004 | Quasidynamic Layout Optimizations for Improving Data Locality · IEEE Trans. Parallel Distributed Syst. 2004 |
Compilers and program optimization
memory optimization |
0.0 | 1 | 2004 | A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 |
Embedded and real-time systems › embedded software › embedded operating systems
embedded memory management |
0.0 | 1 | 2004 | A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 |
Energy-efficient computing › power management
low-power modes |
0.0 | 1 | 2004 | Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004 |
Memory systems › DRAM › DRAM architecture
memory bank |
0.0 | 1 | 2004 | Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004 |
Energy-efficient computing › power management
memory power management |
0.0 | 1 | 2004 | Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004 |
Energy-efficient computing
memory system energy |
0.0 | 1 | 2004 | Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004 |
Embedded and real-time systems › embedded software › embedded operating systems › embedded memory management
scratchpad memory management |
0.0 | 1 | 2004 | A compiler-based approach for dynamically managing scratch-pad memories in embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 |
Energy-efficient computing › energy-aware software
energy-aware compilation |
0.0 | 1 | 2002 | An integer linear programming based approach for parallelizing applications in On-chip multiprocessors · DAC 2002 |
Memory systems
cache |
0.0 | 2 | 2007 | Modeling and improving data cache reliability · SIGMETRICS 2007 Quasidynamic Layout Optimizations for Improving Data Locality · IEEE Trans. Parallel Distributed Syst. 2004 |
Operating systems › resource management › memory management
dynamic memory allocation |
0.0 | 1 | 2001 | Dynamic Management of Scratch-Pad Memory Space · DAC 2001 |
Compilers and program optimization › compiler optimization › compiler-directed memory management
scratch-pad memory management |
0.0 | 1 | 2001 | Dynamic Management of Scratch-Pad Memory Space · DAC 2001 |
Hardware reliability and fault tolerance
reliability modeling |
0.0 | 1 | 2007 | Modeling and improving data cache reliability · SIGMETRICS 2007 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 2004 | Access Pattern Restructuring for Memory Energy · IEEE Trans. Parallel Distributed Syst. 2004 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.3loop transformation · 0.2data transformation · 0.2translation registers · 0.1integer linear programming · 0.1compile-time analysis · 0.1code transformation · 0.1source-to-source translation · 0.1polyhedral compilation · 0.1soft error modeling · 0.1static layout optimizer · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Coherency Traffic Reduction in Manycore SystemsabstractWith the increasing number of cores in manycore accelerators and chip multiprocessors (CMPs), it gets more challenging to provide cache coherency efficiently. Although the snooping-based protocols are appropriate solutions to small-scale systems, they are inefficient for large systems because of the limited bandwidth. Therefore, large-scale manycores require directory-based solutions where a hardware structure called directory holds the information. This directory keeps track of all memory blocks and which cache stores a copy of these blocks. The directory sends messages only to caches that store relevant blocks and also coordinate simultaneous accesses to a cache block. As directory-based protocols scale to many cores, performance, network-on-chip (NoC) traffic, and bandwidth become major problems. In this paper, we present software mechanisms to improve the effectiveness of directory-based cache coherency in manycore and multicore systems with shared memory. In multithreaded applications, some of the data accesses do not disrupt cache coherency, but they still produce coherency messages among cores such as read-only (private) data. However, if data is accessed by at least two cores and at least one of them is a write operation, it is called shared data and requires cache coherency. In our proposed system, private data and shared data are determined at compile time, and cache coherency protocol only applies to shared data. We implement our approach in two stages. First, we use Andersen's static pointer analysis to analyze the program and mark its private instructions, i.e., instructions that load or store private data. Then, we use these analyses to decide if cache coherency protocol will be applied or not at runtime. Our simulation results on parallel benchmarks show that our approach reduces cycle count, dynamic random access memory (DRAM) accesses, and coherency traffic up to 13%. Erdem Derebasoglu, Ismail Kadayif, Ozcan Ozturk 0001 |
DSD | 2 |
| 2018 | Classifying Data Blocks at Subpage Granularity With an On-Chip Page Table to Improve Coherence in Tiled CMPsabstractAs shown in some prior studies, a significant percentage of data blocks accessed in parallel codes are private, and not keeping track of those blocks can improve the effectiveness of directory structures in Chip multiprocessors (CMPs). In this paper, we have two major contributions. First, we showed that compared to the classification of cache blocks at page granularity, data block classification (DBC) at subpage level helps to detect considerably more private data blocks. Based on this idea, we propose two different approaches for enhancing the effectiveness of directory caches in tiled CMPs. In the first approach, which is called quasi-dynamic subpage level DBC (QDBC), a data block is assumed to be private from the beginning of the program execution and stays private as long as the corresponding subpage is accessed by only one core. Our second approach, which is called dynamic subpage level DBC, turns a data block into private again after all blocks within the corresponding subpage are evicted from private cache hierarchy. Memory block classification at subpage level, however, may increase the frequency of the operating system involvement in updating the maintenance bits in page table entries. To overcome this, we propose, as a second contribution, a distributed table called as on-chip page table (o-CPT), which stores recently accessed page translations in the system. Our simulation results show that, compared to page level data classification, QDBC and DBC approaches relying on the o-CPT can detect significantly more private data blocks and considerably improve system performance. Mohammadreza Soltaniyeh, Ismail Kadayif, Ozcan Ozturk 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Hardware/software approaches for reducing the process variation impact on instruction fetchesabstractAs technology moves towards finer process geometries, it is becoming extremely difficult to control critical physical parameters such as channel length, gate oxide thickness, and dopant ion concentration. Variations in these parameters lead to dramatic variations in access latencies in Static Random Access Memory (SRAM) devices. This means that different lines of the same cache may have different access latencies. A simple solution to this problem is to adopt the worst-case latency paradigm. While this egalitarian cache management is simple, it may introduce significant performance overhead during instruction fetches when both address translation (instruction Translation Lookaside Buffer (TLB) access) and instruction cache access take place, making this solution infeasible for future high-performance processors. In this study, we first propose some hardware and software enhancements and then, based on those, investigate several techniques to mitigate the effect of process variation on the instruction fetch pipeline stage in modern processors. For address translation, we study an approach that performs the virtual-to-physical page translation once, then stores it in a special register, reusing it as long as the execution remains on the same instruction page. To handle varying access latencies across different instruction cache lines, we annotate the cache access latency of instructions within themselves to give the circuitry a hint about how long to wait for the next instruction to become available. Ismail Kadayif, Mahir Turkcan, Seher Kiziltepe, Ozcan Ozturk 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2007 | Modeling and improving data cache reliabilityabstractSoft errors arising from energetic particle strikes pose a significant reliability concern for computing systems, especially for those running in noisy environments. Technology scaling and aggressive leakage control mechanisms make the problem caused by these transient errors even more severe. Therefore, it is very important to employ reliability enhancing mechanisms in processor/memory designs to protect them against soft errors. To do so, we first need to model soft errors, and then study cost/reliability tradeoffs among various reliability enhancing techniques based on the model so that system requirements could be met. Ismail Kadayif, Mahmut T. Kandemir |
SIGMETRICS | 1 |
| 2007 | Reducing Data TLB Power via Compiler-Directed Address GenerationabstractAddress translation using the translation lookaside buffer (TLB) consumes as much as 16% of the chip power on some processors because of its high associativity and access frequency. While prior work has looked into optimizing this structure at the circuit and architectural levels, this paper takes a different approach to optimizing its power by reducing the number of data TLB (dTLB) lookups for data references. The main idea is to keep translations in a set of translation registers (TRs) and intelligently use them in software to directly generate the physical addresses without going through the dTLB. The software has to work within the confines of the TRs provided by the hardware and has to maximize the reuse of such translations to be effective. The authors propose strategies and code transformations for achieving this in array-based and pointer-based codes, looking to optimize data accesses. Results with a suite of Spec95 array-based and pointer-based codes show dTLB energy savings of up to 73% and 88%, respectively, compared to directly using the dTLB for all references. Despite the small increase in instructions executed with the mechanisms, the approach can, in fact, provide performance benefits in certain cache-addressing strategies Ismail Kadayif, Partho Nath, Mahmut T. Kandemir, Anand Sivasubramaniam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Prefetching-aware cache line turnoff for saving leakage energyabstractWhile numerous prior studies focused on performance and energy optimizations for caches, their interactions have received much less attention. This paper studies this interaction and demonstrates how performance and energy optimizations can affect each other. More importantly, we propose three optimization schemes that turn off cache lines in a prefetching-sensitive manner. These schemes treat prefetched cache lines differently from the lines brought to the cache in a normal way (i.e., through a load operation) in turning off the cache lines. Our experiments with applications from the SPEC2000 suite indicate that the proposed approaches save significant leakage energy with very small degradation on performance. Ismail Kadayif, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 1 |
| 2005 | Studying interactions between prefetching and cache line turnoffabstractWhile lots of prior studies focused on performance and energy optimizations for caches, their interactions have received much less attention. This is unfortunate since in general the performance-oriented techniques influence energy behavior of the cache, and the energy-oriented techniques usually increase program execution cycles. The overall energy and performance behavior of caches in embedded systems when multiple techniques co-exist remains an open research problem. This paper studies this interaction and illustrates how performance and energy optimizations affect each other. We also point out several potential optimizations that could be based on this study. Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen |
ASP-DAC | 1 |
| 2005 | Compiling for memory emergencyabstractThere has been a continued growth in the sales of mobile and embedded devices, in spite of the economic recession in many parts of the world. Many of these devices operate under tight memory bounds. Past research dealt with this problem and proposed both hardware and software solutions oriented toward reducing memory space requirements of embedded and mobile applications. One of the common characteristics of most of these prior efforts is that they assume the memory space available to an embedded application is fixed for the entire execution. Unfortunately, this is not a valid assumption in many execution environments since a typical embedded/mobile platform can have multiple applications executing concurrently and sharing a common memory space. As a result, the amount of memory available to a particular application can vary during execution. A particularly interesting scenario is what we call "memory emergency", where the size of the memory available to an application suddenly drops. If the application is not written to cope with this emergency scenario, the result would normally be a premature termination due to insufficient memory. In this paper, we propose compiler-based solutions to this memory emergency problem. The proposed compiler support modifies a given application code assuming a memory emergency model and reduces memory space demand (when necessary) by recomputing data values, thereby performing a tradeoff between memory space reduction and performance overhead. Our goal is to be able to work with the reduced memory space but minimize the performance overhead it brings. We evaluate the proposed approaches using twelve array-based embedded benchmarks. Our experimental analysis shows that the proposed approaches are very successful in responding to many memory emergency scenarios. Mahmut T. Kandemir, Guangyu Chen, Ismail Kadayif |
LCTES | 3 |
| 2005 | An integer linear programming-based tool for wireless sensor networks
Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin |
J. Parallel Distributed Comput. | 1 |
| 2005 | Data space-oriented tiling for enhancing localityabstractImproving locality of data references is becoming increasingly important due to increasing gap between processor cycle times and off-chip memory access latencies. Improving data locality not only improves effective memory access time but also reduces memory system energy consumption due to data references. An optimizing compiler can play an important role in enhancing data locality in array-intensive embedded media applications with regular data access patterns.This paper presents a compiler-based data space-oriented tiling approach (DST). In this strategy, the data space (e.g., an array of signals) is logically divided into chunks (called data tiles) and each data tile is processed in turn. In processing a data tile, our approach traverses the entire iteration space of all nests in the code and executes all iterations (potentially coming from different nests) that access the data tile being processed. In doing so, it also takes data dependences into account. Since a data space is common across all nests that access it, DST can potentially achieve better results than traditional iteration space (loop) tiling by exploiting internest data locality.We also present an example application of DST for improving the effectiveness of a scratch pad memory (SPM) for data accesses. SPMs are alternatives to conventional cache memories in embedded computing world. These small on-chip memories, like caches, provide fast and low-power access to data; but, they differ from conventional data caches in that their contents are managed by compiler instead of hardware. We have implemented DST in a source-to-source translator and quantified its benefits using a simulator. Our preliminary results with several array-intensive applications and varying input sizes show that our approach outperforms classical iteration space-oriented tiling as well as a data-oriented approach that considers each nest in isolation. Ismail Kadayif, Mahmut T. Kandemir |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2005 | Compiler-directed high-level energy estimation and optimizationabstractThe demand for high-performance architectures and powerful battery-operated mobile devices has accentuated the need for power optimization. While many power-oriented hardware optimization techniques have been proposed and incorporated in current systems, the increasingly critical power constraints have made it essential to look for software-level optimizations as well. The compiler can play a pivotal role in addressing the power constraints of a system as it wields a significant influence on the application's runtime behavior. This paper presents a novel Energy-Aware Compilation (EAC) framework that estimates and optimizes energy consumption of a given code, taking as input the architectural and technological parameters, energy models, and energy/performance/code size constraints. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time. As shown in this paper, EAC allows compiler writers and system designers to investigate power-performance tradeoffs of traditional compiler optimizations and to develop energy-conscious high-level code transformations. Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2005 | Optimizing instruction TLB energy using software and hardware techniquesabstractPower consumption and power density for the Translation Look-aside Buffer (TLB) are important considerations not only in its design, but can have a consequence on cache design as well. After pointing out the importance of instruction TLB (iTLB) power optimization, this article embarks on a new philosophy for reducing the number of accesses to this structure. The overall idea is to keep a translation currently being used in a register and avoid going to the iTLB as far as possible---until there is a page change. We propose four different approaches for achieving this, and experimentally demonstrate that one of these schemes that uses a combination of compiler and hardware enhancements can reduce iTLB dynamic power by over 85% in most cases.The proposed approaches can work with different instruction-cache (iL1) lookup mechanisms and achieve significant iTLB power savings without compromising on performance. Their importance grows with higher iL1 miss rates and larger page sizes. They can work very well with large iTLB structures that can possibly consume more power and take longer to lookup, without the iTLB getting into the common case. Further, we also experimentally demonstrate that they can provide performance savings for virtually indexed, virtually tagged iL1 caches, and can even make physically indexed, physically tagged iL1 caches a possible choice for implementation. Ismail Kadayif, Anand Sivasubramaniam, Mahmut T. Kandemir, Gokul B. Kandiraju, Guangyu Chen |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2005 | Optimizing Array-Intensive Applications for On-Chip MultiprocessorsabstractWith energy consumption becoming one of the first-class optimization parameters in computer system design, compilation techniques that consider performance and energy simultaneously are expected to play a central role. In particular, compiling a given application code under performance and energy constraints is becoming an important problem. In this paper, we focus on an on-chip multiprocessor architecture and present a set of code optimization strategies. We first evaluate an adaptive loop parallelization strategy (i.e., a strategy that allows each loop nest to execute using a different number of processors if doing so is beneficial) and measure the potential energy savings when unused processors during execution of a nested loop are shut down (i.e., placed into a power-down or sleep state). Our results show that shutting down unused processors can lead to as much as 67 percent energy savings at the expense of up to 17 percent performance loss in a set of array-intensive applications. To eliminate this performance penalty, we also discuss and evaluate a processor preactivation strategy based on compile-time analysis of nested loops. Based on our experiments, we conclude that an adaptive loop parallelization strategy combined with idle processor shut down and preactivation can be very effective in reducing energy consumption without increasing execution time. We then generalize our strategy and present an application parallelization strategy based on integer linear programming (ILP). Given an array-intensive application, our optimization strategy determines the number of processors to be used in executing each loop nest based on the objective function and additional compilation constraints provided by the user/programmer. Our initial experience with this constraint-based optimization strategy shows that it is very successful in optimizing array-intensive applications on on-chip multiprocessors under multiple energy and performance constraints. Ismail Kadayif, Mahmut T. Kandemir, Guilin Chen, Ozcan Ozturk 0001, Mustafa Karaköy, Ugur Sezer |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Tuning In-Sensor Data Filtering to Reduce Energy Consumption in Wireless Sensor NetworksabstractIn recent years, research on wireless sensor networks has been undergoing a revolution, promising to have significant impact on a broad range of applications from military to health care to food safety. An important problem in many sensor network applications is to decide the amount of computation (or filtering) that needs to be done in the sensor nodes before the data are shifted to a central base station. Right amount of data filtering in the sensor nodes can lead to large savings in network-wide energy consumption. The main goal of this paper is to develop an automated strategy for data filtering in wireless sensor nodes. Assuming that one needs to reduce the overall energy consumption (as opposed to reducing just computation energy or communication energy), the proposed strategy attempts to strike a balance between computation energy consumption and communication energy consumption. Our experimental results clearly indicate that the proposed data filtering strategy generates substantial energy savings in practice. Ismail Kadayif, Mahmut T. Kandemir |
DATE | 1 |
| 2004 | Exploiting Processor Workload Heterogeneity for Reducing Energy Consumption in Chip MultiprocessorsabstractAdvances in semiconductor technology are enabling designs with several hundred million transistors. Since building sophisticated single processor based systems is a complex process from design, verification, and software development perspectives, the use of chip multiprocessing is inevitable in future microprocessors. In fact, the abundance of explicit loop-level parallelism in many embedded applications helps us identify chip multiprocessing as one of the most promising directions in designing systems for embedded applications. Another architectural trend that we observe in embedded systems, namely, multi-voltage processors, is driven by the need of reducing energy consumption during program execution. Practical implementations such as Transmeta's Crusoe and Intel's XScale tune processor voltage/frequency depending on current execution load. Considering these two trends, chip multiprocessing and voltage/frequency scaling, this paper presents an optimization strategy for an architecture that makes use of both chip parallelism and voltage scaling. In our proposal, the compiler takes advantage of heterogeneity in parallel execution between the loads of different processors and assigns different voltages/frequencies to different processors if doing so reduces energy consumption without increasing overall execution cycles significantly. Our experiments with a set of applications show that this optimization can bring large energy benefits without much performance loss. Ismail Kadayif, Mahmut T. Kandemir, Ibrahim Kolcu |
DATE | 1 |
| 2004 | Compiler-Guided Code Restructuring for Improving Instruction TLB Energy Behavior
Ismail Kadayif, Mahmut T. Kandemir, I. Demirkiran |
Euro-Par | 1 |
| 2004 | Compiler-directed physical address generation for reducing dTLB powerabstractAddress translation using the Translation Lookaside Buffer (TLB) consumes as much as 16% of the chip power on some processors because of its high associativity and access frequency. While prior work has looked into optimizing this structure at the circuit and architectural levels, this paper takes a different approach of optimizing its power by reducing the number of data TLB (dTLB) lookups for data references. The main idea is to keep translations in a set of translation registers, and intelligently use them in software to directly generate the physical addresses without going through the dTLB. The software has to work within the confines of the translation registers provided by the hardware, and has to maximize the reuse of such translations to be effective. We propose strategies and code transformations for achieving this in array-based and pointer-based codes, looking to optimize data accesses. Results with a suite of Spec95 array-based and pointer-based codes show dTLB energy savings of up to 73% and 88%, respectively, compared to directly using the dTLB for all references. Despite the small increase in instructions executed with our mechanisms, the approach can in fact provide performance benefits in certain cases. Ismail Kadayif, Partho Nath, Mahmut T. Kandemir, Anand Sivasubramaniam |
ISPASS | 1 |
| 2004 | A compiler-based approach for dynamically managing scratch-pad memories in embedded systemsabstractOptimizations aimed at improving the efficiency of on-chip memories in embedded systems are extremely important. Using a suitable combination of program transformations and memory design space exploration aimed at enhancing data locality enables significant reductions in effective memory access latencies. While numerous compiler optimizations have been proposed to improve cache performance, there are relatively few techniques that focus on software-managed on-chip memories. It is well-known that software-managed memories are important in real-time embedded environments with hard deadlines as they allow one to accurately predict the amount of time a given code segment will take. In this paper, we propose and evaluate a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework. Our framework includes an optimization suite that uses loop and data transformations, an on-chip memory partitioning step, and a code-rewriting phase that collectively transform an input code automatically to take advantage of the on-chip SPM. Compared with previous work, the proposed scheme is dynamic, and allows the contents of the SPM to change during the course of execution, depending on the changes in the data access pattern. Experimental results from our implementation using a source-to-source translator and a generic cost model indicate significant reductions in data transfer activity between the SPM and off-chip memory. Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2004 | Access Pattern Restructuring for Memory EnergyabstractImproving memory energy consumption of programs that manipulate arrays is an important problem as these codes spend large amounts of energy in accessing off-chip memory. We propose a data-driven strategy to optimize the memory energy consumption in a banked memory system. Our compiler-based strategy modifies the original execution order of loop iterations in array-dominated applications to increase the length of the time period(s) in which memory banks are idle (i.e., not accessed by any loop iteration). To achieve this, it first classifies loop iterations according to their bank accesses patterns and then, with the help of a polyhedral tool, tries to bring the iterations with similar bank access patterns close together. Increasing the idle periods of memory banks brings two major benefits: first, it allows us to place more memory banks into low-power operating modes and, second, it enables us to use a more aggressive (i.e., more energy saving) operating mode (hence, saving more energy) for a given bank (instead of a less aggressive mode). The proposed strategy can reduce memory energy consumption in both sequential and parallel applications. Our strategy has been implemented in an experimental compiler using a polyhedral tool and evaluated using nine array-dominated applications on both a cacheless system and a system with cache memory. Our experimental results indicate that the proposed strategy is very successful in reducing the memory system energy and improves the memory energy by as much as 36.8 percent over a strategy that uses low-power modes without optimizing data access pattern. Our results also show that optimizations that target reducing off-chip memory energy can generate very different results from those that target at improving only cache locality. Victor M. DeLaLuz, Ismail Kadayif, Mahmut T. Kandemir, Ugur Sezer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | Quasidynamic Layout Optimizations for Improving Data LocalityabstractCompiler-directed locality optimization techniques are effective in reducing the number of cycles spent in off-chip memory accesses. Recently, methods have been developed that transform memory layouts of data structures at compile-time to improve spatial locality of nested loops beyond current control-centric (loop nest-based) optimizations. Most of these data-centric transformations use a single static (program-wide) memory layout for each array. A disadvantage of these static layout-based locality enhancement strategies is that they might fail to optimize codes that manipulate arrays, which demand different layouts in different parts of the code. We introduce a new approach, which extends current static layout optimization techniques by associating different memory layouts with the same array in different parts of the code. We call this strategy "quasidynamic layout optimization." In this strategy, the compiler determines memory layouts (for different parts of the code) at compile time, but layout conversions occur at runtime. We show that the possibility of dynamically changing memory layouts during the course of execution adds a new dimension to the data locality optimization problem. Our strategy employs a static layout optimizer module as a building block and, by repeatedly invoking it for different parts of the code, it checks whether runtime layout modifications bring additional benefits beyond static optimization. Our experiments indicate significant improvements in execution time over static layout-based locality enhancing techniques. Ismail Kadayif, Mahmut T. Kandemir |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Compiler-directed scratch pad memory optimization for embedded multiprocessorsabstractThis paper presents a compiler strategy to optimize data accesses in regular array-intensive applications running on embedded multiprocessor environments. Specifically, we propose an optimization algorithm that targets at reducing extra off-chip memory accesses caused by interprocessor communication. This is achieved by increasing the application-wide reuse of data that resides in scratch-pad memories of processors. Our results obtained using four array-intensive image processing applications indicate that exploiting interprocessor data sharing can reduce energy-delay product significantly on a four-processor embedded system. Mahmut T. Kandemir, Ismail Kadayif, Alok N. Choudhary, Ibrahim Kolcu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Generalized Data Transformations for Enhancing Cache Behavior
Victor M. DeLaLuz, Mahmut T. Kandemir, Ismail Kadayif, Ugur Sezer |
DATE | 3 |
| 2003 | An Integrated Approach for Improving Cache Behavior
Gokhan Memik, Mahmut T. Kandemir, Alok N. Choudhary, Ismail Kadayif |
DATE | 4 |
| 2003 | Compiler-Directed Management of Instruction AccessesabstractWe present a compiler-oriented strategy to reduce the memory system energy consumption due to instruction accesses and increase performance by exploiting scratch pad memories. Scratch pad memories (SPMs) are alternatives to conventional cache memories in embedded computing. These small on-chip memories, like caches, provide fast and low-power access to data and instructions; but, they differ from caches in that their contents are managed by software instead of hardware. Our compiler framework keeps the most frequently used instructions in SPM and dynamically changes the contents of the SPM as the (instruction) working set of the application changes. Guilin Chen, Guangyu Chen, Ismail Kadayif, Wei Zhang 0002, Mahmut T. Kandemir, Ibrahim Kolcu, Ugur Sezer |
DSD | 3 |
| 2003 | CCC: Crossbar Connected Caches for Reducing Energy Consumption of On-Chip MultiprocessorsabstractWith shrinking feature size of silicon fabrication technology, architects are putting more and more logic into a single die. While one might opt to use these transistors for building complex single processor based architectures, recent trends indicate a shift towards on-chip multiprocessor systems since they are simpler to implement and can provide better performance. An important problem in on-chip multiprocessors is energy consumption. In particular, on-chip cache structures can be major energy consumers. In this work, we study energy behavior of different cache architectures, and propose a new architecture, where processors share a single, banked cache using crossbar interconnects. Our detailed cycle-accurate simulations show that this cache architecture brings energy benefits ranging from 9% to 26% (over an architecture where each processor has a private cache). Lin Li 0002, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Ismail Kadayif |
DSD | 5 |
| 2003 | An Energy-Oriented Evaluation of Communication Optimizations for Microcensor Networks
Ismail Kadayif, Mahmut T. Kandemir, Alok N. Choudhary, Mustafa Karaköy |
Euro-Par | 1 |
| 2002 | Optimizing inter-nest data localityabstractBy examining data reuse patterns of four array-intensive embedded applications, we found that these codes exhibit a significant amount of inter-nest reuse (i. e., the data reuse that occurs between different nests). While traditional compiler techniques that target array-intensive applications can exploit intra-nest data reuse, there has not been much success in the past in taking advantage of internest data reuse. In this paper, we present a compiler strategy that optimizes inter-nest reuse using loop (iteration space) transformations. Our approach captures the impact of execution of a nest on cache contents using an abstraction called footprint vector. Then, it transforms a given nest such that the new (transformed) access pattern reuses the data left in cache by the previous nest in the code. In optimizing inter-nest locality, our approach also tries to achieve good intra-nest locality. Our simulation results indicate large performance improvements. In particular, inter-nest loop optimization generates competitive results with intra-nest loop and data optimizations. Mahmut T. Kandemir, Ismail Kadayif, Alok N. Choudhary, Joseph Zambreno |
CASES | 2 |
| 2002 | Influence of Loop Optimizations on Energy Consumption of Multi-bank Memory Systems
Mahmut T. Kandemir, Ibrahim Kolcu, Ismail Kadayif |
CC | 3 |
| 2002 | An energy saving strategy based on adaptive loop parallelizationabstractIn this paper, we evaluate an adaptive loop parallelization strategy (i.e., a strategy that allows each loop nest to execute using different number of processors if doing so is beneficial) and measure the potential energy savings when unused processors during execution of a nested loop in a multi-processor on-a-chip (MPoC) are shut down (i.e., placed into a power-down or sleep state). Our results show that shutting down unused processors can lead to as much as 67% energy savings with up to 17% performance loss in a set of array-intensive applications. We also discuss and evaluate a processor pre-activation strategy based on compile-time analysis of nested loops. Based on our experiments, we conclude that an adaptive loop parallelization strategy combined with idle processor shut-down and pre-activation can be very effective in reducing energy consumption without increasing execution time. Ismail Kadayif, Mahmut T. Kandemir, Mustafa Karaköy |
DAC | 1 |
| 2002 | An integer linear programming based approach for parallelizing applications in On-chip multiprocessorsabstractWith energy consumption becoming one of the first-class optimization parameters in computer system design, compilation techniques that consider performance and energy simultaneously are expected to play a central role. In particular, compiling a given application code under performance and energy constraints is becoming an important problem. In this paper, we focus on an on-chip multiprocessor architecture and present a parallelization strategy based on integer linear programming. Given an array-intensive application, our optimization strategy determines the number of processors to be used in executing each nest based on the objective function and additional compilation constraints provided by the user. Our initial experience with this strategy shows that it is very successful in optimizing array-intensive applications on on chip multiprocessors under energy and performance constraints. Ismail Kadayif, Mahmut T. Kandemir, Ugur Sezer |
DAC | 1 |
| 2002 | EAC: A Compiler Framework for High-Level Energy Estimation and OptimizationabstractThis paper presents a novel Energy-Aware Compilation (EAC) framework that can estimate and optimize energy consumption of a given code taking as input the architectural and technological parameters, energy models, and energy/performance constraints,. The framework has been validated using a cycle-accurate architectural-level energy simulator and found to be within 6% error margin while providing significant estimation speedup. The estimation speed of EAC is the key to the number of optimization alternatives that can be explored within a reasonable compilation time. Ismail Kadayif, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam |
DATE | 1 |
| 2002 | Generating physical addresses directly for saving instruction TLB energyabstractPower consumption and power density for the Translation Lookaside Buffer (TLB) are important considerations not only in its design, but can have a consequence on cache design as well. This paper embarks on a new philosophy for reducing the number of accesses to the instruction TLB (iTLB) for power and performance optimizations. The overall idea is to keep a translation currently being used in a register and avoid going to the iTLB as far as possible - until there is a page change. We propose four different approaches for achieving this, and experimentally demonstrate that one of these schemes that uses a combination of compiler and hardware enhancements can reduce iTLB dynamic power by over 85% in most cases. These mechanisms can work with different instruction-cache (iLl) lookup mechanisms and achieve significant iTLB power savings without compromising on performance. Their importance grows with higher iLl miss rates and larger page sizes. They can work very well with large iTLB structures, that can possibly consume more power and take longer to lookup, without the iTLB getting into the common case. Further, we also experimentally demonstrate that they can provide performance savings for virtually-indexed, virtually-tagged iLl caches, and can even make physically-indexed, physically-tagged iLl caches a possible choice for implementation. Ismail Kadayif, Anand Sivasubramaniam, Mahmut T. Kandemir, Gokul B. Kandiraju, Guangyu Chen |
MICRO | 1 |
| 2001 | Dynamic Management of Scratch-Pad Memory SpaceabstractOptimizations aimed at improving the efficiency of on-chip memories are extremely important. We propose a compiler-controlled dynamic on-chip scratch-pad memory (SPM) management framework that uses both loop and data transformations. Experimental results obtained using a generic cost model indicate significant reductions in data transfer activity between SPM and off-chip memory. Mahmut T. Kandemir, J. Ramanujam, Mary Jane Irwin, Narayanan Vijaykrishnan, Ismail Kadayif, Amisha Parikh |
DAC | 5 |
| 2001 | vEC: virtual energy countersabstractEnergy has become a critical issue in processor design, especially in embedded environments. Thus, there is a need for tools, which provide an accurate and fast estimation of energy. In this paper, we present the design and use of a tool, Virtual Energy Counters (vEC), for estimating the energy consumption of user programs. vEC is built on top of the Perfmon user library for the UltraSPARC platform, and provides a user interface, which can be used within user programs to estimate the energy consumption. The energy estimates are provided for those consumed in the data, instruction and extended caches, main memory, address bus, data bus, address pads, and data pads. Ismail Kadayif, T. Chinoda, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin, Anand Sivasubramaniam |
PASTE | 1 |