EDBT 2026 Demo / reviewers in the wild / expert
Peter Petrov
dblp:96/1532
· DBLP profile ↗
35ranked-venue papers
12as first author
0since 2021 · last 2017
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 12 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Memory systems · 48% Processor architecture and microarchitecture · 19% Embedded and real-time systems · 18% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache management
cache partitioning |
0.1 | 2 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.1 | 1 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 |
Energy-efficient computing › memory energy efficiency
low-power cache design |
0.1 | 2 | 2005 | Energy-effcient physically tagged caches for embedded processors with virtual memory · DAC 2005 Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Compilers and program optimization
register allocation |
0.1 | 1 | 2008 | Compiler-driven register re-assignment for register file power-density and temperature reduction · DAC 2008 |
Memory systems
cache coherence |
0.1 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.1 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Memory systems
cache design |
0.1 | 2 | 2004 | Tag compression for low power in dynamically customizable embedded processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Processor architecture and microarchitecture › multithreading
context switching |
0.1 | 1 | 2006 | Rapid and low-cost context-switch through embedded processor customization for real-time and control applications · DAC 2006 |
Embedded and real-time systems
real-time operating systems |
0.1 | 1 | 2006 | Rapid and low-cost context-switch through embedded processor customization for real-time and control applications · DAC 2006 |
Embedded and real-time systems
real-time scheduling |
0.1 | 1 | 2006 | Rapid and low-cost context-switch through embedded processor customization for real-time and control applications · DAC 2006 |
Memory systems
cache |
0.1 | 1 | 2005 | Energy-effcient physically tagged caches for embedded processors with virtual memory · DAC 2005 |
Energy-efficient computing › low-power design
power optimization |
0.0 | 1 | 2004 | Tag compression for low power in dynamically customizable embedded processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 |
Processor architecture and microarchitecture › branch handling
branch folding |
0.0 | 1 | 2001 | Speeding Up Control-Dominated Applications through Microarchitectural Customizations in Embedded Processors · DAC 2001 |
Memory systems › cache › cache organization
configurable cache |
0.0 | 1 | 2001 | Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Memory systems › cache design
partitioned cache |
0.0 | 1 | 2001 | Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Compilers and program optimization
program transformation |
0.0 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Hardware reliability and fault tolerance › reliability analysis
thermal reliability |
0.0 | 1 | 2008 | Compiler-driven register re-assignment for register file power-density and temperature reduction · DAC 2008 |
Embedded and real-time systems › embedded processor
embedded processor design |
0.0 | 2 | 2004 | Tag compression for low power in dynamically customizable embedded processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 Performance and power effectiveness in embedded processors customizable partitioned caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2001 |
Embedded and real-time systems › embedded processor
embedded processor customization |
0.0 | 1 | 2006 | Rapid and low-cost context-switch through embedded processor customization for real-time and control applications · DAC 2006 |
Embedded and real-time systems
embedded processor |
0.0 | 1 | 2001 | Speeding Up Control-Dominated Applications through Microarchitectural Customizations in Embedded Processors · DAC 2001 |
Distributed systems › fault tolerance
checkpointing |
0.0 | 1 | 1991 | On the Optimal Total Processing Time Using Checkpoints · IEEE Trans. Software Eng. 1991 |
Methods — techniques the papers use, named apart from their topics
cross-layer customization · 0.2compiler-driven code transformation · 0.2algorithmic heuristic · 0.2NP-hardness proof · 0.2application-driven partitioning · 0.1register liveness analysis · 0.1compile-time switch point identification · 0.1static code/data layout analysis · 0.0VLSI implementation · 0.0application profiling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Validity of Automated Inferences in Mapping of Anatomical Ontologies
Milko Krachunov, Peter Petrov, Maria Nisheva-Pavlova, Dimitar Vassilev |
ISMIS | 2 |
| 2015 | Preliminary field results of soil moisture from Kuwait desert as a core validation site of SMAP satelliteabstractField work was conducted in two SMAP 36×36 km grid cells, B and D, located in the west and north of Kuwait, respectively. The in-situ gravimetric sampling field work activity in Grid Cell D indicates a variation of volumetric soil moisture from 0.17 m3m-3in January, 2014 to 0.015 m3m-3in June, 2014. Field work in Grid Cell B indicates a variation from 0.0352 m3m-3in December 2014 to 0.0168 m3m-3in May 2015. Soil roughness was estimated in grid cell D using a pin profilometer and was found to vary from 0.2 to 0.7 RMS with a correlation length ranging from 91 cm to 93 cm. The first weather station was installed in grid cell B in April 2015. Hala Khalid AlJassar, Peter Petrov, Dara Entekhabi, Marouane Temimi, Nevil Kodiyan, Mohamed Shuaib Ansari |
IGARSS | 2 |
| 2011 | Dynamically Adaptive I-Cache Partitioning for Energy-Efficient Embedded MultitaskingabstractThe ever increasing importance of battery-powered devices coupled with high performance requirements and shrinking process geometries have further exacerbated the problem of energy efficiency in modern embedded systems. The cache memories are a major contributor to the system power consumption, and as such have been a primary target for energy reduction techniques. Recent advances in configurable cache architectures have enabled an entirely new set of approaches for application-driven energy- and cost-efficient cache resource utilization. We propose a run-time and adaptive instruction cache partitioning methodology, which leverages configurable cache architectures to achieve an energy- and performance-conscious adaptive mapping of instruction cache resources to tasks in dynamic multi-task workloads sharing a processor core trough preemptive multitasking. Sizable leakage and dynamic power reductions are achieved with only a negligible and system-controlled performance impact. The methodology assumes no prior information regarding the dynamics and the structure of the workload. As the proposed dynamic cache partitioning alleviates the adverse effects of cache interference, performance is maintained very close to the baseline case, while achieving 50%-80% reductions in dynamic and leakage power for the on-chip instruction cache memory. Mathew Paul, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platformsabstractWe present a methodology for off-chip memory bandwidth minimization through application-driven L2 cache partitioning in multi-core systems. A major challenge with multi-core system design is the widening gap between the memory demand generated by the processor cores and the limited off-chip memory bandwidth and memory service speed. This severely restricts the number of cores that can be integrated into a multi-core system and the parallelism that can be actually achieved and efficiently exploited for not only memory demanding applications, but also for workloads consisting of many tasks utilizing a large number of cores and thus exceeding the available off-chip bandwidth. Chenjie Yu, Peter Petrov |
DAC | 2 |
| 2010 | Context-aware TLB preloading for interference reduction in embedded multi-tasked systemsabstractRapid system responsiveness and execution time predictability are of significant importance for a large class of real-time embedded systems. Multi-tasking leads to interference in the shared processor resources such as caches and TLBs, which in turn results in not only deteriorated performance but also, and for some applications even more importantly, highly suboptimal worst-case execution time (WCET) estimates due to the interference unpredictability. We present a methodology for task-aware D-TLB interference reduction and preloading through an application-specific task's state introspection at context-switch time for embedded multitasking. The proposed technique addresses the problem through a synergistic cooperation between the compiler, for an application-specific analysis of the task's context, and the OS, for a run-time introspection of the context and an efficient identification of TLB entries of current (live) and of near-future usage. Ilya Chukhman, Peter Petrov |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Cache partitioning for energy-efficient and interference-free embedded multitaskingabstractWe propose a technique that leverages configurable data caches to address the problem of energy inefficiency and intertask interference in multitasking embedded systems. Data caches are often necessary to provide the required memory bandwidth. However, caches introduce two important problems for embedded systems. Caches contribute to a significant amount of power as they typically occupy a large part of the chip and are accessed frequently. In nanometer technologies, such large structures contribute significantly to the total leakage power as well. Additionally, cache outcomes in multitasking environments are notoriously difficult to predict, if not impossible, thus resulting in poor real-time guarantees. We study the effect of multiprogramming workloads on the data cache in a preemptive multitasking environment, and propose a technique which leverages configurable cache architectures to not only eliminate intertask cache interference, but also to significantly reduce both dynamic and leakage power. By mapping tasks to different cache partitions, interference is completely eliminated. Dynamic and leakage power are significantly reduced as only a subset of the cache is active at any moment. We introduce a profile-based, off-line algorithm, which identifies a beneficial cache partitioning. The OS configures the data cache during context-switch by activating the corresponding partition. Our experiments on a large set of multitasking benchmarks demonstrate that our technique not only efficiently eliminates intertask interference, but also significantly reduces both dynamic and leakage power. Rakesh Reddy, Peter Petrov |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | Energy- and Performance-Efficient Communication Framework for Embedded MPSoCs through Application-Driven Release ConsistencyabstractWe present a framework for performance-, bandwidth-, and energy-efficient intercore communication in embedded MultiProcessor Systems-on-a-Chip (MPSoC). The methodology seamlessly integrates compiler, operating system, and hardware support to achieve a low-cost communication between synchronized producers and consumers. The technique is especially beneficial for data-streaming applications exploiting pipeline parallelism with computational phases mapped to separate cores. Code transformations utilizing a simple ISA support ensure that producer writes are propagated to consumers with a single interconnect transaction per cache block just prior to the producer exiting its synchronization region. Furthermore, in order to completely eliminate misses to shared data caused by interference with private data and also to minimize the cache energy, we integrate to the proposed framework a cache way partitioning policy based on a simple cache configurability support, which isolates the shared buffers from other cache traffic. This mechanism results in significant power savings since only a subset of the cache ways needs to be looked up for each cache access. The end result of the proposed framework is a single communication transaction per shared cache block between a producer and a consumer with no coherence misses on the consumer caches. Our experiments demonstrate significant reductions in interconnect traffic, cache misses, and energy for a set of multiprocessor benchmarks. Chenjie Yu, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Low-Cost and Energy-Efficient Distributed Synchronization for Embedded MultiprocessorsabstractWe present a framework for a distributed and lowcost implementation of synchronization mechanisms for embedded shared-memory multiprocessors. The proposed architecture effectively implements the queued-lock semantics in a completely decentralized manner through low-cost and distributed synchronization controllers performing distributed synchronization management protocols. The proposed approach achieves three major benefits. First, it completely eliminates the overwhelming bus contention traffic when multiple cores compete for a synchronization variable. Second, it exhibits extremely low best-case latency of lock acquisition (with zero bus transactions). Third, the approach enables multiple venues for high energy efficiency as the local synchronization controllers can efficiently determine, without any bus transactions or local cache spinning, the exact timing of when a lock is made available to or a barrier enabled at the local processor. It becomes possible for the system software or the thread library to employ various low-power policies. Chenjie Yu, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Cross-layer customization for rapid and low-cost task preemption in multitasked embedded systemsabstractPreemptive multitasking is widely used in many low-cost and real-time embedded applications for its superior hardware utilization. The frequent and asynchronous context switches, however, require the preservation and restoration of the task state, thus resulting in a large number of memory transfer instructions. As a consequence, task responsiveness and application throughput can be significantly deteriorated. To address this problem we propose a cross-layer customization framework which through the close cooperation of compiler, OS, and hardware architecture achieves rapid and low-cost task switch. Application information extracted during compile-time regarding state liveness is exploited in order to preserve a minimal amount of task state on task preemption. We introduce two complementary techniques to implement the application-aware state preservation. The first technique utilizes compiler-generated custom routines which preserve/restore an extremely small live context at judiciously selected points in the application code. The second technique requires more sophisticated hardware support. It employs an OS-controlled register file mapping to achieve a rapid context switch. By mapping a small fraction of the register file in a single clock cycle, a context switch is achieved requiring no memory transfers for the majority of cases to preserve/restore the live state. The effect of aggressively replicated register files, where each task is given its own replica, is achieved with the hardware cost of only adding from 25% to 50% extra physical registers. Through the utilization of these novel mechanisms, a significant improvement on task response time is achieved as the context-switch cost is minimized. Xiangrong Zhou, Peter Petrov |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2009 | Temperature-aware register reallocation for register file power-density minimizationabstractIncreased chip temperature has been known to cause severe reliability problems and to significantly increase leakage power. The register file has been previously shown to exhibit the highest temperature compared to all other hardware components in a modern high-end embedded processor, which makes it particularly susceptible to faults and elevated leakage power. We show that this is mostly due to the highly clustered register file accesses where a set of few registers physically placed close to each other are accessed with very high frequency. We propose compile-time temperature-aware register reallocation methodologies for breaking such groups of registers and to uniformly distribute the accesses to the register file. This is achieved with no performance and no hardware overheads . We show that the underlying problem is NP-hard, and subsequently introduce and evaluate two efficient algorithmic heuristics. Our extensive experimental study demonstrates the efficiency of the proposed methodology. Xiangrong Zhou, Chenjie Yu, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2009 | Low-Power Snoop Architecture for Synchronized Producer-Consumer Embedded MultiprocessingabstractWe introduce a cross-layer customization methodology where application knowledge regarding data sharing in producer-consumer relationships is used in order to aggressively eliminate unnecessary and predictable snoop-induced cache lookups even for references to shared data, thus, achieving significant power reductions with minimal hardware cost. The technique exploits application-specific information regarding the exactproducer-consumerrelationshipsbetween tasks as well as information regardingtheprecisetimingofsynchronizedaccessesto shared memory buffers by their corresponding producers and/or consumers. Snoop-induced cache lookups for accesses to the shared data are eliminated when it is ensured that such lookups will not result in extra knowledge regarding the cache state in respect to the other caches and the memory. Our experiments show average power reductions of more than 80% compared to a general-purpose snoop protocol. Chenjie Yu, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessorsabstractWe present a cross-layer customization methodology for latency and bandwidth efficient inter-core communication in embedded multiprocessors. The methodology integrates compiler, operating system, and hardware support to achieve a bandwidth efficient, snoop-free, and coherence cache miss-free shared memory communication between synchronized producer and consumers cores. A compiler-driven code transformation is introduced that utilizes a simple ISA support in the form of a special write-through store instruction. It ensures that producer writes are propagated to the consumers with a single bus transaction per cache block when the producer performs the last write to that cache line before exiting its synchronization region. Information regarding the shared buffers involved in the communications is captured by the OS and provided to the cores with the purpose of filtering bus traffic and performing remote updates when necessary. The end result of the proposed methodology is a single bus transaction per shared cache block and snoop-free communication between a producer and a set of consumers with no intervening coherence misses on the consumer caches. Our experiments demonstrate the significant reductions in both bus traffic and cache misses for a set of multiprocessor benchmarks. Chenjie Yu, Peter Petrov |
DAC | 2 |
| 2008 | Compiler-driven register re-assignment for register file power-density and temperature reductionabstractTemperature hot-spots have been known to cause severe reliability problems and to significantly increase leakage power. The register file has been previously shown to exhibit the highest temperature compared to all other hardware components in a modern high-end embedded processor, which makes it particularly susceptible to faults and elevated leakage power. We show that this is mostly due to the highly clustered register file accesses where a set of few registers physically placed close to each other are accessed with very high frequency. In this paper we propose a compiler-based register reassignment methodology, which purpose is to break such groups of registers and to uniformly distribute the accesses to the register file. This is achieved with no performance and no hardware overheads. We show that the underlying problem is NP-hard, and subsequently introduce an efficient algorithmic heuristic. Xiangrong Zhou, Chenjie Yu, Peter Petrov |
DAC | 3 |
| 2008 | Direct address translation for virtual memory in energy-efficient embedded systemsabstractThis article presents a methodology for virtual memory support in energy-efficient embedded systems. A holistic approach is proposed, where the combined efforts of compiler, operating system, and hardware architecture achieve a significant system power reductions. The application information extracted and analyzed by the compiler is utilized dynamically by the microarchitecture and the operating system to perform energy-efficient and, for many memory references, time-deterministic address translations. We demonstrate that by using application information regarding virtual memory layout, an efficient and conflict-free translation process can be implemented through the utilization of a small hardware direct translation table (DTT) accessed in an application-specific manner. The set of virtual pages is partitioned into groups, such that for each group only a few of the least significant bits are used as an index to obtain the physical page number. We outline an efficient compile-time algorithm for identifying these groups and allocate their translation entries optimally into the DTT. The introduced hardware is minimal in terms of area, performance, and power overhead, while offering the flexibility of software programmability. This is achieved through a small set of registers and tables, which are made software accessible. We have quantitatively evaluated the proposed methodology on a number of embedded applications, including voice, image, and video processing. Xiangrong Zhou, Peter Petrov |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | Heterogeneously tagged caches for low-power embedded systems with virtual memory supportabstractAn energy-efficient data cache organization for embedded processors with virtual memory is proposed. Application knowledge regarding memory references is used to eliminate most tag translations. A novel tagging scheme is introduced, where both virtual and physical tags coexist. Physical tags and special handling of superset index bits are only used for references to shared regions in order to avoid cache inconsistency. By eliminating the need for most address translations on cache access, a significant power reduction is achieved. We outline an efficient hardware architecture, where the application information is captured in a reprogrammable way and the cache is minimally modified. Xiangrong Zhou, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | Application-aware snoop filtering for low-power cache coherence in embedded multiprocessorsabstractMaintaining local caches coherently in shared-memory multiprocessors results in significant power consumption. The customization methodology we propose exploits the fact that in embedded systems, important knowledge is available to the system designers regarding memory sharing between tasks. We demonstrate how the snoop-induced cache probings can be significantly reduced by identifying and exploiting in a deterministic way the shared memory regions between the processors. Snoop activity is enabled only for the accesses referring to known shared regions. The hardware support is not only cost efficient, but also software programmable, which allows for reprogrammability and customization across different tasks and applications. Xiangrong Zhou, Chenjie Yu, Alokika Dash, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2008 | Guest Editorial Special Section on Application Specific ProcessorsabstractThis special section on application specific processors consists of eight articles, four of which are extended versions of papers presented at the Fifth Workshop on Application Specific Processors (WASP) held in 2007. Paolo Ienne, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Eliminating inter-process cache interference through cache reconfigurability for real-time and low-power embedded multi-tasking systemsabstractWe propose a technique which leverages configurable data caches to address the problem of cache interference in multitasking embedded systems. Data caches are often necessary to provide the required memory bandwidth. However, caches introduce two important problems for embedded systems. Cache outcomes in multi-tasking environments are notoriously difficult to predict, if not impossible, thus resulting in poor real-time guarantees. Additionally, caches contribute to a significant amount of power. These issues are key factors for many embedded systems. We study the effect of multiple tasks on the data cache, and propose a technique which leverages configurable cache architectures to eliminate inter-task cache interference. By mapping tasks to different cache partitions, interference is completely eliminated with only a minimal impact on performance. Furthermore, dynamic and leakage power are significantly reduced as only a subset of the cache is active at any moment. We introduce a profile-based, static analysis algorithm, which identifies a beneficial cache partitioning. The OS configures the data cache during context-switch by activating the corresponding partition.Our experiments on a large set of multitasking benchmarks demonstrate that our technique not only efficiently eliminates inter-task interference but also significantly reduces both dynamic and leakage power. Rakesh Reddy, Peter Petrov |
CASES | 2 |
| 2006 | Rapid and low-cost context-switch through embedded processor customization for real-time and control applicationsabstractIn this paper, we present a methodology for low-cost and rapid context switch for multithreaded embedded processors with real-time guarantees. Context-switch, which involves saving and restoring the thread state, has constituted not only a large performance overhead for many multithreaded embedded systems, but also an obstacle creating a significant delay in the response time for many time-critical control applications. The proposed technique exploits application information extracted during compile time to make sure that only a minimal amount of thread state is saved and subsequently restored on preemption. The register liveness within the application inner loops is analyzed and a few points, referred to as switch points, are identified where the program has minimal number of live registers. At run-time the preemption point is deferred to a switch point and the Real-Time Operating System (RTOS) kernel invokes a switch point specific code generated by the compiler to save and restore the thread state in a custom fashion. Through the utilization of these novel mechanisms, a drastic improvement on both performance and response time is achieved. The presented experimental results demonstrate the effectiveness of the proposed technique on a number of widely-used computational kernels and embedded applications. Xiangrong Zhou, Peter Petrov |
DAC | 2 |
| 2006 | Energy-Efficient Cache Coherence for Embedded Multi-Processor Systems through Application-Driven Snoop FilteringabstractMaintaining local caches coherent in bus-based multiprocessor systems results in significantly elevated power consumption, as the bus snooping protocols require local cache lookups for each memory reference placed on the common bus. Such a conservative approach is warranted in general-purpose systems, where no prior knowledge regarding the communication structure between threads or processes is available. In such a general-purpose context the assumption is that each memory request is potentially a reference to a shared memory region, which may result in cache inconsistency, if no correcting activities are undertaken. The approach we propose exploits the fact that in embedded systems, important knowledge is available to the system designers regarding communication activities between tasks allocated to the different processor nodes. We demonstrate how the snoop-related cache probing activity can be drastically reduced by identifying in a deterministic way all the shared memory regions and the communication patterns between the processor nodes. Cache snoop activity is enabled only for the fraction of the bus transactions, which refer to locations belonging to known shared memory regions for each processor node; for the remaining larger part of memory references known to be of no relation to the given processor node, snoop probings in the local cache are completely disabled, thus saving a large amount of power. The experiments which we have performed on a number of important applications demonstrate the effectiveness of the proposed approach Alokika Dash, Peter Petrov |
DSD | 2 |
| 2006 | Low-power cache organization through selective tag translation for embedded processors with virtual memory supportabstractIn this paper we present a novel cache architecture for energy-efficient data caches in embedded processors with virtual memory. Application knowledge regarding the nature of memory references is used to eliminate tag address translations for most of the cache accesses. We introduce a novel cache tagging scheme, where both virtual and physical tags co-exist in the cache tag arrays. Physical tags and special handling for the super-set cache index bits are used for references to shared data regions in order to avoid cache consistency problems. By eliminating the need for address translation on cache access for the majority of references, a significant power reduction is achieved. We outline an efficient hardware architecture for the proposed approach, where the application information is captured in a reprogrammable way and the cache architecture is minimally modified. Our experimental results show energy reductions for the address translation hardware in the range of 90%, while the reduction for the entire cache architecture is within the range of 25%-30%. Xiangrong Zhou, Peter Petrov |
ACM Great Lakes Symposium on VLSI | 2 |
| 2005 | Energy-effcient physically tagged caches for embedded processors with virtual memoryabstractIn this paper we present a low-power tag organization for physically tagged caches in embedded processors with virtual memory support. An exceedingly small subset of tag bits is identified for each application hot-spot so that only these tag bits are used for cache access with no performance sacrifice as they provide complete address resolution. The minimal subset of physical tag bits, i.e. the compressed tag, is dynamically updated following the changes in the physical address space of the application. Special support from the operating system (OS) is introduced in order to maintain the compressed tag during program execution. The compressed tag is updated by the OS to match the current set of physical memory pages allocated to the application. We have proposed efficient algorithms that are incorporated within the memory allocator and the dynamic linker in order to achieve dynamic update of the compressed tags in the cases where the mapping between virtual and physical addresses is modified; such cases include memory allocation/deallocation and swapping physical pages on the secondary memory storage. The only hardware support needed within the l/D-caches is the support for disabling bitlines of the tag arrays. An extensive set of experimental results demonstrates the efficacy of the proposed approach. Peter Petrov, Daniel Tracy, Alex Orailoglu |
DAC | 1 |
| 2005 | A reprogrammable customization framework for efficient branch resolution in embedded processorsabstractWe present a customization framework for embedded processors which employs the utilization of application-specific information, thus specializing the processor's microarchitecture to the application needs. The increased processor utilization leads to a low-cost system implementation with no sacrifice in performance requirements and to reduced custom hardware in a typical SOC. We illustrate these ideas through the branch resolution problem, known to impose severe performance degradation on control-dominated embedded applications. A customization approach for early branch resolution and subsequent folding is presented. The application-specific information is captured by the microarchitecture through a low-cost reprogrammable hardware, thus attaining the twin benefits of processor standardization and application-specific customization. Experimental results show that for a representative set of control-dominated applications a reduction in the range of 3--22% in processor cycles can be achieved, thus extending the scope of low-cost embedded processors in complex codesigns for control intensive systems. Peter Petrov, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2004 | Tag compression for low power in dynamically customizable embedded processorsabstractWe present a methodology for power reduction by instruction/data cache-tag compression for low-power embedded processors. By statically analyzing the code/data memory layouts for the application hot spots, a variety of proposed schemes for effective tag-size reduction can be employed for power minimization in instruction and data caches. The schemes rely on significantly reducing the number of tag bits stored in the tag arrays for cache-conflict identification, thus considerably decreasing the number of active bitlines, sense amps, and comparator cells. We present a set of tag compression techniques and evaluate each of them separately in terms of efficiency and required hardware support. A detailed very large scale integrated implementation has been performed and a number of experimental results on a set of embedded applications is reported for each technique. Energy dissipation decreases of up to 95% can be observed for the tag arrays, implying significant energy reductions in the range of 50% when amortized across the overall cache subsystem. Peter Petrov, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Low-power instruction bus encoding for embedded processorsabstractAbstract—This paper presents a low-power encoding framework for embedded processor instruction buses. The encoder is capable of adjusting its encoding not only to suit applications but furthermore to suit different aspects of particular program execution. It achieves this by exploiting application-specific knowledge regarding program hot-spots, and thus identifies efficient instruction transformations so as to minimize the bit transitions on the instruction bus lines. Not only is the switching activity on the individual bus lines considered but so is the coupling activity across adjacent bus lines, a foremost contributor to the total power dissipation in the case of nanometer technologies. Low-power codes are utilized in a reprogrammable application specific manner. The restriction to two well-selected classes of simply computable, functional transformations delivers significant storage benefits and ease of reprogrammability, in the process obtaining significant power savings. The microarchitectural support enables reprogrammability of the encoding transformations in order to track code particularities effectively. Such reprogrammability is achieved by utilizing small tables that store relevant application information. The few transformations that result in optimal power reductions for each application hot-spot are selected by utilizing short indices stored into a table, which is accessed only once at the beginning of the transformed bit sequence. Extensive experimental results show significant power reductions ranging up to 80 % for switching activity on bus lines and up to 70 % when bus coupling effects are also considered. I. Peter Petrov, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | Power Efficiency through Application-Specific Instruction Memory Transformations
Peter Petrov, Alex Orailoglu |
DATE | 1 |
| 2003 | Customizable Embedded Processor ArchitecturesabstractIn this paper, we present a framework for dynamic application customization for high-performance and low-power embedded processors. The proposed architecture is capable of utilizing application information to boost the performance and lower the power consumption of the most important microarchitectural components such as instruction/data caches and the memory subsystem. We present a design framework, including CAD support infrastructure and reprogrammable hardware support, for a dynamically customizable microarchitecture. We outline the underlying algorithms for compile-time extraction of the utilized application properties and we present the architectural principles of the hardware support. Extensive experimental results confirm the efficacy of this novel embedded processor architecture. Peter Petrov, Alex Orailoglu |
DSD | 1 |
| 2003 | Low-power Branch Target Buffer for Application-Specific Embedded ProcessorsabstractIn this paper we present a methodology for a low-power branch identification mechanism, which enables the design of extremely power efficient branch predictors for embedded processors. The proposed technique utilizes application-specific information regarding the control-flow structure of the program major loops. Such information is used to completely eliminate the power hungry branch target buffer (BTB) lookups which normally occur at every execution cycle. Exact application knowledge regarding the control-flow structure of the program obviates the power expensive BTB operations, thus enabling the utilization of contemporary branch predictors in high-end, yet power-sensitive embedded processors. The utilization of exact application knowledge results not only in the complete elimination of the power hungry BTB structure but also in a perfect branch and target address identification. Cost-efficient and programmable hardware architecture for capturing the control-flow structure of the program is presented thereafter. The hardware complexity of the proposed architecture is carefully analyzed in terms of power, performance and area overhead. The proposed technique delivers power reductions in excess of 90% for a set of embedded benchmarks. Peter Petrov, Alex Orailoglu |
DSD | 1 |
| 2003 | Compiler-Based Register Name Adjustment for Low-Power Embedded Processors
Peter Petrov, Alex Orailoglu |
ICCAD | 1 |
| 2003 | Virtual Page Tag Reduction for Low-power TLBsabstractWe present a methodology for a power-optimized, software-controlled translation lookaside buffer (TLB) organization. A highly reduced number of virtual page number (VPN) bits sufficient to perform physical address translation is efficiently identified and used when performing TLB lookups, delivering significant power reductions. Information regarding the virtual address space of the program code and data provided by the compiler is augmented with information regarding the dynamically linked libraries and data allocated run-time by the loader, the dynamic linker, and the memory manager. The hardware support needed is constrained to disabling bitlines of the tag arrays associated to the 1-TLB and the D-TLB. Algorithms for identifying the reduced VPNs for power optimized TLB operations together with the required OS support are presented. Peter Petrov, Alex Orailoglu |
ICCD | 1 |
| 2002 | Power Efficient Embedded Processor Ip's through Application-Specific Tag Compression in Data CachesabstractIn this paper, we present a methodology for power minimization by data cache tag compression. The set of tags being accessed by the major application loops is analyzed statically during compile time and an efficient and optimal compression scheme is proposed Only a very limited number of tag bits are stored in the tag array for cache conflict identification, thus achieving a significant reduction in the number of active bitlines, sense amps, and comparator cells. The underlying hardware support for dynamically compressing the tags consists of a highly cost and power efficient programmable encoder which lies outside the cache access path, thus not affecting the processor cycle time. A detailed VLSI implementation has been performed and a number of experimental results on a set of embedded applications and numerical kernels is reported Energy dissipation decreases of up to 95% can be observed for the tag arrays, while significant energy reductions in the range of 10%-50% are observed when amortized across the overall cache subsystem. Peter Petrov, Alex Orailoglu |
DATE | 1 |
| 2001 | Faults in Processor Control Subsystems: Testing Correctness and Performance Faults in the Data Prefetching UnitabstractThe processor control subsystems have for a long time been recognized as a bottleneck in the process of achieving complete fault coverage through various functional test propagation approaches. The difficult-to-test corner cases are further accentuated in fault-resilient control subsystems as no functional effect is incurred as a result of the fault, even though performance suffers. We investigate the construction of software programs, capable of providing full fault coverage at minimal hardware cost, for one such fault resilient subsystem in processor architecture: the data prefetching unit. Experimental results confirm the efficacy of the proposed method. Sobeeh Almukhaizim, Peter Petrov, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2001 | Speeding Up Control-Dominated Applications through Microarchitectural Customizations in Embedded ProcessorsabstractWe present a methodology for microarchitectural customization of embedded processors by exploiting application information, thus attaining the twin benefits of processor standardization and application-specific customization. Such powerful techniques enable increased application fragments to be placed on the processor, with no sacrifice in system requirements, thus reducing the custom hardware and the concomitant area requirements in SOCs. We illustrate these ideas through the branch resolution problem, known to impose severe performance degradation on control-dominated embedded applications. A low-cost late customizable hardware that uses application information to fold out a set of frequently executed branches is described. Experimental results show that for a representative set of control dominated applications a reduction in the range of 7%-22% in processor cycles can be achieved, thus extending the scope of low-cost embedded processors in complex co-designs for control intensive systems. Peter Petrov, Alex Orailoglu |
DAC | 1 |
| 2001 | Performance and power effectiveness in embedded processors customizable partitioned cachesabstractThis paper explores an application-specific customization technique for the data cache, one of the foremost area/power consuming and performance determining microarchitectural features of modern embedded processors. The automated methodology for. customizing the processor microarchitecture that we propose results in increased performance, reduced power consumption and improved determinism of critical system parts while the fixed design ensures processor standardization. The resulting improvements help to enlarge the significant role of embedded processors in modern hardware-software codesign techniques by leading to increased processor utilization and reduced hardware cost. A novel methodology for static analysis and a microarchitecturally field-reprogrammable implementation of a customizable cache controller that implements a partitioned cache structure is proposed. Partitioning the load/store instructions eliminates cache interference; hence, precise knowledge about the hit/miss behavior of the references within each partition becomes available, resulting in significant reduction in tag reads and comparisons. Moreover, eliminating cache interference naturally leads to a significant reduction in the miss rate. The paper presents an algorithm for defining cache partitions, hardware support for customizable cache partitions, and a set of experimental results. The experimental results indicate significant improvements in both power consumption and miss rate. Peter Petrov, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1991 | On the Optimal Total Processing Time Using CheckpointsabstractThe authors investigate the problem of optimizing the expected blocking time duration by providing a schedule of checkpoints during the required job processing time. They give a general approach for determining the optimal checkpoint schedule and derive some cases when the optimal checkpointing is uniform. The model has applications in unreliable computing systems, multiclient computer service, data transmissions, etc.> Boyan Dimitrov, Zohel Khalil, Nikolai Kolev, Peter Petrov |
IEEE Trans. Software Eng. | 4 |