Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mirko Loghi

dblp:01/292 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
0since 2021 · last 2018
0000-0001-7876-3612ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 12 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Energy-efficient computing · 35% Memory systems · 26% Hardware reliability and fault tolerance · 20%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
leakage power reduction
0.322014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking · IEEE Trans. Computers 2010
Energy-efficient computing
power management
0.322014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking · IEEE Trans. Computers 2010
Hardware reliability and fault tolerance
aging
0.212014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Memory systems
cache
0.212014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Memory systems › cache design
cache indexing
0.212014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Hardware reliability and fault tolerance › aging › transistor aging
negative bias temperature instability
0.212014
Dynamic Indexing: Leakage-Aging Co-Optimization for Caches · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Electronic design automation › high-level synthesis › memory synthesis
memory partitioning
0.112010
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking · IEEE Trans. Computers 2010
Memory systems › on-chip memory
scratchpad memory
0.112010
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking · IEEE Trans. Computers 2010
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Energy-efficient computing
power-performance tradeoff
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Parallel and multicore computing
programming models
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Parallel and multicore computing › parallel programming models › hybrid programming models
shared memory and message passing
0.112007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007
Embedded and real-time systems › embedded software › embedded operating systems
embedded memory management
0.012010
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking · IEEE Trans. Computers 2010
Parallel and multicore computing
parallel programming models
0.012007
Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support · IEEE Trans. Computers 2007

Methods — techniques the papers use, named apart from their topics

dynamic indexing · 0.2power state modeling · 0.1exhaustive search · 0.1hardware-software tuning · 0.1comparative analysis · 0.1
YearPublicationVenuePosition
2018 A Portable 3-D Imaging FMCW MIMO Radar Demonstrator With a $24\times 24$ Antenna Array for Medium-Range Applications
abstract
Multiple-input multiple-output (MIMO) radars have been shown to improve target detection for surveillance applications thanks to their proven high-performance properties. In this paper, the design, implementation, and results of a complete 3-D imaging frequency-modulated continuous-wave MIMO radar demonstrator are presented. The radar sensor working frequency range spans between 16 and 17 GHz, and the proposed solution is based on a 24-transmitter and 24-receiver MIMO radar architecture, implemented by time-division multiplexing of the transmit signals. A modular approach based on conventional low-cost printed circuit boards is used for the transmit and receive systems. Using digital beamforming algorithms and radar processing techniques on the received signals, a high-resolution 3-D sensing of the range, azimuth, and elevation can be calculated. With the current antenna configuration, an angular resolution of 2.9° can be reached. Furthermore, by taking advantage of the 1-GHz bandwidth of the system, a range resolution of 0.5 m is achieved. The radio-frequency front-end, digital system and radar signal processing units are here presented. The medium-range surveillance potential and the high-resolution capabilities of the MIMO radar are proved with results in the form of radar images captured from the field measurements.
Alexander Rudolf Ganis, Enric Miralles Navarro, Bernhard Schoenlinner, Ulrich Prechtel, Askold Meusling, Christoph Heller, Thomas Spreng, Jan Mietzner, Christian Krimmer, Babette Haeberle, Steffen Lutz, Mirko Loghi, Angel Belenguer, Héctor Esteban, Volker Ziegler
IEEE Trans. Geosci. Remote. Sens.12
2014 Dynamic Indexing: Leakage-Aging Co-Optimization for Caches
abstract
Traditional implementations of low-power states based on voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused by negative bias temperature instability (NBTI), which can be explained in terms of the intuitive correlation between the idleness and the reduced workload of a system. Such a joint benefit has been exploited only partially because of the different nature of energy and aging as cost functions: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually result in limited aging reductions. In this paper, we address this problem in the context of power-managed caches, which represent a critical target for NBTI-reduced aging: given their symmetric structure, SRAM structures are, in particular, sensitive to NBTI effects because they cannot take advantage of the value-dependent recovery typical of NBTI. We propose a strategy called dynamic indexing, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the various power managed units (e.g., lines). This distribution allows fully using the leakage optimization potential and extending the lifetime of a cache. We explore various alternatives, in particular different granularities of the power managed units as well as different reindexing functions. Experimental analysis shows that it is possible to simultaneously reduce leakage power and aging in caches, with minimal power consumption overhead.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2014 Energy/Lifetime Cooptimization by Cache Partitioning With Graceful Performance Degradation
abstract
Aging of transistors can adversely impact the long-term reliability of devices in subnanometric technologies. Without any countermeasure, the first component that becomes unreliable will determine the life span of an entire device. The effect is more susceptible in memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. In this paper, we propose a reliability management technique based on the idea of cache partitioning, which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various subblocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. A coarse-grain implementation of this approach, with the use of a smart aging-driven partitioning algorithm, provides a lifetime extension of more than 2× . On the other hand, a fine-grain strategy with a single cache line as a unit of power management, stretch the lifetime to its maximum limits with an addition of small hardware overhead.
Haroon Mahmood, Mirko Loghi, Massimo Poncino, Enrico Macii
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Application-specific memory partitioning for joint energy and lifetime optimization
abstract
Power management of caches based on turning idle cache lines into a low-energy state is also beneficial for the aging effects caused by Negative Bias Temperature Instability (NBTI), provided that idleness is correctly exploited; unlike energy, aging, being a measure of delay, is in fact a worst-case metric.
Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii
DATE3
2012 Energy-optimal caches with guaranteed lifetime
abstract
This work addresses the aging of the memory sub-system due to NBTI (Negative Bias Temperature Instability) in systems that have to provide a guaranteed level of service, and specifically, a guaranteed lifetime.
Mirko Loghi, Haroon Mahmood, Andrea Calimera, Massimo Poncino, Enrico Macii
ISLPED1
2012 Aging-aware caches with graceful degradation of performance
abstract
Aging of transistors can substantially shorten the lifetime of devices in sub-nanometric technologies. Without any countermeasure, the first component which becomes unreliable will determine the life span of an entire device. This problem is even more relevant for memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. Traditional implementation of power management by turning idle cache lines into a low-energy state can also mitigate the aging effects caused by Negative Bias Temperature Instability (NBTI) provided that idleness is correctly exploited. In this work, we propose a cache structure which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various sub-blocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. Coupling such aging mitigation with the resulting energy reduction techniques we can obtain up to 2.5x lifetime extension and 40% energy savings with respect to a power managed cache.
Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii
VLSI-SoC3
2011 Partitioned cache architectures for reduced NBTI-induced aging
abstract
Conventional power management knobs such as voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused Negative Bias Temperature Instability (NBTI). Such a benefit can be especially exploited in SRAM memories, which are particularly sensitive to NBTI effects: given their symmetric structure, they cannot in fact take advantage of value-dependent recovery. We propose an architectural solutions that is based on the idea of partitioning a memory into multiple banks of identical size. While this organization has been widely used for reducing both dynamic and static power, its exploitation for aging benefits requires proper management of the existing idleness of the various banks. This can be achieved by means of a sort of time-varying addressing scheme in which addresses are mapped to different banks over time in such a way that the idleness is uniformly distributed over all the banks. Experimental analysis shows that it is possible to simultaneously reducing leakage power and aging in caches, with minimal overhead and without modifying the internal structure of the SRAM arrays.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
DATE2
2011 Buffering of frequent accesses for reduced cache aging
abstract
Previous works have shown that typical power management knobs such as voltage scaling or power gating can also be exploited to reduce aging phenomena caused by Negative Bias Temperature Instability (NBTI). We propose a scheme for power-managed caches that allows to significantly improving the aging of the cache thanks to the use of a small buffer that stores a copy of the lines that are most critical for aging, that is, the ones with the least opportunity of being power-managed; by using the buffer instead of the cache when accessing these critical lines, the original cache is preserved and its lifetime is significantly prolonged. As a side effect, this scheme improves total power since the less energy-hungry buffer is accessed most of the time. Experimental analysis shows this scheme allows to achieve significant (>3x on average) lifetime extensions for the cache, with a concurrent energy saving between 18 and 24%, depending on cache size.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2010 Aging effects of leakage optimizations for caches
abstract
Besides static power consumption, sub-90nm devices have to account for NBTI effects, which are one of the major concerns about system reliability.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2010 Dynamic indexing: concurrent leakage and aging optimization for caches
abstract
Previous works have shown that the traditional implementations of power management (i.e., using power gating or voltage scaling) can also mitigate the aging effect induced by Negative Bias Temperature Instability (NBTI), due to the partial recovery that occurs during the idle intervals used by power management. However, such a potential has been exploited only partially because of the different nature of energy and aging: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually turn into limited aging reductions. We address this problem in the context of caches, for which idleness is related to their access pattern. We propose a dynamic indexing scheme, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the cache lines. In this way it is possible to fully use the leakage optimization potential and to extend the lifetime of a cache. Experimental analysis shows that it is possible to obtain caches that are effectively aging-free, without any penalty in leakage energy reduction.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ISLPED2
2010 Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking
abstract
Partitioning a memory into multiple blocks that can be independently accessed is a widely used technique to reduce its dynamic power. For embedded systems, its benefits can be even pushed further by properly matching the partition to the memory access patterns. When leakage energy comes into play, however, idle memory blocks must be put into a proper low-leakage sleep state to actually save energy when not accessed. In this case, the matching becomes an instance of the power management problem, because moving to and from this sleep state requires additional energy. In this work, we propose an effective solution to the problem of the leakage-aware partitioning of a memory into disjoint subblocks; in particular, we target scratchpad memories, which are commonly used in some embedded systems as a replacement for caches. We show that, although the solution space is extremely large (for a N--block partition, all the combinations of N-1 address boundaries) and nonconvex, it is possible to prove a nontrivial property that considerably reduces the number of partition boundaries to be enumerated, therefore, making exhaustive exploration feasible. We are thus able to provide an optimal solution to the leakage-aware partitioning problem. Experiments on a different sets of embedded applications have shown that total energy savings larger than 60 percent on average can be obtained, with a marginal overhead in execution time, thanks to an effective implementation of the low-leakage sleep state.
Mirko Loghi, Olga Golubeva, Enrico Macii, Massimo Poncino
IEEE Trans. Computers1
2009 Energy-optimal synchronization primitives for single-chip multi-processors
abstract
Synchronization among tasks accounts for a sizable fraction of the energy consumption and execution time of applications running on Multi-Processor Systems-on-Chips platforms. In order to achieve fast and energy-efficient operations, it is therefore essential to implement efficient and power-frugal synchronization primitives. The design of such primitives is complicated by several software and hardware issues, such as: processors running at different speeds, different implementations of the waiting phase upon entering the critical section, and the ratio between static and dynamic power. In this work, we compare a set of classical implementations (i.e., based on busy waiting, or on sleep states) of mutex semaphores, and propose a hybrid (wait/sleep) semaphore in which the sleep state is entered only after a number of busywait cycles. The proposed scheme provides the best overall energy-delay product with respect to previously proposed schemes. Furthermore, we identify an optimal length of the busy-wait cycles, which is empirically shown to depend on the time required to switch from the sleep to the active state.
Cesare Ferri, R. Iris Bahar, Mirko Loghi, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2009 A cosimulation methodology for HW/SW validation and performance estimation
abstract
Cosimulation strategies allow us to simulate and verify HW/SW embedded systems before the real platform is available. In this field, there is a large variety of approaches that rely on different communication mechanisms to implement an efficient interface between the SW and the HW simulators. However, the literature lacks a comprehensive methodology which addresses the need for integrating and synchronizing heterogeneous simulators, like, for example, the SystemC simulation kernel for HW modules and an instruction set simulator for SW applications, without being intrusive for the HW and SW descriptions involved in the simulation. In this context, this article presents, compares, and integrates in a system-level framework two different co-simulation strategies for modeling, analyzing, and validating the performance of a HW/SW embedded system. Moreover, for both of them, a mechanism is proposed to provide an accurate time synchronization of the HW/SW communication. The first strategy is intended to provide an early cosimulation environment where HW/SW interaction can be validated without involving the operating system. The communication is implemented between a single SW task and a SystemC description of an HW module by exploiting the features of the remote debugging interface of a debugger (the GNU GDB), and by modifying the SystemC simulation kernel. On the other hand, the second strategy is intended to be used in further development steps, when the operating system is introduced to validate the cosimulation between HW modules and multitasking SW applications. In this approach, the communication is implemented via interrupts by using the features offered by the operating system. Experimental results are reported on two different case studies to analyze and compare the effectiveness of both the approaches.
Franco Fummi, Mirko Loghi, Massimo Poncino, Graziano Pravadelli
ACM Trans. Design Autom. Electr. Syst.2
2009 Tag Overflow Buffering: Reducing Total Memory Energy by Reduced-Tag Matching
abstract
We propose a novel energy-efficient cache architecture based on a matching mechanism that uses a reduced number of tag bits. The idea behind the proposed architecture is based on moving a large subset of the tag bits from the cache into an external register (called theTagOverflowBuffer) that serves as an identifier of the current locality of the memory references. Dynamic energy efficiency is achieved by accessing, for most of the memory references, a reduced-tag cache; furthermore, because of the reduced number of tag bits, leakage energy is also reduced as a by-product. We achieve average energy savings ranging from 16% to 40% (depending on different cache structural parameters) on total (i.e., static and dynamic) cache energy, and measured on a standard suite of embedded applications.
Mirko Loghi, Paolo Azzoni, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Architectural leakage-aware management of partitioned scratchpad memories
Olga Golubeva, Mirko Loghi, Massimo Poncino, Enrico Macii
DATE2
2007 On the energy efficiency of synchronization primitives for shared-memory single-chip multiprocessors
abstract
Applications running on Multiprocessor Systems-on-Chips (MP-SoCs) exhibit complex interaction patterns, resulting in significant amounts of time spent while synchronizing for mutually exclusive access to shared resources. Such an overhead is expected to increase with the degree of parallelism and with the mutual correlation of concurrent tasks, thus becoming in a severe obstacle to the full exploitation of a system potential. Although the topic has been extensively studied in the literature, in MPSoC architectures, which exhibit different tradeoffs with respect to traditional multi-processors, the available results may not be valid or hold only partially. Furthermore, the strict energy budget of MPSoCs requires also the evaluation of the energy efficiency of such synchronization primitives. In this work we survey various state-of-the-art implementations of synchronization primitives, in order to assess their impact on performance and on energy consumption. The results of our analysis show that some commonly accepted intuitions in the multiprocessor domain do not hold in the context of MPSoCs.
Olga Golubeva, Mirko Loghi, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2007 Locality-driven architectural cache sub-banking for leakage energy reduction
abstract
In most processors, caches account for the largest fraction of onchip transistors, thus being a primary candidate for tackling the leakage problem. Existing architectural solutions usually rely on customized cache structures, which are needed to implement some kind of power management policy. Memory arrays, however, are carefully developed and finely tuned by foundries, and their internal structure is typically non accessible to system designers.
Olga Golubeva, Mirko Loghi, Enrico Macii, Massimo Poncino
ISLPED2
2007 Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural Support
abstract
In today's multiprocessor SoCs (MPSoCs), parallel programming models are needed to fully exploit hardware capabilities and to achieve the 100 Gops/W energy efficiency target required for ambient intelligence applications. However, mapping abstract programming models onto tightly power-constrained hardware architectures imposes overheads which might seriously compromise performance and energy efficiency. The objective of this work is to perform a comparative analysis of message passing versus shared memory as programming models for single-chip multiprocessor platforms. Our analysis is carried out from a hardware-software viewpoint: we carefully tune hardware architectures and software libraries for each programming model. We analyze representative application kernels from the multimedia domain, and identify application-level parameters that heavily influence performance and energy efficiency. Then, we formulate guidelines for the selection of the most appropriate programming model and its architectural support
Francesco Poletti, Antonio Poggiali, Davide Bertozzi, Luca Benini, Paul Marchal, Mirko Loghi, Massimo Poncino
IEEE Trans. Computers6
2007 Power macromodeling of MPSoC message passing primitives
abstract
Estimating the energy consumption of software in multiprocessor systems-on-chip (MPSoCs) is crucial for enabling quick evaluations of both software and hardware optimizations. However, high-level estimations should be applicable at software level, possibly constructing effective power models depending on parameters that can be extracted directly from the application characteristics. We propose a methodology for accurate analysis of power consumption of message-passing primitives in a MPSoC, and, in particular, an energy model which, in spite of its simplicity, allows to model the traffic-dependent nature of energy consumption through the use of a single, abstract parameter, namely, the size of the message exchanged.
Mirko Loghi, Luca Benini, Massimo Poncino
ACM Trans. Embed. Comput. Syst.1
2006 ISS-centric modular HW/SW co-simulation
abstract
Modular design is an important requirement in modern embedded system design flows because of the widespread acceptance of new paradigms such as IP core reuse and platform-based design. Co-simulation frameworks must thus support modular design, since programmable devices, ad-hoc HW components, and the interconnect infrastructure must be easily interchangeable in order to allow design exploration while keeping the SW portion unchanged or only marginally changed. The proposed co-simulation framework implements such a modular approach to co-simulation by means of a novel paradigm in which HW models can be modified on the fly by keeping the SW parts unchanged. This is achieved through an ISS-centric co-simulation strategy in which modularity is provided in terms of (i) the replacement of HW components thanks to the use of a common interface based on the device address space, or (ii) the use of different ISS's, thanks to a re-configurable simulator. We demonstrate our approach onto an industrial-strength embedded application, showing that the proposed co-simulation strategy provides both high speed and accuracy.
Franco Fummi, Giovanni Perbellini, Mirko Loghi, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2006 Synchronization-driven dynamic speed scaling for MPSoCs
abstract
Equalizing the ratios between workloads and speeds of processing elements provides the optimal speed allocation. Based on that principle, this work describes a dynamic speed setting policy for multiprocessor systems-on-chip (MPSoCs) that relies on the estimation of processor idle times specifically due to the synchronization work. The policy provides two advantages: first, it does not rely on any assumption about the communication pattern of the application executed by the system. Second, it is purely architectural; it automatically detects changes in the system workload and sets processors speeds accordingly by means of a custom hardware block.Results on a parallel MPEG video decoding application show an EDP saving above 55%, averaged over several datasets, corresponding to an energy saving above 50%, and a corresponding penalty in performance below 8%.
Mirko Loghi, Massimo Poncino, Luca Benini
ISLPED1
2006 Cache coherence tradeoffs in shared-memory MPSoCs
abstract
Shared memory is a common interprocessor communication paradigm for single-chip multiprocessor platforms. Snoop-based cache coherence is a very successful technique that provides a clean shared-memory programming abstraction in general-purpose chip multiprocessors, but there is no consensus on its usage in resource-constrained multiprocessor systems on chips (MPSoCs) for embedded applications. This work aims at providing a comparative energy and performance analysis of cache-coherence support schemes in MPSoCs. Thanks to the use of a complete multiprocessor simulation platform, which relies on accurate technology-homogeneous power models, we were able to explore different cache-coherent shared-memory communication schemes for a number of cache configurations and workloads.
Mirko Loghi, Massimo Poncino, Luca Benini
ACM Trans. Embed. Comput. Syst.1
2005 Virtual Hardware Prototyping through Timed Hardware-Software Co-Simulation
abstract
Designers of factory automation applications increasingly demand tools for rapid prototyping of hardware extensions to existing systems and verification of resulting behaviors through hardware and software co-simulation. The paper presents a framework for the timing-accurate co-simulation of HDL models and their verification against hardware and software running on an actual embedded device of which only a minimal knowledge of the current design is required. Experiments on real-life applications show that early architectural and design decisions can be taken by measuring the expected performance on the models realized using the proposed framework.
Franco Fummi, Mirko Loghi, Stefano Martini, Marco Monguzzi, Giovanni Perbellini, Massimo Poncino
DATE2
2005 Tag Overflow Buffering: An Energy-Efficient Cache Architecture
abstract
We propose a novel energy-efficient memory architecture which relies on the use of a cache with a reduced number of tag bits. The idea behind the proposed architecture is based on moving a large number of the tag bits from the cache into an external register (tag overflow buffer) that identifies the current locality of the memory references; additional hardware allows us to dynamically update the value of the reference locality contained in the buffer. Energy efficiency is achieved by using, for most of the memory accesses, a reduced-tag cache. This architecture is minimally intrusive for existing designs, since it assumes the use of a regular cache, and does not require any special circuitry internal to the cache such as row or column activation mechanisms. Average energy savings are 51% on tag energy, corresponding to about 20% saving on total cache energy, measured on a set of typical embedded applications.
Mirko Loghi, Paolo Azzoni, Massimo Poncino
DATE1
2005 Exploring Energy/Performance Tradeoffs in Shared Memory MPSoCs: Snoop-Based Cache Coherence vs. Software Solutions
abstract
Shared memory is a common interprocessor communication paradigm for single-chip multi-processor platforms. Snoop-based cache coherence is a very successful technique that provides a clean shared-memory programming abstraction in general-purpose chip multiprocessors, but there is no consensus on its usage in resource-constrained multiprocessor systems on chips (MPSoC) for embedded applications. This work aims at providing a comparative energy and performance analysis of cache coherence support schemes in MPSoC. Thanks to the use of a complete multiprocessor simulation platform, which relies on accurate technology-homogeneous power models, we were able to explore different cache-coherent shared-memory communication schemes for a number of cache configurations and workloads.
Mirko Loghi, Massimo Poncino
DATE1
2005 Exploring the energy efficiency of cache coherence protocols in single-chip multi-processors
abstract
The performance of the various cache coherence protocols proposed in the literature have been extensively analyzed in the context of high-performance multi-processor systems.A similar analysis for Multi-Processor Systems-on-Chips (MP-SoCs), where energy is at least as important as performace, and for which strict constraints on hardware and software resources do exist, has not been done yet.This work provides an effort in that sense, showing energy/performance tradeoffs for different snoop-based protocols on a realistic MPSoC architecture. The analysis leverage a multi-processor simulation platform, augmented with accurate power models, that allows cycle-accurate simulations.Our analysis show that (i) cache write policy is actually more important than the actual cache coherence protocol, and (ii) matching the programming model and style to the architecture may have dramatic effects on the energy and performance of the system.
Mirko Loghi, Martin Letis, Luca Benini, Massimo Poncino
ACM Great Lakes Symposium on VLSI1
2004 Analyzing On-Chip Communication in a MPSoC Environment
abstract
This work focuses on communication architecture analysis for multi-processor systems-on-chips (MPSoCs), and it leverages a SystemC-based platform to simulate a complete multi-processor system at the cycle-accurate and signal-accurate level. These features allow to stimulate the communication sub-system with functional traffic generated by real applications running on top of a configurable number of ARM processors. This opens up the possibility for communication infrastructure exploration and for the investigation of its impact on system performance at the highest level of accuracy. Our simulation environment proved capable of a detailed comparative analysis between two industry-standard communication architectures, under realistic workloads and different system configurations, pointing out the impact of fine grained architectural mismatches on macroscopic performance differences.
Mirko Loghi, Federico Angiolini, Davide Bertozzi, Luca Benini, Roberto Zafalon
DATE1
2004 Cycle-accurate power analysis for multiprocessor systems-on-a-chip
abstract
Developing energy-aware software for multiprocessor systems-on-chip (MPSoCs) is a difficult task, which requires the knowledge of the distribution of the power consumption among several heterogeneous devices (cores, memories, busses, etc.). In this work we analyze the power breakdowns of power consumption for a complete MPSoC platform, under several application workloads and operating conditions. We leverage a complete-system simulation platform with accurate power models for all key hardware modules. Our analysis shows that caches and system interconnect dominate in the power breakdown, pointing out how software locality is meaningful not only for performance but also for energy optimization.
Mirko Loghi, Massimo Poncino, Luca Benini
ACM Great Lakes Symposium on VLSI1
2004 Analyzing Power Consumption of Message Passing Primitives in a Single-Chip Multiprocessor
abstract
In this work, we propose a methodology for the accurate analysis of the power consumption of interprocessor communication in an MPSoC, and the construction of high-level power macromodels. The models leverage a complete MPSoC power estimation environment, that allows to evaluate the power consumption of various software functions, including message passing primitives, which could not be fully characterized in single-processor analysis framework developed in the past. Based on this data we built power macromodels that achieve average estimation errors below 5%.
Mirko Loghi, Luca Benini, Massimo Poncino
ICCD1