VLDB 2026 Research / reviewers in the wild / expert
Davide Bertozzi
dblp:48/5390
· DBLP profile ↗
91ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0001-7462-4551ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 82 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 27 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Special Session: Optimizing Edge AI - Current Challenges and the Neuromorphic OutlookabstractThe increasing deployment of AI (artificial intelligence) on edge devices presents major challenges due to strict constraints on computation, memory, energy, and latency. Effective Edge AI systems thus require multi-objective optimization that balances accuracy, hardware efficiency, and reliability. The Horizon Twinning project AIDA4Edge tackles these challenges by developing methods for efficient and reliable AI on resource-constrained platforms. This paper presents key approaches explored within the project, including neural network quantization, hardware-aware neural architecture search, dynamic neural networks, and self-adaptive resilient AI architectures. Finally, these strategies are placed within a broader, biologically inspired paradigm, highlighting neuromorphic computing as a natural continuation of Edge AI efforts toward highly efficient and resilient intelligent systems. Milan R. Dincic, Zoran H. Peric, Davide Bertozzi, Alice Bizzarri, Rizwan Tariq Syed, Edward G. Jones, Riccardo Zese, Marko S. Andjelkovic, Fabian Vargas 0001, Milos Krstic, Oliver Rhodes, Modhe Almelihi, Tamara Milovanovic, Ivan Popovic, Sofija Peric |
DDECS | 3 |
| 2026 | Cross-Layer Reliability Analysis of Slimmable Neural Networks under Permanent Faults
Nikolaos Zazatis, Alessandro Veronesi, Konstantinos Varakliotis, Pelopidas Tsoumanis, Christos P. Sotiriou, Letícia Maria Veiras Bolzani, Marko S. Andjelkovic, Davide Bertozzi |
VTS | 8 |
| 2025 | EMBER: A Cycle-based Framework for Early-Stage Reliability Assessment in Parametric RTL DesignsabstractModern trends towards higher architectural complexity and smaller technology nodes do not align with the requirement for reliability that many application domains expose. While system robustness remains one of the most critical aspects for missioncritical application domains, the currently available tools allow designers to accurately investigate the reliability only at the late design stages, limiting the effectiveness of their intervention in addressing architectural vulnerabilities.In this context, this paper presents a novel framework for early-stage reliability assessment that enables fast evaluations to guide the RTL design process, without compromising analysis quality or relying on imprecise high-level fault injection models. In the presented analysis, the paper shows how the proposed framework leads to a significant time saving when targeting highly parametric hardware designs, achieving up to 79 x compile time reduction, and up to $37.6 x$ simulation time reduction when compared to other state-of-the-art approaches. Alessandro Veronesi, Letícia Maria Veiras Bolzani, Michele Favalli, Milos Krstic, Davide Bertozzi |
ATS | 5 |
| 2025 | Multi-Partner Project: Twinning for Excellence in Reliable Electronics (TWIN-RELECT)abstractReliable electronics plays a major role in shaping our daily lives, being a key enabler for critical applications, such as space missions, avionics, automotive, medicine, banking, automated industry, wireless communication networks, etc. However, design of highly reliable electronic systems remains a challenge with the advances in semiconductor technology and increase in integrated circuit (IC) complexity. In this work, we introduce the Horizon Europe Twinning project TWIN-RELECT, aimed at strengthening the scientific expertise in designing reliable integrated circuits. The paper presents the general project concept and objectives, and main directions of the joint research activities. The primary scientific goal is to contribute to the development of novel, more efficient, European Electronic Design Automation (EDA) tool-chain for design of reliable chips. Marko S. Andjelkovic, Fabian Vargas 0001, Milos Krstic, Luigi Dilillo, Alain Michez, Frédéric Wrobel, Davide Bertozzi, Mikel Luján, Christos Georgakidis, Nikolaos Chatzivangelis, Katerina Tsilingiri, Nikolaos Zazatis, Georgios Ioannis Paliaroutis, Pelopidas Tsoumanis, Christos P. Sotiriou |
DATE | 7 |
| 2025 | AIDA4Edge: Twinning for Excellence in Adaptive Edge Artificial IntelligenceabstractThe growing demand for deployment of Artificial Intelligence (AI) on resource-constrained edge devices has motivated extensive research on the design of efficient edge-compatible AI hardware accelerators. One of the most promising solutions are the self-adaptive AI accelerators, capable of optimizing in real time their performance and energy consumption according to application requirements. This work introduces the EU-funded project Twinning for Excellence in Adaptive Edge Artificial Intelligence (AIDA4Edge), aimed to advance the state-of-the-art in the design of adaptive neural network accelerators for edge applications. The main goal is to develop a novel hybrid self-adaptive neural network architecture combining spiking and artificial neural networks, and supporting runtime adaptation of network functionality, precision and reliability. Furthermore, we aim to enhance the neural network training by incorporating hardware and quantization constraints in an automated tuning engine. Marko S. Andjelkovic, Rizwan Tariq Syed, Alessandro Veronesi, Fabian Vargas 0001, Markus Ulbricht 0002, Letícia Maria Veiras Bolzani, Milos Krstic, Davide Bertozzi, Edward G. Jones, Oliver Rhodes, Riccardo Zese, Michele Favalli, Alice Bizzarri, Evelina Lamma, Marco Gavanelli, Elena Bellodi, Zoran H. Peric, Jelena Nikolic, Milan R. Dincic, Aleksandra Jovanovic 0001, Dejan Ciric, Nikola Vucic, Sofija Peric, Jelena Jovanovic 0006, Milica Stojanovic, Tatjana R. Nikolic, Goran Nikolic, Jelena Nedeljkovic, Danijel Dankovic, Emilija Zivanovic, Milos Marjanovic, Sandra Veljkovic, Nikola Mitrovic, Bratislav Predic, Tamara Milovanovic |
DSD | 8 |
| 2025 | An Efficient Multicast Addressing Encoding Scheme for Multi-Core Neuromorphic ProcessorsabstractMulti-Core neuromorphic processors are becoming increasingly significant due to their energy-efficient local computing and scalable modular architecture, particularly for event-based processing applications. However, minimizing the cost of inter-core communication, which accounts for the majority of energy usage, remains a challenging issue. Beyond optimizing circuit design at lower abstraction levels, an efficient multicast addressing scheme is crucial. We propose a hierarchical bit string encoding scheme that largely expands the addressing capability of state-of-the-art symbol-based schemes for the same number of routing bits. When put at work with a real neuromorphic task, this hierarchical bit string encoding achieves a reduction in area cost by approximately 29% and decreases energy consumption by about 50%. Aron Bencsik, Giacomo Indiveri, Davide Bertozzi |
ISCAS | 4 |
| 2024 | Adaptive Localization for Autonomous Racing Vehicles with Resource-Constrained Embedded PlatformsabstractModern autonomous vehicles have to cope with the consolidation of multiple critical software modules processing huge amounts of real-time data on power- and resource-constrained embedded MPSoCs. In such a highly-congested and dynamic scenario, it is extremely complex to ensure that all components meet their quality-of-service requirements (e.g., sensor frequencies, accuracy, responsiveness, reliability) under all possible working conditions and within tight power budgets. One promising solution consists of taking advantage of complementary resource usage patterns of software components by implementing dynamic resource provisioning. A key enabler of this paradigm consists of augmenting applications with dynamic reconfiguration capability, thus adaptively modulating quality-of-service based on resource availability or proactively demanding resources based just on the complexity of the input at hand. The goal of this paper is to explore the feasibility of such a dynamic model of computation for the critical localization function of self-driving vehicles, so that it can burden on system resources just for what is needed at any point in time or gracefully degrade accuracy in case of resource shortage. We validate our approach in a harsh scenario, by implementing it in the localization module of an autonomous racing vehicle. Experiments show that we can adapt to variations in operational conditions such as the system workload, and that we can also achieve an overall reduction of platform utilization and power consumption for this computation-greedy software module by up to$1.6\times$and$1.5\times$, respectively, for roughly the same quality of service. Federico Gavioli, Gianluca Brilli, Paolo Burgio, Davide Bertozzi |
DATE | 4 |
| 2024 | Cross-Layer Reliability Analysis of NVDLA Accelerators: Exploring the Configuration SpaceabstractInvestigating the effects of Single Event Upset in domain-specific accelerators represents one of the key enablers to deploy Deep Neural Networks (DNNs) in mission-critical edge applications. Currently, reliability analyses related to DNNs mainly focus either on the DNNs model, at application level, or on the hardware accelerator, at architecture level. This paper presents a systematic cross-layer reliability analysis of NVIDIA Deep-Learning Accelerator, a popular family of industry-grade, open and free DNN accelerators. The goals are i) to analyze the propagation of faults from the hardware to the application level, and ii) to compare different architectural configurations. Our investigation delivers new insights into the performance-accuracy-reliability trade-off spanned by the configuration space of Deep Learning accelerators. In particular, the Failure in Time can be reduced up to 4.3x for the same DNN model accuracy and by up to 9.4x for the same performance, while accounting 6.5x inference latency and 1.1% accuracy drop, respectively. Alessandro Veronesi, Alessandro Nazzari, Dario Passarello, Milos Krstic, Michele Favalli, Luca Cassano, Antonio Miele, Davide Bertozzi, Cristiana Bolchini |
ETS | 8 |
| 2024 | LNOI Wireless Switches Based on Optical Phased Arrays for On-Chip CommunicationabstractOn-chip optical wireless links are attracting a great deal of interest as they can provide a possible solution to overcome the drawbacks associated with wired connections. In this paper, we propose a new approach for on-chip communication using optical wireless switches based on thin-film lithium niobate on insulator (LNOI) technology. The optical wireless switches exploit reconfigurable optical phased arrays (OPAs) both at the transmitter and at the receivers. We investigate the radiation characteristics and design criteria of the LN antenna element serving as a unit radiator in the OPAs. We then demonstrate the implementation of an on-chip optical wireless switch in a simple infinite homogeneous host medium and assuming a realistic multilayer structure configuration. Moreover, we examine the impact of various geometrical parameters and the fabrication imperfections on the device performance and provide a discussion on the design optimization. Our findings assess the feasibility of optical wireless switches based on LNOI technology for on-chip wireless communication. Simone Ferraresi, Gaetano Bellanca, Marina Barbiroli, Franco Fuschini, Velio Tralli, Davide Bertozzi, Vincenzo Petruzzelli, Giovanna Calò |
IEEE J. Sel. Areas Commun. | 7 |
| 2022 | Exploring Software Models for the Resilience Analysis of Deep Learning Accelerators: the NVDLA Case StudyabstractDeep learning accelerator models described with software imperative languages are frequently used for their large-scale reliability analysis in order to overcome the prohibitive simulation times of logic-level and RTL models. However, they are faced with the challenge of preserving consistency between software-visible variables and faulty microarchitectural states. The goal of this work is to determine a suitable accelerator modelling that enables analysis without overloading the simulation engine. Toward this goal, the paper explores different accelerator modelling strategies featuring increasing levels of hardware visibility. They are compared in their capability to gain insights into the reliability of the multiply-and-accumulate (MAC) pipeline of an industry-standard deep learning accelerator from NVIDIA. Our results show that subtle microarchitectural details that are typically overlooked by competing approaches play a relevant role in determining accelerator reliability. Alessandro Veronesi, Francesco Dall'Occo, Davide Bertozzi, Michele Favalli, Milos Krstic |
DDECS | 3 |
| 2020 | Cross-Layer Hardware/Software Assessment of the Open-Source NVDLA Configurable Deep Learning AcceleratorabstractThe Nvidia Deep Learning Accelerator (NVDLA) is a free and open architecture that aims at promoting a standard way of designing deep neural network (DNN) inference engines. The analogy between open-source software and hardware points to FPGAs as ideal implementation platforms for open hardware accelerators. However, the instantiation flexibility enabled by reconfigurable logic should be correlated to the capacity of cost-effective devices. This paper explores the resource utilization-performance trade-offs spanned by the main precompiled NVDLA accelerator configurations on top of the mainstream Zynq UltraScale+ MPSoC. For the sake of comprehensive end-to-end performance characterization, the inference rate of the software stack is matched to that of the accelerator hardware, thus identifying current bottlenecks and promising optimization directions. Alessandro Veronesi, Milos Krstic, Davide Bertozzi |
VLSI-SOC | 3 |
| 2020 | PSION+: Combining Logical Topology and Physical Layout Optimization for Wavelength-Routed ONoCsabstractOptical networks-on-chip (ONoCs) are a promising solution for high-performance multicore integration with better latency and bandwidth than traditional electrical NoCs. Wavelength-routed ONoCs (WRONoCs) offer yet additional performance guarantees. However, WRONoC design presents new EDA challenges which have not yet been fully addressed. So far, most topology analysis is abstract, i.e., overlooks layout concerns, while for layout the tools available perform place and route (P&R) but no topology optimization. Thus, a need arises for a novel optimization method combining both aspects of WRONoC design. In this article, such a method, PSION+, is laid out. This new procedure uses a linear programming model to optimize a WRONoC physical layout template to optimality. This template-based optimization scheme is a new idea in this area that seeks to minimize problem complexity while keeping design flexibility. A simple layout template format is introduced and explored. Finally, multiple model reduction techniques to reduce solver run-time are also presented and tested. When compared to the state-of-the-art design procedure, results show a decrease in maximum optical insertion loss of 41%. Alexandre Truppel, Tsun-Ming Tseng, Davide Bertozzi, José Carlos Alves, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | PSION: Combining Logical Topology and Physical Layout Optimization for Wavelength-Routed ONoCsabstractOptical Networks-on-Chip (ONoCs) are a promising solution for high-performance multi-core integration with better latency and bandwidth than traditional Electrical NoCs. Wavelength-routed ONoCs (WRONoCs) offer yet additional performance guarantees. However, WRONoC design presents new EDA challenges which have not yet been fully addressed. So far, most topology analysis is abstract, i.e., overlooks layout concerns, while for layout the tools available perform Place & Route (P&R) but no topology optimization. Thus, a need arises for a novel optimization method combining both aspects of WRONoC design. In this paper such a method, PSION, is laid out. When compared to the state-of-the-art design procedure, results show a 1.8x reduction in maximum optical insertion loss. Alexandre Truppel, Tsun-Ming Tseng, Davide Bertozzi, José Carlos Alves, Ulf Schlichtmann |
ISPD | 3 |
| 2018 | A Boolean model for delay fault testing of emerging digital technologies based on ambipolar devicesabstractEmerging nanotechnonologies such as ambipolar carbon nanotube field effect transistors (CNTFETs) and silicon nanowire FETs (SiNFETs) provide ambipolar devices allowing the design of more complex logic primitives than those found in today's typical CMOS libraries. When switching, such devices show a behavior not seen in simpler CMOS and FinFET cells, making unsuitable the existing delay fault testing approaches. We provide a Boolean model of switching ambipolar devices to support delay fault testing of logic cells based on such devices both in Boolean and Pseudo-Boolean satisfiability engines. Marcello Dalpasso, Davide Bertozzi, Michele Favalli |
DATE | 2 |
| 2018 | Wavelength-Routed Optical Networks-on-Chip: Design Methods and Tools to Bridge the Gap Between Logic Topologies and Physical Ones in 3D ArchitecturesabstractSilicon photonics is gaining momentum as a candidate technology platform for future intra- and inter-chip communications. However, its industrial uptake depends not only on technology maturity, but also on the capability to bridge the abstraction gap between technology developers and system designers. This paper presents an early-stage cross-layer refinement methodology of wavelength-routed optical network-on-chip topologies, linking logic topology synthesis to the physical implementation steps. Davide Bertozzi, Marco Gavanelli, Maddalena Nonato |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Interfacing 3D-stacked Electronic and Optical NoCs with Mixed CMOS-ECL Bridges: a Realistic Preliminary AssessmentabstractThe combination of optical networks-on-chip and 3D stacking represents the most promising system integration framework to overcome the communication bottleneck of future many-core processors. From an architecture viewpoint, the availability of an energy-efficient, low-latency bridge connecting the electronic network-on-chip with the optical one is as important as the maturity of the optical interconnect technology. The key design challenge consists of overcoming the inherent serial nature of optical communications, which is typically pursued by increasing either the data rate or the bit-level parallelism, or by a combination thereof. This paper explores an hybrid CMOS-ECL technology platform for bridge implementation by means of a complete logic synthesis effort. By spanning the wider configuration space of the hybrid bridge with respect to fully-CMOS realizations, the paper identifies the most energy-efficient configurations and provides a comparative assessment of achievable quality metrics. Derived results represent a solid and realistic starting point for future optimizations and for the refinement into an actual layout. Mahdi Tala, Oliver Schrape, Milos Krstic, Davide Bertozzi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | CustomTopo: a topology generation method for application-specific wavelength-routed optical NoCsabstractOptical network-on-chip (NoC) is a promising platform beyond electronic NoCs. In particular, wavelength-routed optical network-on-chip (WRONoC) is renowned for its high bandwidth and ultra-low signal delay. Current WRONoC topology generation approaches focus on full-connectivity, i.e. all masters are connected to all slaves. This assumption leads to wasted resources for application-specific designs. In this work, we propose CustomTopo: a general solution to the topology generation problem on WRONoCs that supports customized connectivity. CustomTopo models the topology structure and its communication behavior as an integer-linear-programming (ILP) problem, with an adjustable optimization target considering the number of add-drop filters (ADFs), the number of wavelengths, and insertion loss. The time for solving the ILP problem in general positively correlates with the network communication densities. Experimental results show that CustomTopo is applicable for various communication requirements, and the resulting customized topology enables a remarkable reduction in both resource usage and insertion loss. Mengchu Li, Tsun-Ming Tseng, Davide Bertozzi, Mahdi Tala, Ulf Schlichtmann |
ICCAD | 3 |
| 2018 | Understanding the Design Space of Wavelength-Routed Optical NoC Topologies for Power-Performance OptimizationabstractSilicon photonics is the most promising emerging technology to deliver on- and off-chip communication performance and power that vastly exceed the capabilities of electronics. However, a significant abstraction gap does exist between novel devices and circuits and the higher-order switching structures that system designers need to instantiate. Currently, designers mostly rely on their intuition to bridge this abstraction gap. This paper lays the groundwork for a more rigorous and effective approach, by vertically-integrating the most advanced design methods and tools for topology synthesis and refinement in the context of a novel performance analysis framework. As a result, we can extract the highest aggregate bandwidth out of an optical network-on-chip topology, and provide an early-stage analysis of its static power, thus unveiling unexplored portions of the design space and interpreting its characteristics. Mahdi Tala, Davide Bertozzi |
VLSI-SoC | 2 |
| 2018 | Special session on overcoming reliability and energy-efficiency challenges with silicon photonics for future manycore computingabstractSilicon photonics has emerged in recent years as one of the most promising solutions to overcome the challenge of worsening chip-scale communication performance with technology scaling. Recent breakthroughs in silicon photonic device fabrication and CMOS integration have presented computer designers with an opportunity to devise on-chip optical networks that have significant advantages in bandwidth density, energy-efficiency, and propagation delay over traditional electrical solutions. Not surprisingly, the challenge of designing chip-scale silicon photonic communication fabrics is today actively being pursued by a number of researchers worldwide. Many semiconductor companies (e.g., Intel, IBM) have begun investing heavily into silicon photonics and are releasing functional prototypes. However, silicon photonic interconnects have high susceptibility to faults due to several factors such as homodyne and heterodyne crosstalk, process variations, and thermal fluctuations. Moreover, photonic devices can have a significant power dissipation footprint, which can increase further when compensating for potential faults. New network-centric circuits, architectures, tools, and protocols are required to overcome these challenges. This special session focuses on the overarching goals of enabling high fault resilience and energy-efficiency in emerging silicon photonic on-chip networks. Sudeep Pasricha, Davide Bertozzi, Hui Li 0034 |
VTS | 2 |
| 2017 | A tool for synthesizing power-efficient and custom-tailored wavelength-routed optical ringsabstractOut of all the optical network-on-chip topologies, the ring has been proved to be far superior to its competitors: the contention-free all-to-all communications offer the lowest latency possible, while its clean physical design with few crossings and ring resonators provides unmatchable power results. The ring implements simultaneous communications by using a communication matrix that sets a distinctive waveguide-wavelength pair for each of them. That communication matrix has a high impact on energy consumption, but so far there have been very few efforts towards optimizing and automating its design. As far as we know, we propose the best optical ring design algorithm, which produces rings with the lowest number of wavelengths and waveguides in the literature. The algorithm is completed with a layout-aware and fully automated laser power calculation framework to help the user choose the most power-efficient design point. Marta Ortín-Obón, Luca Ramini, Víctor Viñals, Davide Bertozzi |
ASP-DAC | 4 |
| 2017 | An asynchronous NoC router in a 14nm FinFET library: Comparison to an industrial synchronous counterpartabstractAn asynchronous high-performance low-power 5-port network-on-chip (NoC) router is introduced. The proposed router integrates low-latency input buffers using a circular FIFO design, and a novel end-to-end credit-based virtual channel (VC) flow control for a replicated switch architecture. This asynchronous router is then compared to an AMD synchronous router, in a realistic advanced 14nm FinFET library. This is the first such comparison, to the best of our knowledge, using a real synchronous router baseline already fabricated in several commercial products. Initial post-synthesis pre-layout experiments show dominating results for the asynchronous router, when compared to the synchronous router. In particular, 55% less area and 28% latency improvement are observed for the asynchronous implementation. Also, 88% and 58% savings in idle and active power, respectively, are obtained. Weiwei Jiang 0002, Davide Bertozzi, Gabriele Miorandi, Steven M. Nowick, Wayne P. Burleson, Greg Sadowski |
DATE | 2 |
| 2017 | Logic programming approaches for routing fault-free and maximally parallel wavelength-routed optical networks-on-chip (Application paper)abstractAbstract One promising trend in digital system integration consists of boosting on-chip communication performance by means of silicon photonics, thus materializing the so-called Optical Networks-on-Chip. Among them, wavelength routing can be used to route a signal to destination by univocally associating a routing path to the wavelength of the optical carrier. Such wavelengths should be chosen so to minimize interferences among optical channels and to avoid routing faults. As a result, physical parameter selection of such networks requires the solution of complex constrained optimization problems. In previous work, published in the proceedings of the International Conference on Computer-Aided Design, we proposed and solved the problem of computing the maximum parallelism obtainable in the communication between any two endpoints while avoiding misrouting of optical signals. The underlying technology, only quickly mentioned in that paper, is Answer Set Programming. In this work, we detail the Answer Set Programming approach we used to solve such problem. Another important design issue is to select the wavelengths of optical carriers such that they are spread across the available spectrum, in order to reduce the likelihood that, due to imperfections in the manufacturing process, unintended routing faults arise. We show how to address such problem in Constraint Logic Programming on Finite Domains. Marco Gavanelli, Maddalena Nonato, Andrea Peano, Davide Bertozzi |
Theory Pract. Log. Program. | 4 |
| 2017 | Contrasting Laser Power Requirements of Wavelength-Routed Optical NoC Topologies Subject to the Floorplanning, Placement, and Routing Constraints of a 3-D-Stacked SystemabstractA realistic assessment of optical networks-on-chip (ONoCs) can be performed only in the context of a comprehensive floorplanning strategy for the system as a whole, especially when the 3-D stacking of electronic and optical layers is implemented. This paper fosters layout-aware ONoC design by developing a physical mapping methodology for wavelength-routed ONoC topologies subject to the floorplanning, placement, and routing constraints that arise in a 3-D-stacked environment. As a result, this paper is able to compare the power efficiency and signal-to-noise ratio of ring-based versus filter-based wavelength-routed topologies as determined by their physical design flexibility. Marta Ortín-Obón, Mahdi Tala, Luca Ramini, Víctor Viñals, Davide Bertozzi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Design technology for fault-free and maximally-parallel wavelength-routed optical networks-on-chipabstractThe recent interest in emerging interconnect technologies is bringing the issue of a proper EDA support for them to the forefront, so to tackle the design complexity. A relevant case study is provided by wavelength-routed optical NoCs (WRONoCs), which add communication performance guarantees to the typical latency, throughput and power benefits of an optical link, thus providing an appealing technology for the photonic integration of high-end embedded systems. Typically, only abstract WRONoC models are considered to figure out architecture-level performance, and logic connectivity patterns for the quantification of the required signal strength (i.e., static power). However, this design practice overlooks the needed refinement step, where key physical parameters are assigned such as wavelengths of the optical channels, and size of the optical filters. This step is unfortunately not decoupled from the architectural evaluation, since its main constraint (i.e., avoiding routing faults) turns out to be a key limiter for both the network scale and the achievable communication parallelism. By proposing a formal methodology to select WRONoC parameters while avoding the routing fault concern, this paper aims at maximizing the levels of connectivity and/or of bit parallelism that WRONoCs can achieve, while relating their upper bounds to the uncertainty of the manufacturing process. Andrea Peano, Luca Ramini, Marco Gavanelli, Maddalena Nonato, Davide Bertozzi |
ICCAD | 5 |
| 2016 | A built-in self-testing framework for asynchronous bundled-data NoC switches resilient to delay variationsabstractMost multi- and many-core integrated systems are currently designed by following a globally asynchronous locally synchronous paradigm. Asynchronous interconnection networks are promising candidates to interconnect IP cores operating at potentially different frequencies. Nevertheless, post-fabrication testing is a big challenge to bring asynchronous NoCs to the market due to a lack of testing methodologies and support for them. In particular, the unpredictable delay variability introduced by the manufacturing process may differentiate the delay of nominally-balanced I/O timing paths, thus making the order of the input patterns unpredictable and precluding the correct behaviour of signature-based test compactors. This paper tackles this challenge by proposing a testing framework for asynchronous NoCs which works effectively despite delay variations in and across timing paths of the NoC under test. Moreover, in order to mitigate the growing test application costs in modern ICs, we come up with a built-in self-testing infrastructure which automatically controls and delivers the outcome of the testing process without the intervention of an external automatic test equipment (ATE). Gabriele Miorandi, Alberto Celin, Michele Favalli, Davide Bertozzi |
NOCS | 4 |
| 2016 | Populating and exploring the design space of wavelength-routed optical network-on-chip topologies by leveraging the add-drop filtering primitiveabstractEmerging technologies often carry different logic primitives for which contemporary logic synthesis techniques are not suitable. For this reason, the design space of circuit-level solutions relying on such technologies is often largely unexplored. One domain where this trend is evident consists of topologies for wavelength-routed optical networks-on-chip (WRONoCs). Current literature only reports isolated design points that are inspired only by the intuition of researchers: there are currently no methodologies that enable designers to populate the design space in a consistent way. As a result, the possibility of an automated synthesis flow guiding the designer to the most promising solution in the design space is far from coming. This paper aims at bridging the former gap by identifying the basic primitive which is at the core of each topology. Then, a methodology is consistently derived to combine basic primitives together, thus potentially yielding all points of the topology design space. To our knowledge, this is the first time the design space of wavelength-routed topologies is populated in a potentially exhaustive way through a systematic methodology. As a result, it becomes evident that for a specified quality metric, there exist better solutions than the topologies that have been found out so far by designers' intuition. This paper is thus the stepping stone for future work targeting automatic synthesis of WRONoC topologies. Mahdi Tala, Marco Castellari, Marco Balboni, Davide Bertozzi |
NOCS | 4 |
| 2016 | PROTON+: A Placement and Routing Tool for 3D Optical Networks-on-Chip with a Single Optical LayerabstractOptical Networks-on-Chip (ONoCs) are a promising technology to overcome the bottleneck of low bandwidth of electronic Networks-on-Chip. Recent research discusses power and performance benefits of ONoCs based on their system-level design, while layout effects are typically overlooked. As a consequence, laser power requirements are inaccurately computed from the logic scheme but do not consider the layout. In this article, we propose PROTON+, a fast tool for placement and routing of 3D ONoCs minimizing the total laser power. Using our tool, the required laser power of the system can be decreased by up to 94% compared to a state-of-the-art manually designed layout. In addition, with the help of our tool, we study the physical design space of ONoC topologies. For this purpose, topology synthesis methods (e.g., global connectivity and network partitioning) as well as different objective function weights are analyzed in order to minimize the maximum insertion loss and ultimately the system’s laser power consumption. For the first time, we study optimal positions of memory controllers. A comparison of our algorithm to a state-of-the-art placer for electronic circuits shows the need for a different set of tools custom-tailored for the particular requirements of optical interconnects. Anja von Beuningen, Luca Ramini, Davide Bertozzi, Ulf Schlichtmann |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2015 | Synergistic use of multiple on-chip networks for ultra-low latency and scalable distributed routing reconfiguration
Marco Balboni, José Flich, Davide Bertozzi |
DATE | 3 |
| 2015 | SSDExplorer: A Virtual Platform for Performance/Reliability-Oriented Fine-Grained Design Space Exploration of Solid State DrivesabstractCurrently available electronic design automation tools for design space exploration of solid state drives (SSDs) are not able to assess: 1) the device architecture inefficiencies; 2) architecture overdesign for a target performance; and 3) performance degradation caused by the disk usage. These tools feature either an overly high abstraction modeling strategy or lack the required flexibility to perform design exploration. To overcome these problems, this paper proposes SSDExplorer, a tool for fine-grained yet reasonably fast design space exploration of different SSD architectures highlighting possible bottlenecks. To prove its accuracy SSDExplorer has been validated with two real SSDs. SSDExplorer efficiency has been assessed by evaluating the impact of the NAND flash read retry algorithm impact on the SSD performance as a function of its internal architecture. Lorenzo Zuolo, Cristian Zambelli, Rino Micheloni, Marco Indaco, Stefano Di Carlo, Paolo Prinetto, Davide Bertozzi, Piero Olivo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2015 | Performance and Reliability Analysis of Cross-Layer Optimizations of NAND Flash ControllersabstractNAND flash memories are becoming the predominant technology in the implementation of mass storage systems for both embedded and high-performance applications. However, when considering data and code storage in Non-Volatile Memories (NVMs), such as NAND flash memories, reliability and performance become a serious concern for systems designers. Designing NAND flash-based systems based on worst-case scenarios leads to waste of resources in terms of performance, power consumption, and storage capacity. This is clearly in contrast with the request for runtime reconfigurability, adaptivity, and resource optimization in modern computing systems. There is a clear trend toward supporting differentiated access modes in flash memory controllers, each one setting a differentiated tradeoff point in the performance-reliability optimization space. This is supported by the possibility of tuning the NAND flash memory performance, reliability, and power consumption through several tuning knobs such as the flash programming algorithm and the flash error correcting code. However, to successfully exploit these degrees of freedom, it is mandatory to clearly understand the effect that the combined tuning of these parameters has on the full NVM subsystem. This article performs a comprehensive quantitative analysis of the benefits provided by the runtime reconfigurability of an MLC NAND flash controller through the combined effect of an adaptable memory programming circuitry coupled with runtime adaptation of the ECC correction capability. The full NVM subsystem is taken into account, starting from a characterization of the low-level circuitry to the effect of the adaptation on a wide set of realistic benchmarks in order to provide readers a clear view of the benefit this combined adaptation may provide at the system level. Davide Bertozzi, Stefano Di Carlo, Salvatore Galfano, Marco Indaco, Piero Olivo, Paolo Prinetto, Cristian Zambelli |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | A vertically integrated and interoperable multi-vendor synthesis flow for predictable noc design in nanoscale technologiesabstractWe deliver a design flow for the synthesis and convergence of application-specific networks-on-chip. The flow comes with novel features that can better address nanoscale design challenges: front-end driven floorplanning, dynamic IR-drop minimization, fast and accurate system-level power grid modeling, predictable link design. Above all, such features are addressed by different prototype engines, even from different vendors, that can be smoothly integrated into the flow by means of a common specification format called Communication Exchange Format (CEF), that enables unprecedented tool interactions. This flow is validated by means of an extensive demonstration framework. Alberto Ghiribaldi, Hervé Tatenguem, Federico Angiolini, Mikkel Bystrup Stensgaard, Tobias Bjerregaard, Davide Bertozzi |
ASP-DAC | 6 |
| 2014 | Guided Participatory Research on Parallel Computer Architectures for K-12 Students Through a Narrative ApproachabstractThe approach to computer science (CS) education is typically geared towards the knowledge of the principles behind information technology, but there are social indicators that it overlooks some important educative aspects such as thinking competences and social attitudes. Such aspects play a fundamental role when bringing CS education to the K-12 level. In order to enable a truly educational experience, we propose to bring specific CS research problems within reach of K-12 students, because the active knowledge construction process that takes place during research requires children to be engaged with all of their knowledge, skills and attitudes. This poses the challenge of overcoming the knowledge gap of students, which we address by means of a synergistic cooperation of CS experts and educators. More specifically, we propose the narrative approach as the key enabler for CS participatory research with K-12 students. Valentina Mazzoni, Luigina Mortari, Federico Corni, Davide Bertozzi |
CSEDU (3) | 4 |
| 2014 | Assessing the energy break-even point between an optical NoC architecture and an aggressive electronic baselineabstractMany crossbenchmarking results reported in the open literature raise optimistic expectations on the use of optical networks-on-chip (ONoCs) for high-performance and low-power on-chip communication. However, most of those previous works ultimately fail to make a compelling case for chip-level nanophotonic NoCs, especially for the lack of aggressive electronic baselines (ENoC), and the poor accuracy in physical- and architecture-layer analysis of the ONoC. This paper aims at providing the guidelines and minimum requirements so that nanophotonic emerging technology may become of practical relevance. The key differentiating factor of this work consists of contrasting ONoC solutions with an aggressive ENoC architecture with realistic complexity, performance, and power figures, synthesized on an industrial 40nm low-power technology. At the same time, key physical design issues and network interface architecture requirements for the ONoC under test are carefully assessed, thus paving the way for a well-grounded definition of the requirements for the emerging ONoC technology to achieve the energy break-even point with respect to pure electronic interconnect solutions in future multi- and many-core systems. Luca Ramini, Alberto Ghiribaldi, Paolo Grani, Sandro Bartolini, Hervé Tatenguem, Davide Bertozzi |
DATE | 6 |
| 2014 | SSDExplorer: A virtual platform for fine-grained design space exploration of Solid State DrivesabstractSolid State Drives (SSDs) are gaining particular momentum in various frameworks such as multimedia, large data centers and cloud environments. Unfortunately, efficient CAD tools for SSD design space exploration able to assess the optimization of the device microarchitecture w.r.t. the target performance are still missing. This paper tries to close this gap by proposing SSDExplorer, a tool for fine-grained and fast design space exploration of SSD devices. SSDExplorer provides unprecedented insights into the architecture behavior and subcomponent interaction efficiency, while avoiding the need for the actual implementation of an FTL or of key hardware components. This is achieved by the introduction of suitable abstractions of the different components. This is confirmed by the thorough validation of SSDExplorer against a commercial SSD device. Lorenzo Zuolo, Cristian Zambelli, Rino Micheloni, Salvatore Galfano, Marco Indaco, Stefano Di Carlo, Paolo Prinetto, Piero Olivo, Davide Bertozzi |
DATE | 9 |
| 2014 | A complete electronic network interface architecture for global contention-free communication over emerging optical networks-on-chipabstractAlthough many valuable research works have investigated the properties of optical networks-on-chip (ONoCs), the vast majority of them lack an accurate exploration of the network interface architecture (NI) required to support optical communications on the silicon chip. The complexity of this architecture is especially critical for a specific kind of ONoCs: wavelength-routed ones. From a logical viewpoint, they can be considered as full nonblocking crossbars, thus the control complexity is implemented at the NIs. To our knowledge, this paper proposes the first complete NI architecture for wavelength-routed optical NoCs, by coping with the intricacy of networking issues such as flow control, buffering strategy, deadlock avoidance, serialization, and above all, with their codesign in a complete architecture. Marta Ortín-Obón, Luca Ramini, Hervé Tatenguem, Víctor Viñals, Davide Bertozzi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2014 | Augmenting manycore programmable accelerators with photonic interconnect technology for the high-end embedded computing domainabstractThere is today consensus on the fact that optical interconnects can relieve bandwidth density concerns at integrated circuit boundaries. However, when it comes to the extension of this emerging interconnect technology to on-chip communication as well, such consensus seems to fall apart. The main reason consists of a fundamental lack of compelling cases proving the superior performance and/or energy properties yielded by devices of practical interest, when re-architected around a photonically-integrated communication fabric. This paper takes its steps from the consideration that manycore computing platforms are gaining momentum in the high-end embedded computing domain in the form of general-purpose programmable accelerators. Hence, the performance and energy implications when augmenting these devices with optical interconnect technology are derived by means of an accurate benchmarking framework against an aggressively optimized electrical counterpart. Marco Balboni, Marta Ortín-Obón, Alessandro Capotondi, Hervé Tatenguem, Alberto Ghiribaldi, Luca Ramini, Víctor Viñals, Andrea Marongiu, Davide Bertozzi |
NOCS | 9 |
| 2014 | Crossbar replication vs. sharing for virtual channel flow control in asynchronous NoCs: A comparative studyabstractIn on-chip interconnection networks, performance optimization techniques can be often achieved in two opposite ways: by making control logic more complex inside switches, or by pushing design complexity to the switch boundaries. The implementation of virtual channel (VC) flow control is an important application domain of this design trade-off. The data path of VC switches typically exhibits replicated buffers. The underlying philosophy (i.e., resource replication) can be pushed to the limit, thus incuring an apparently high area cost, while simplifying the switch control path. On the other hand, unreplicated resources require complex control logic for the sake of their efficient sharing among virtual networks. Investigating this design tradeoff is especially important for asynchronous networks, where the synthesis of complex control circuits is a challenge. This paper is a first step toward a design space exploration of VC implementation techniques for transition-signalling bundled-data asynchronous NoCs, and contrasts a VC switch with replicated crossbars against a unified-crossbar architecture relying on multistage switch allocation. Gabriele Miorandi, Alberto Ghiribaldi, Steven M. Nowick, Davide Bertozzi |
VLSI-SoC | 4 |
| 2014 | Capturing the sensitivity of optical network quality metrics to its network interface parametersabstractSUMMARY Optical networks‐on‐chip (ONoCs) are gaining momentum as a way to improve energy consumption and bandwidth scalability in the next generation multicore and many‐core systems. Although many valuable research works have investigated their properties, the vast majority of them lack an accurate exploration of the network interface architecture required to support optical communications on the silicon chip. The complexity of this architecture is especially critical for a specific kind of ONoCs: the wavelength‐routed ones. These are capable of delivering contention‐free all‐to‐all connectivity without the need for path reservation, unlike space‐routed ONoCs. From a logical viewpoint, they can be considered as full nonblocking crossbars; thus, the control complexity is implemented at the network interfaces. To our knowledge, this paper proposes the first complete network interface architecture for wavelength‐routed optical NoCs, by coping with the intricacy of networking issues such as flow control, buffering strategy, deadlock avoidance, serialization, and above all, their codesign in a complete architecture. The evaluation methodology spans from area and energy analysis via actual synthesis runs in 40‐nm technology to RTL‐equivalent (register‐transfer level) SystemC modelling of the network architecture and aims at verifying whether the projected benefits of ONoCs versus their electrical counterparts are still preserved when the complexity of their network interface is considered in the analysis. Copyright © 2014 John Wiley & Sons, Ltd. Marta Ortín-Obón, Luca Ramini, Víctor Viñals, Davide Bertozzi |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | FLARES: An Aging Aware Algorithm to Autonomously Adapt the Error Correction Capability in NAND Flash MemoriesabstractWith the advent of solid-state storage systems, NAND flash memories are becoming a key storage technology. However, they suffer from serious reliability and endurance issues during the operating lifetime that can be handled by the use of appropriate error correction codes (ECCs) in order to reconstruct the information when needed. Adaptable ECCs may provide the flexibility to avoid worst-case reliability design, thus leading to improved performance. However, a way to control such adaptable ECCs' strength is required. This article proposes FLARES, an algorithm able to adapt the ECC correction capability of each page of a flash based on a flash RBER prediction model and on a measurement of the number of errors detected in a given time window. FLARES has been fully implemented within the YAFFS 2 filesystem under the Linux operating system. This allowed us to perform an extensive set of simulations on a set of standard benchmarks that highlighted the benefit of FLARES on the overall storage subsystem performances. Stefano Di Carlo, Salvatore Galfano, Marco Indaco, Paolo Prinetto, Davide Bertozzi, Piero Olivo, Cristian Zambelli |
ACM Trans. Archit. Code Optim. | 5 |
| 2013 | A transition-signaling bundled data NoC switch architecture for cost-effective GALS multicore systemsabstractAsynchronous networks-on-chip (NoCs) are an appealing solution to tackle the synchronization challenge in modern multicore systems through the implementation of a GALS paradigm. However, they have found only limited applicability so far due to two main reasons: the lack of proper design tool flows as well as their significant area footprint over their synchronous counterparts. This paper proposes a largely unexplored design point for asynchronous NoCs, relying on transition-signaling bundled data, which contributes to break the above barriers. Compared to an existing lightweight synchronous switch architecture, xpipesLite, the post-layout asynchronous switch achieved a 71% reduction in area, up to 85% reduction in overall power consumption, and a 44% average reduction in energy-per-flit, while mastering the more stringent timing assumptions of this solution with a semi-automated synthesis flow. Alberto Ghiribaldi, Davide Bertozzi, Steven M. Nowick |
DATE | 2 |
| 2013 | Contrasting wavelength-routed optical NoC topologies for power-efficient 3D-stacked multicore processors using physical-layer analysisabstractOptical networks-on-chip (ONoCs) are currently still in the concept stage, and would benefit from explorative studies capable of bridging the gap between abstract analysis frameworks and the constraints and challenges posed by the physical layer. This paper aims to go beyond the traditional comparison of wavelength-routed ONoC topologies based only on their abstract properties, and for the first time assesses their physical implementation efficiency in an homogeneous experimental setting of practical relevance. As a result, the paper can demonstrate the significant and different deviation of topology layouts from their logic schemes under the effect of placement constraints on the target system. This becomes then the preliminary step for the accurate characterization of technology-specific metrics such as the insertion loss critical path, and to derive the ultimate impact on power efficiency and feasibility of each design. Luca Ramini, Paolo Grani, Sandro Bartolini, Davide Bertozzi |
DATE | 4 |
| 2013 | Topic 13: High-Performance Networks and Communication - (Introduction)
Olav Lysne, Torsten Hoefler, Pedro López 0001, Davide Bertozzi |
Euro-Par | 4 |
| 2013 | PROTON: an automatic place-and-route tool for optical networks-on-chipabstractOptical Networks-on-Chip (ONoCs) are considered a promising way of improving power and bandwidth limitations in next generation multi- and many-core integrated systems. Today, most related research acknowledges the key role of the physical layer in assessing ONoC topologies (e.g., insertion loss), but overlooks the placement and routing stage in the design process, hence applying physical design considerations to topology logic schemes. Such a mismatch is fundamentally due to the lack of mature CAD tools for placement and routing of optical NoCs. The objective of this work is to bridge this gap: We propose PROTON, a fast tool for automatic placement and routing of ONoC topologies, which can support designers in quantifying the degradation of design quality metrics when moving from topology logic schemes to their physical implementation. This gap is especially relevant for Wavelength-Routed ONoCs (WRONoCs), where logic schemes typically make unrealistic assumptions about the placement of initiators and targets. For this reason, we put PROTON to work with the most promising WRONoC topologies and explore their physical design space given the placement and routing constraints of a 3D stacked system. We also compare automatically generated layouts with handcrafted ones reported in the literature for the same topologies and target system, and prove an insertion loss improvement by up to 150x. With PROTON the exploration of the physical design space of ONoC topologies is possible as well as their scalability analysis considering the layout. Anja Boos, Luca Ramini, Ulf Schlichtmann, Davide Bertozzi |
ICCAD | 4 |
| 2013 | A complete self-testing and self-configuring NoC infrastructure for cost-effective MPSoCsabstractNetworks-on-chip need to survive to manufacturing faults in order to sustain yield. An effective testing and configuration strategy however implies two opposite requirements. One one hand, a fast and scalable built-in self-testing and self-diagnosis procedure has to be carried out concurrently at NoC switches. On the other hand, programming the NoC routing mechanism to go around faulty links and switches can be optimally performed by a centralized controller with global network visibility. To the best of our knowledge, this article proposes for the first time a global network testing and configuration strategy that meets the opposite requirements by means of a fault-tolerant dual network architecture and a fast configuration algorithm for the most common failure patterns. Experimental results report an area overhead as low as 12.5% with respect to the baseline switch architecture while achieving a high degree of fault tolerance. In fact, even when multiple stuck-at faults are considered, the capability of fault masking by the dual network is always over 80%, and the support for multiple link failures is more than 90% in presence of two unusable links in the main network with minimum set-up times. Alberto Ghiribaldi, Daniele Ludovici, Francisco Triviño, Alessandro Strano, José Flich, José L. Sánchez 0002, Francisco J. Alfaro, Michele Favalli, Davide Bertozzi |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2013 | An efficient, low-cost routing framework for convex mesh partitions to support virtualizationabstractAt the core of an efficient chip multiprocessors (CMP) is support for unicast and multicast routing, low implementation costs, and the ability to isolate concurrent applications with maximum utilization of the CMP. We present an efficient logic-based unicast and multicast routing algorithm that guarantees isolation of local application traffic within any near-convex region on the chip, and the algorithms to recognize supported partitions and configure the cores accordingly. Evaluations show that the routing algorithm has a 57% more compact implementation than a recent multicast solution with the same coverage, and it achieves 5% higher throughput with 13% lower latency. Frank Olaf Sem-Jacobsen, Samuel Rodrigo, Tor Skeie, Alessandro Strano, Davide Bertozzi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | Enabling power efficiency through dynamic rerouting on-chipabstractNetworks-on-chip (NoCs) are key components in many-core chip designs. Dynamic power-awareness is a new challenge present in NoCs that must be efficiently handled by the routing functionality as it introduces irregularities in the commonly used 2-D meshes. In this article, we propose a logic-based routing algorithm, iFDOR, oriented towards dynamic powering down one region within every application partition on the chip through dynamic rerouting, with low implementation costs. Results show that we can successfully shutdown an arbitrary rectangular region within an application partition without significant impact on network performance. Frank Olaf Sem-Jacobsen, Samuel Rodrigo, Alessandro Strano, Tor Skeie, Davide Bertozzi, Francisco Gilabert Villamón |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2012 | Design of a collective communication infrastructure for barrier synchronization in cluster-based nanoscale MPSoCsabstractBarrier synchronization is a key programming primitive for shared memory embedded MPSoCs. As the core count increases, software implementations cannot provide the needed performance and scalability, thus making hardware acceleration critical. In this paper we describe an interconnect extension implemented with standard cells and with a mainstream industrial toolflow. We show that the area overhead is marginal with respect to the performance improvements of the resulting hardware-accelerated barriers. We integrate our HW barrier into the OpenMP programming model and discuss synchronization efficiency compared with traditional software implementations. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, Davide Bertozzi, Daniele Bortolotti, Andrea Marongiu, Luca Benini |
DATE | 4 |
| 2012 | A cross-layer approach for new reliability-performance trade-offs in MLC NAND flash memoriesabstractIn spite of the mature cell structure, the memory controller architecture of Multi-level cell (MLC) NAND Flash memories is evolving fast in an attempt to improve the uncorrected/miscorrected bit error rate (UBER) and to provide a more flexible usage model where the performance-reliability trade-off point can be adjusted at runtime. However, optimization techniques in the memory controller architecture cannot avoid a strict trade-off between UBER and read throughput. In this paper, we show that co-optimizing ECC architecture configuration in the memory controller with program algorithm selection at the technology layer, a more flexible memory sub-system arises, which is capable of unprecedented trade-offs points between performance and reliability. Cristian Zambelli, Marco Indaco, Michele Fabiano, Stefano Di Carlo, Paolo Prinetto, Piero Olivo, Davide Bertozzi |
DATE | 7 |
| 2012 | Cost-Effective Contention Avoidance in a CMP with Shared Memory Controllers
Samuel Rodrigo, Frank Olaf Sem-Jacobsen, Hervé Tatenguem, Tor Skeie, Davide Bertozzi |
Euro-Par | 5 |
| 2012 | A retrospective look at xpipes: The exciting ride from a design experience to a design platform for nanoscale networks-on-chipabstractThis paper gives a retrospective look at the xpipes framework, and documents its evolution from a promising network-on-chip (NoC) design experience to a comprehensive design platform for the next-generation of nanoscale NoCs. Since the early days of xpipes, its cross-layer approach to NoC design has given a significant contribution to bridge the gap between the NoC concept and an industry-relevant interconnect technology. Davide Bertozzi, Luca Benini |
ICCD | 1 |
| 2012 | Xpipes: A latency insensitive parameterized network-on-chip architecture for multi-processor SoCsabstractThe growing complexity of customizable embedded multi-processor architectures for digital media processing will soon require highly scalable network-on-chip based communication infrastructures. In this paper, we propose xpipes, a scalable and high-performance NoC architecture for multi-processor SoCs, consisting of soft macros that can be turned into instance-specific network components at instantiation time. The flexibility of its components allows our NoC to support both homogeneous and heterogeneous architectures. The interface with IP cores at the periphery of the network is standardized (OCP-based). Links can be pipelined with a flexible number of stages to decouple data introduction speed from worst-case link delay. Switches are lightweight and support reliable communication for arbitrary link pipeline depths (latency insensitive operation). xpipes has been described in synthesizable SystemC, at the cycle-accurate and signal-accurate level. Matteo Dall'Osso, Gianluca Biccari, Luca Giovannini, Davide Bertozzi, Luca Benini |
ICCD | 4 |
| 2012 | Engineering a Bandwidth-Scalable Optical Layer for a 3D Multi-core Processor with Awareness of Layout ConstraintsabstractThe performance of future chip multi-processors will only scale with the number of integrated cores if there is a corresponding increase in memory access efficiency. The focus of this paper on a 3D-stacked wavelength-routed optical layer for high bandwidth and low latency processor-memory communication goes in this direction and complements ongoing efforts on photonically integrated bandwidth-rich DRAM devices. This target environment dictates layout constraints that make the difference in discriminating between alternative design choices of the optical layer. This paper assesses network partitioning options and bandwidth scalability techniques with deep technology and layout awareness, the main contribution lying in the characterization and precise quantification of such interaction effects between the technology platform, the layout constraints and the network-level quality metrics of a passive optical NoC. Luca Ramini, Davide Bertozzi, Luca P. Carloni |
NOCS | 2 |
| 2011 | Exploiting structural redundancy of SIMD accelerators for their built-in self-testing/diagnosis and reconfigurationabstractProcess scaling has given designers billions of transistors to work with. As feature sizes near the atomic scale, extensive variation and wear-out inevitably make margining uneconomical or impossible. In this context, new design approaches are required. The inherent regularity and redundancy of SIMD architectures make them suitable to address the challenges posed by new semiconductor technologies at the architecture level. This paper proposes a built-in self-test/self-diagnosis procedure for a SIMD processor. Concurrent BIST operations are carried out after reset at each PE, thus resulting in scalable test application time with processor size. The key principle consists of exploiting the inherent structural redundancy of the SIMD architecture in a cooperative way, thus strongly reducing the testing framework latency and area overhead. Once the faults are detected, a reconfiguration technique is then proposed in order to preserve correct operation. Testing of single stuck-at faults is performed at-speed in 240 cycles regardless of the accelerator size, with a hardware overhead of less than 10%. Finally, the fault-tolerant tile integrating both BIST, reconfiguration logic and spare PE requires a 25% of total area overhead. Alessandro Strano, Davide Bertozzi, Arnaud Grasset, Sami Yehia |
ASAP | 2 |
| 2011 | Exploiting Network-on-Chip structural redundancy for a cooperative and scalable built-in self-test architectureabstractThis paper proposes a built-in self-test/self-diagnosis procedure at start-up of an on-chip network (NoC). Concurrent BIST operations are carried out after reset at each switch, thus resulting in scalable test application time with network size. The key principle consists of exploiting the inherent structural redundancy of the NoC architecture in a cooperative way, thus detecting faults in test pattern generators too. At-speed testing of stuck-at faults can be performed in less than 1200 cycles regardless of their size, with an hardware overhead of less than 11%. Alessandro Strano, Crispín Gómez Requena, Daniele Ludovici, Michele Favalli, María Engracia Gómez, Davide Bertozzi |
DATE | 6 |
| 2011 | System-level infrastructure for boot-time testing and configuration of networks-on-chip with programmable routing logicabstractNetworks-on-chip need to survive to manufacturing faults in order to sustain yield. An effective testing and configuration strategy however implies two opposite requirements. On one hand, a fast and scalable built-in self-testing and self-diagnosis procedure has to be carried out concurrently at NoC switches. On the other hand, programming the NoC routing mechanism to go around faulty links and switches can be optimally performed by a centralized controller with global network visibility. This paper proposes a global hardware infrastructure that meets such requirements by means of a fault-tolerant dual network architecture and a configuration strategy for reprogramming the routing mechanism of each switch. This is the first complete infrastructure for testing and reconfiguring a NoC based on reprogrammable routing logic. Alberto Ghiribaldi, Daniele Ludovici, Michele Favalli, Davide Bertozzi |
VLSI-SoC | 4 |
| 2011 | Cost-Efficient On-Chip Routing Implementations for CMP and MPSoC SystemsabstractThe high-performance computing domain is enriching with the inclusion of networks-on-chip (NoCs) as a key component of many-core (CMPs or MPSoCs) architectures. NoCs face the communication scalability challenge while meeting tight power, area, and latency constraints. Designers must address new challenges that were not present before. Defective components, the enhancement of application-level parallelism, or power-aware techniques may break topology regularity, thus, efficient routing becomes a challenge. This paper presents universal logic-based distributed routing (uLBDR), an efficient logic-based mechanism that adapts to any irregular topology derived from 2-D meshes, instead of using routing tables. uLBDR requires a small set of configuration bits, thus being more practical than large routing tables implemented in memories. Several implementations of uLBDR are presented highlighting the tradeoff between routing cost and coverage. The alternatives span from the previously proposed LBDR approach (with 30% of coverage) to the uLBDR mechanism achieving full coverage. This comes with a small performance cost, thus exhibiting the tradeoff between fault tolerance and performance. Power consumption, area, and delay estimates are also provided highlighting the efficiency of the mechanism. To do this, different router models (one for CMPs and one for MPSoCs) have been designed as a proof concept. Samuel Rodrigo, José Flich, Antoni Roca 0001, Simone Medardoni, Davide Bertozzi, Jesús Camacho Villanueva, Federico Silla, José Duato |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2010 | Design space exploration of a mesochronous link for cost-effective and flexible GALS NOCsabstractThere is today little doubt on the fact that a high-performance and cost-effective Network-on-Chip can only be designed in 45nm and beyond under a relaxed synchronization assumption. In this direction, this paper focuses on a GALS system where the NoC and its end-nodes have independent clocks (unrelated in frequency and phase) and are synchronized via dual-clock FIFOs at network interfaces. Within the network, we assume mesochronous synchronization implemented with hierarchical clock tree distribution. This paper contributes two essential components of any practical design automation support for network instantiation in the target system. On one hand, it introduces a switch design which greatly reduces the overhead for mesochronous synchronization and can be adapted to meet different layout constraints. On the other hand, the paper illustrates a design space exploration framework of mesochronous links that can direct the selection of synchronization options on a port-by-port basis for all the switches in the NoC, based on timing and layout constraints. A final case study illustrates how a cost-effective GALS NoC can be assembled, placed and routed by exploiting the flexibility of the architecture and the outcomes of the exploration framework, thus proving the viability and effectiveness of the design platform. Daniele Ludovici, Alessandro Strano, Georgi Gaydadjiev, Luca Benini, Davide Bertozzi |
DATE | 5 |
| 2010 | Addressing Manufacturing Challenges with Cost-Efficient Fault Tolerant RoutingabstractThe high-performance computing domain is enriching with the inclusion of Networks-on-chip (NoCs) as a key component of many-core (CMPs or MPSoCs) architectures. NoCs face the communication scalability challenge while meeting tight power, area and latency constraints. Designers must address new challenges that were not present before. Defective components, the enhancement of application-level parallelism or power-aware techniques may break topology regularity, thus, efficient routing becomes a challenge.In this paper, uLBDR (Universal Logic-Based Distributed Routing) is proposed as an efficient logic-based mechanism that adapts to any irregular topology derived from 2D meshes, being an alternative to the use of routing tables (either at routers or at end-nodes). uLBDR requires a small set of configuration bits, thus being more practical than large routing tables implemented in memories. Several implementations of uLBDR are presented highlighting the trade-off between routing cost and coverage. The alternatives span from the previously proposed LBDR approach (with 30\% of coverage) to the uLBDR mechanism achieving full coverage. This comes with a small performance cost, thus exhibiting the trade-off between fault tolerance and performance. Samuel Rodrigo, José Flich, Antoni Roca 0001, Simone Medardoni, Davide Bertozzi, Jesús Camacho Villanueva, Federico Silla, José Duato |
NOCS | 5 |
| 2010 | Improved Utilization of NoC Channel Bandwidth by Switch Replication for Cost-Effective Multi-processor Systems-on-ChipabstractVirtual channels are an appealing flow control technique for on-chip interconnection networks (NoCs), in that they can potentially avoid deadlock and improve link utilization and network throughput. However, their use in the resource constrained multi-processor system-on-chip (MPSoC) domain is still controversial, due to their significant overhead in terms of area, power and cycle time degradation. This paper proposes a simple yet efficient approach to VC implementation, which results in more area- and power-saving solutions than conventional design techniques. While these latter replicate only buffering resources for each physical link, we replicate the entire switch and prove that our solution is counter intuitively more area/power efficient while potentially operating at higher speeds. This result builds on a well-known principle of logic synthesis for combinational circuits (the area-performance trade-off when inferring a logic function into a gate-level netlist), and proves that when a designer is aware of this, novel architecture design techniques can be conceived. Francisco Gilabert Villamón, María Engracia Gómez, Simone Medardoni, Davide Bertozzi |
NOCS | 4 |
| 2009 | Designing Regular Network-on-Chip Topologies under Technology, Architecture and Software ConstraintsabstractRegular multi-core processors are appearing in the embedded system market as high performance software programmable solutions. The use of regular interconnect fabrics for them allows fast design time, ease of routing, predictability of electrical parameters and good scalability. k-ary n-mesh topologies are candidate solutions for these systems, borrowed from the domain of off-chip interconnection networks. However, the on-chip integration has to deal with unique challenges at different levels of abstraction. From a technology viewpoint, interconnect reverse scaling causes critical paths to go across global links. Poor interconnect performance might also impact IP core speed depending on the synchronization mechanism at the interface. Finally, this might also conflict with the requirements that communication libraries employed in the MPSoC domain pose on the underlying interconnect fabric. This paper provides a comprehensive overview of these topics, by characterizing physical feasibility of representative k-ary n-mesh topologies and by providing silicon-aware system-level performance figures. Francisco Gilabert Villamón, Daniele Ludovici, Simone Medardoni, Davide Bertozzi, Luca Benini, Georgi Gaydadjiev |
CISIS | 4 |
| 2009 | Assessing fat-tree topologies for regular network-on-chip design under nanoscale technology constraintsabstractMost of past evaluations of fat-trees for on-chip interconnection networks rely on oversimplifying or even irrealistic architecture and traffic pattern assumptions, and very few layout analyses are available to relieve practical feasibility concerns in nanoscale technologies. This work aims at providing an in-depth assessment of physical synthesis efficiency of fat-trees and at extrapolating silicon-aware performance figures to back-annotate in the system-level performance analysis. A 2D mesh is used as a reference architecture for comparison, and a 65 nm technology is targeted by our study. Finally, in an attempt to mitigate the implementation cost of k-ary n-tree topologies, we also review an alternative unidirectional multi-stage interconnection network which is able to simplify the fat-tree architecture and to minimally impact performance. Daniele Ludovici, Francisco Gilabert Villamón, Simone Medardoni, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, Georgi Gaydadjiev, Davide Bertozzi |
DATE | 8 |
| 2009 | Effectiveness of adaptive supply voltage and body bias as post-silicon variability compensation techniques for full-swing and low-swing on-chip communication channelsabstractAdaptive body bias (ABB) and adaptive supply voltage (ASV) have been showed to be effective methods for post-silicon tuning of circuit properties to reduce variability. While their properties have been compared on generic combinational circuits or microprocessor circuit sub-blocks, the advent of multi-core systems is bringing a new application domain forefront. Global interconnects are evolving to complex communication channels with drivers and receivers, in an attempt to mitigate the effects of reverse scaling and reduce power. The characterization of the performance spread of these links and the exploration of effective and power-aware compensation techniques for them is becoming a key design issue. This work compares the variability compensation efficiency of ABB vs ASV when put at work in two representative link architectures of today's ICs: a traditional full-swing interconnect and a low-swing signaling scheme for low-power communication. We provide guidelines for the post-silicon variability compensation of these communication channels. Giacomo Paci, Davide Bertozzi, Luca Benini |
DATE | 2 |
| 2009 | Capturing topology-level implications of link synthesis techniques for nanoscale networks-on-chipabstractIn the context of nanoscale networks-on-chip (NoCs), each link implementation solution is not just a specific synthesis optimization technique with local performance and power implications, but gives rise to a well-differentiated point in the architecture design space. This in an effect of the tight interaction existing between architecture and physical design layers in nanoscale technologies. Daniele Ludovici, Georgi Gaydadjiev, Davide Bertozzi, Luca Benini |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Comparing tightly and loosely coupled mesochronous synchronizers in a NoC switch architectureabstractWith the advent of networks-on-chip (NoCs), the interest for mesochronous synchronizers is again on the rise due to the intricacies of skew-controlled chip-wide clock tree distribution. Recently proposed schemes agree on a source synchronous design style with some form of ping-pong buffering to counter timing and metastability concerns. However, the integration issues of such synchronizers in a NoC setting are still largely uncovered. Most schemes are in fact placed between communicating switches, thus neglecting the abrupt increase of buffering resources needed at switch input stages. This paper goes a step forward and aims at deep integration of the synchronizer in the switch architecture, thus merging key tasks such as synchronization, buffering and flow control into a unique architecture block. This paper compares the integrated and the loosely coupled solutions from a performance and area viewpoint, while devoting special attention to their robustness with respect to physical design parameters. Daniele Ludovici, Alessandro Strano, Davide Bertozzi, Luca Benini, Georgi Gaydadjiev |
NOCS | 3 |
| 2009 | Reducing the Abstraction and Optimality Gaps in the Allocation and Scheduling for Variable Voltage/Frequency MPSoC PlatformsabstractThis paper proposes a novel approach to solve the allocation and scheduling problems for variable voltage/frequency multiprocessor systems-on-chip, which minimizes overall system energy dissipation. The optimality of derived system configurations is guaranteed, while the computation efficiency of the optimizer allows for solving problem instances that were traditionally considered beyond reach for exact solvers (optimality gap). Furthermore, this paper illustrates the development- and run-time software infrastructures that assist the user in developing applications and implementing optimizer solutions. The proposed approach guarantees a high level of power, performance, and constraint satisfaction predictability as from validation on the target platform, thus bridging the abstraction gap. Martino Ruggiero, Davide Bertozzi, Luca Benini, Michela Milano, Alexandru Andrei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Variation tolerant NoC design by means of self-calibrating linksabstractWe present the implementation and analysis of a variation tolerant version of a switch-to-switch link in a NoC. The goal is to tolerate the effects of process variations on NoC architectures using self-correcting links that automatically detect delay variations and compensate them. The correction is applied without increasing the switch-to-switch latency by substituting the output flip-flops of the sending switch with a self-correcting flip-flop followed by an adaptive voltage swing selector. Higher delay variations will result in a smaller slack in the switch-to-switch path, but the adaptive voltage swing selector could mitigate its impact on the NoC communication by increasing the voltage swing on the link, thus allowing a compensation of the delay variation. As a result, it is possible to tolerate delay variations at the cost of additional power consumption. Simone Medardoni, Marcello Lajolo, Davide Bertozzi |
DATE | 3 |
| 2008 | Process Variation Tolerant Pipeline Design Through a Placement-Aware Multiple Voltage Island Design StyleabstractA common technique to compensate process variation induced performance deviations during post-silicon testing consists of the dynamic adaptation of processor voltage. This however comes at a significant power cost. We envision multi supply voltage design (MSV) as a promising technique to mitigate such power overhead. Voltage islands are widely recognized as the state-of-the-art in MSV design. In this paper, we develop a novel design methodology that leverages voltage islands to compensate process variations through a commercial synthesis flow. Possible violation scenarios of performance requirements in fabricated chips are pre-characterized at design time through statistical static timing analysis. Then, during post-silicon testing the supply voltage of a proper number of voltage islands is raised depending on the actual violation scenario, thus bringing performance back within nominal values. Voltage islands are generated by exploiting cell proximity for minimal perturbation of performance pre-optimized placements. Bonesi Stefano, Davide Bertozzi, Luca Benini, Enrico Macii |
DATE | 2 |
| 2008 | Network Interface Sharing Techniques for Area Optimized NoC ArchitecturesabstractAlthough preliminary analysis frameworks point out the performance speed-ups achievable by on-chip networks with respect to state-of-the-art interconnects, the area concern remains one of the most daunting challenges to make this interconnect technology mainstream. A common approach to relieve the problem consists of sharing most of network interface resources among a number of processor cores. However, buffering resources need to be replicated and control logic reaches a complexity that limits maximum achievable frequency. This paper proposes full sharing of network interface resources, including buffers, thus trading performance for area. While area improvements are significant, a number of physical and system-level effects might mitigate performance degradation, making our technique a promising solution for area efficient network-on-chip realizations across a range of operating conditions. Alberto Ferrante, Simone Medardoni, Davide Bertozzi |
DSD | 3 |
| 2008 | Resource Management Policy Handling Multiple Use-Cases in MPSoC Platforms Using Constraint Programming
Luca Benini, Davide Bertozzi, Michela Milano |
ICLP | 2 |
| 2008 | Exploring High-Dimensional Topologies for NoC Design Through an Integrated Analysis and Synthesis Framework
Francisco Gilabert Villamón, Simone Medardoni, Davide Bertozzi, Luca Benini, María Engracia Gómez, Pedro López 0001, José Duato |
NOCS | 3 |
| 2008 | A multiprocessor system-on-chip for real-time biomedical monitoring and analysis: ECG prototype architectural design space explorationabstractIn this article we focus on multiprocessor system-on-chip (MPSoC) architectures for human heart electrocardiogram (ECG) real time analysis as a hardware/software (HW/SW) platform offering an advance relative to state-of-the-art solutions. This is a relevant biomedical application with good potential market, since heart diseases are responsible for the largest number of yearly deaths. Hence, it is a good target for an application-specific system-on-chip (SoC) and HW/SW codesign. We investigate a symmetric multiprocessor architecture based on STMicroelectronics VLIW DSPs that process in real time 12-lead ECG signals. This architecture improves upon state-of-the-art SoC designs for ECG analysis in its ability to analyze the full 12 leads in real time, even with high sampling frequencies, and its ability to detect heart malfunction for the whole ECG signal interval. We explore the design space by considering a number of hardware and software architectural options. Comparing our design with present-day solutions from an SoC and application point-of-view shows that our platform can be used in real time and without failures. Iyad Al Khatib, Francesco Poletti, Davide Bertozzi, Luca Benini, Mohamed Bechara, Hasan Khalifeh, Axel Jantsch, Rustam Nabiev |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | Interactive presentation: Capturing the interaction of the communication, memory and I/O subsystems in memory-centric industrial MPSoC platformsabstractIndustrial MPSoC platforms exhibit increasing communication needs while not yet reverting to revolutionary solutions such as networks-on-chip. On one hand, the limited scalability of shared busses is being overcome by means of multi-layer communication architectures, which are stressing the role of bridges as key contributors to system performance. On the other hand, technology limitations, data footprint and cost constraints lead to platform instantiations with only few on-chip memory devices and with a global performance bottleneck: the memory controller for access to the off-chip SDRAM memory. The complex interaction among system components and the dependency of macroscopic performance metrics on fine-grain architectural features stress the importance of highly accurate modelling and analysis tools. This paper takes its steps from an extensive modelling effort of a complete industrial MPSoC platform for consumer electronics, including the off-chip memory sub-system. Based on this, relevant design issues concerning the communication, memory and I/O architecture and their interaction are addressed, resulting in guidelines for designers of industry-relevant MPSoCs Simone Medardoni, Martino Ruggiero, Davide Bertozzi, Luca Benini, Giovanni Strano, Carlo Pistritto |
DATE | 3 |
| 2007 | Power-optimal RTL arithmetic unit soft-macro selection strategy for leakage-sensitive technologiesabstractWith the advent of nanoscale technologies, developing power efficient ASICs increasingly requires consideration of static power. An effective approach to make RTL synthesis algorithms and tools leakage-aware consists of the smart inference of RTL macros based on design constraints and optimization directives. This involves exploring the new trade-offs spanned by the design of RTL functional units, as an effect of the features of nanoscale technologies and ofthe power optimizations performed by commercial synthesis tools. This work explores these new trade-offs and proves that making RTL macro selection strategies aware of them results in power savings as high as 43%. Simone Medardoni, Davide Bertozzi, Enrico Macii |
ISLPED | 2 |
| 2007 | Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural SupportabstractIn today's multiprocessor SoCs (MPSoCs), parallel programming models are needed to fully exploit hardware capabilities and to achieve the 100 Gops/W energy efficiency target required for ambient intelligence applications. However, mapping abstract programming models onto tightly power-constrained hardware architectures imposes overheads which might seriously compromise performance and energy efficiency. The objective of this work is to perform a comparative analysis of message passing versus shared memory as programming models for single-chip multiprocessor platforms. Our analysis is carried out from a hardware-software viewpoint: we carefully tune hardware architectures and software libraries for each programming model. We analyze representative application kernels from the multimedia domain, and identify application-level parameters that heavily influence performance and energy efficiency. Then, we formulate guidelines for the selection of the most appropriate programming model and its architectural support Francesco Poletti, Antonio Poggiali, Davide Bertozzi, Luca Benini, Paul Marchal, Mirko Loghi, Massimo Poncino |
IEEE Trans. Computers | 3 |
| 2006 | Allocation, Scheduling and Voltage Scaling on Energy Aware MPSoCs
Luca Benini, Davide Bertozzi, Alessio Guerri, Michela Milano |
CPAIOR | 2 |
| 2006 | A multiprocessor system-on-chip for real-time biomedical monitoring and analysis: architectural design space explorationabstractIn this paper we focus on MPSoC architectures for human heart ECG real-time monitoring and analysis. This is a very relevant bio-medical application, with a huge potential market, hence it is an ideal target for an application-specific SoC implementation. We investigate a symmetric multi-processor architecture based on STMicroelectronics VLIW DSPs that process in real-time 12-lead ECG signals. This architecture improves upon state-of-the-art SoC designs for ECG analysis in its ability to analyze the full 12 leads in real-time, even with high sampling frequencies, and ability to detect heart malfunction. We explore the design space by considering a number of hardware and software architectural options. Iyad Al Khatib, Francesco Poletti, Davide Bertozzi, Luca Benini, Mohamed Bechara, Hasan Khalifeh, Axel Jantsch, Rustam Nabiev |
DAC | 3 |
| 2006 | Supporting task migration in multi-processor systems-on-chip: a feasibility studyabstractWith the advent of multi-processor systems-on-chip, the interest in process migration is again on the rise both in research and in product development. New challenges associated with the new scenario include increased sensitivity to implementation complexity, tight power budgets, requirements on execution predictability, the lack of virtual memory support in many low-end MPSoCs. As a consequence, effectiveness and applicability of traditional transparent migration mechanisms are put in discussion in this context. Our paper proposes a task management software infrastructure that is well suited for the constraints of single chip multiprocessors with distributed operating systems. Load balancing in the system is maintained by means of intelligent initial placement and task migration. We propose a user-managed migration scheme based on code checkpointing and user-level middleware support as an effective solution for many MPSoC application domains. In order to prove the practical viability of this scheme, we also propose a characterization methodology for task migration overhead. We derive the minimum execution time following a task migration event during which the system configuration should be frozen to make up for the migration cost. Stefano Bertozzi, Andrea Acquaviva, Davide Bertozzi, Antonio Poggiali |
DATE | 3 |
| 2006 | Communication-aware allocation and scheduling framework for stream-oriented multi-processor systems-on-chipabstractThis paper proposes a complete allocation and scheduling framework, where an MPSoC virtual platform is used to accurately derive input parameters, validate abstract models of system components and assess constraint satisfaction and objective function optimization. The optimizer implements an efficient and exact approach to allocation and scheduling based on problem decomposition. The allocation subproblem is solved through integer programming while the scheduling one through constraint programming. The two solvers can interact by means of no-good generation, thus building an iterative procedure which has been proven to converge to the optimal solution. Experimental results show significant speedups w.r.t. pure IP and CP exact solution strategies as well as high accuracy with respect to cycle accurate functional simulation. A case study further demonstrates the practical viability of our framework for real-life systems and applications. Martino Ruggiero, Alessio Guerri, Davide Bertozzi, Francesco Poletti, Michela Milano |
DATE | 3 |
| 2005 | Allocation and Scheduling for MPSoCs via Decomposition and No-Good Generation
Luca Benini, Davide Bertozzi, Alessio Guerri, Michela Milano |
CP | 2 |
| 2005 | xpipes Lite: A Synthesis Oriented Design Library For Networks on ChipsabstractThe limited scalability of current bus topologies for systems on chips (SoCs) dictates the adoption of networks on chips (NoCs) as a scalable interconnection scheme. Current SoCs are highly heterogeneous in nature, denoting homogeneous, preconfigured NoCs as inefficient drop-in alternatives. While highly parametric, fully synthesizeable (soft) NoC building blocks appear as a good match for heterogeneous MPSoC architectures, the impact of instantiation-time flexibility on performance, power and silicon cost has not yet been quantified. The paper details /spl times/pipes Lite, a design flow for automatic generation of heterogeneous NoCs. /spl times/pipes Lite is based on highly customizable, high frequency and low latency NoC modules, that are fully synthesizeable. Synthesis results provide modules that are directly comparable, if not better, than the current published state-of-the-art NoCs in terms of area, power latency and target operating frequency measurements. Stergios Stergiou, Federico Angiolini, Salvatore Carta, Luigi Raffo, Davide Bertozzi, Giovanni De Micheli |
DATE | 5 |
| 2005 | Application-Specific Power-Aware Workload Allocation for Voltage Scalable MPSoC PlatformsabstractIn this paper, we address the problem of selecting the optimal number of processing cores and their operating voltage/frequency for a given workload, to minimize overall system power under application-dependent QoS constraints. Selecting the optimal system configuration is non-trivial, since it depends on task characteristics and system-level interaction effects among the cores. For this reason, our QoS-driven methodology for power aware partitioning and frequency selection is based on functional, cycle-accurate simulation on a virtual platform environment. The methodology, being application-specific, is demonstrated on the DES (data encryption system) algorithm, representative of a wider class of streaming applications with independent input data frames and regular work-load. Martino Ruggiero, Andrea Acquaviva, Davide Bertozzi, Luca Benini |
ICCD | 3 |
| 2005 | Allocation and Scheduling for MPSoCs via decomposition and no-good generation
Luca Benini, Davide Bertozzi, Alessio Guerri, Michela Milano |
IJCAI | 2 |
| 2005 | Error control schemes for on-chip communication links: the energy-reliability tradeoffabstractOn-chip interconnection networks for future systems on chip (SoC) will have to deal with the increasing sensitivity of global wires to noise sources such as crosstalk or power supply noise. Hence, transient delay and logic faults are likely to reduce the reliability of across-chip communication. Given the reduced power budgets for SoCs, in this paper, we develop solutions for combined energy minimization and communication reliability control. Redundant bus coding is proved to be an effective technique for trading off energy against reliability, so that the most efficient scheme can be selected to meet predefined reliability requirements in a low signal-to-noise ratio regime. We model on-chip interconnects as noisy channels and evaluate the impact of two error recovery schemes on energy efficiency: correction at the receiver stage versus retransmission of corrupted data. The analysis is performed in a realistic SoC setting, and holds both for shared communication resources and for peer-to-peer links in a network of interconnects. We provide SoC designers with guidelines for the selection of energy efficient error-control schemes for communication architectures. Davide Bertozzi, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | NoC Synthesis Flow for Customized Domain Specific Multiprocessor Systems-on-ChipabstractThe growing complexity of customizable single-chip multiprocessors is requiring communication resources that can only be provided by a highly-scalable communication infrastructure. This trend is exemplified by the growing number of network-on-chip (NoC) architectures that have been proposed recently for system-on-chip (SoC) integration. Developing NoC-based systems tailored to a particular application domain is crucial for achieving high-performance, energy-efficient customized solutions. The effectiveness of this approach largely depends on the availability of an ad hoc design methodology that, starting from a high-level application specification, derives an optimized NoC configuration with respect to different design objectives and instantiates the selected application specific on-chip micronetwork. Automatic execution of these design steps is highly desirable to increase SoC design productivity. This work illustrates a complete synthesis flow, called Netchip, for customized NoC architectures, that partitions the development work into major steps (topology mapping, selection, and generation) and provides proper tools for their automatic execution (SUNMAP, xpipescompiler). The entire flow leverages the flexibility of a fully reusable and scalable network components library called xpipes, consisting of highly-parameterizable network building blocks (network interface, switches, switch-to-switch links) that are design-time tunable and composable to achieve arbitrary topologies and customized domain-specific NoC architectures. Several experimental case studies are presented In the work, showing the powerful design space exploration capabilities of the proposed methodology and tools. Davide Bertozzi, Antoine Jalabert, Srinivasan Murali, Rutuparna Tamhankar, Stergios Stergiou, Luca Benini, Giovanni De Micheli |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Analyzing On-Chip Communication in a MPSoC EnvironmentabstractThis work focuses on communication architecture analysis for multi-processor systems-on-chips (MPSoCs), and it leverages a SystemC-based platform to simulate a complete multi-processor system at the cycle-accurate and signal-accurate level. These features allow to stimulate the communication sub-system with functional traffic generated by real applications running on top of a configurable number of ARM processors. This opens up the possibility for communication infrastructure exploration and for the investigation of its impact on system performance at the highest level of accuracy. Our simulation environment proved capable of a detailed comparative analysis between two industry-standard communication architectures, under realistic workloads and different system configurations, pointing out the impact of fine grained architectural mismatches on macroscopic performance differences. Mirko Loghi, Federico Angiolini, Davide Bertozzi, Luca Benini, Roberto Zafalon |
DATE | 3 |
| 2003 | Transport Protocol Optimization for Energy Efficient Wireless Embedded Systems
Davide Bertozzi, Anand Raghunathan, Luca Benini, Srivaths Ravi 0001 |
DATE | 1 |
| 2003 | xpipes: a Latency Insensitive Parameterized Network-on-chip Architecture For Multi-Processor SoCsabstractThe growing complexity of customizable embedded multiprocessor architectures for digital media processing will soon require highly scalable network-on-chip based communication infrastructures. Here, we propose xpipes, a scalable and high-performance NoC architecture for multiprocessor SoCs, consisting of soft macros that can be turned into instance-specific network components at instantiation time. The flexibility of its components allows our NoC to support both homogeneous and heterogeneous architectures. The interface with IP cores at the periphery of the network is standardized (OCP-based). Links can be pipelined with a flexible number of stages to decouple data introduction speed from worst-case link delay. Switches are lightweight and support reliable communication for arbitrary link pipeline depths (latency insensitive operation). Xpipes has been described in synthesizable SystemC, at the cycle-accurate and signal-accurate level. Matteo Dall'Osso, Gianluca Biccari, Luca Giovannini, Davide Bertozzi, Luca Benini |
ICCD | 4 |
| 2002 | Low Power Error Resilient Encoding for On-Chip Data BusesabstractAs technology scales toward deep submicron, on-chip interconnects are becoming more and more sensitive to noise sources such as power supply noise, crosstalk, radiation induced effects, etc. Transient delay and logic faults are likely to reduce the reliability of data transfers across data-path bus lines. This paper investigates how to deal with these errors in an energy efficient way. We could opt for error correction, which exhibits larger decoding overhead, or for the retransmission of the incorrectly received data word. Provided the timing penalty associated with this latter technique can be tolerated, we show that retransmission strategies are more effective than correction ones from an energy viewpoint, both for the larger detection capability and for the minor decoding complexity. The analysis wits performed by implementing several variants of a Hamming code in the VHDL model of a processor based on the Sparc V8 architecture, and exploiting the characteristics of AMBA bus slave response cycles to carry out retransmissions in a way fully compliant with this standard on-chip bus specification. Davide Bertozzi, Luca Benini, Giovanni De Micheli |
DATE | 1 |
| 2002 | Legacy SystemC Co-Simulation of Multi-Processor Systems-on-ChipabstractWe present a co-simulation environment for multiprocessor architectures, that is based on SystemC and allows a transparent integration of instruction set simulators (ISSs) within the SystemC simulation framework. The integration is based on the well-known concept of bus wrapper, that realizes the interface between the ISS and the simulator. The proposed solution uses an ISS-wrapper interface based on the standard gdb remote debugging interface, and implements two alternative schemes that differ in the amount of communication they require. The two approaches provide different degrees of tradeoff between simulation granularity and speed, and show significant speedup with respect to a micro-architectural, full SystemC simulation of the system description. Luca Benini, Davide Bertozzi, Davide Bruni, Nicola Drago, Franco Fummi, Massimo Poncino |
ICCD | 2 |
| 2002 | Parametric timing and power macromodels for high level simulation of low-swing interconnectsabstractThe impact of global on-chip interconnections on power consumption and speed of integrated circuits is becoming a serious concern. Designers need therefore to quickly estimate how performance and power are affected by a given choice of the interconnection parameters (length, voltage swing, driver and receiver schematics and sizing). This work focuses on the entire communication channel (driver, interconnect, receiver), and provides high level parametric VHDL simulation models for low-swing signaling schemes. These SPICE-derived power and timing macromodels transfer electrical-level information to the RTL simulation in an event-driven fashion, as transitions occur at the input of the interconnect driver. The accuracy reached by this back-annotation technique is within 5% with respect to SPICE results, with only 4% simulation speed penalty in the worst case. Davide Bertozzi, Luca Benini, Bruno Riccò |
ISLPED | 1 |
| 2002 | Power aware network interface management for streaming multimediaabstractA significant fraction of the power in portable devices (handheld terminals, PDAs, laptops, etc.) is drawn by the wireless network interface card (NIC). We test the viability of a novel power management technique, based on the exploitation of the NIC off-mode while a client-controlled streaming multimedia application is in progress. The basic idea is to switch off the card while frames are being played back, until a low-threshold level is reached in the client buffer. In this paper, pessimistic assumptions are made on the timing and power overheads associated with recovering the card from off-mode, but nevertheless a power saving of about 25% is achieved over the average power consumption incurred by the standard IEEE 802.11 mechanism. We also provide design curves showing the minimum buffer size that makes our technique effective, as a function of the network bandwidth and of the card characteristics. Davide Bertozzi, Luca Benini, Bruno Riccò |
WCNC | 1 |