EDBT 2026 Demo / reviewers in the wild / expert
Luigi Raffo
dblp:90/6155
· DBLP profile ↗
48ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0001-9683-009XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SYNtzulA: Open Hardware for Near-Sensor SNN InferenceabstractSpiking Neural Networks (SNNs) exploit event-driven processing to offer high energy efficiency when deploying Artificial Intelligence (AI) on wearable edge devices. However, specialized hardware is needed to fully take advantage of this potential, which, despite recent advances, remains expensive and not widely accessible. To address this, open-source Electronic Design Automation (EDA) tools and Process Design Kits (PDKs) offer a path to democratize the development of neuromorphic hardware. In this work, we present SYNtzulA, a system-on-chip designed for SNN acceleration, developed using the open-source IHP-SG13G2 130 nm PDK and the OpenROAD toolchain. The chip integrates a RISC-V softcore and a dedicated SNN accelerator, occupying approximately 6.8mm2including I/O pads. It operates at up to 125 MHz, reaching a throughput of 2 Giga Synaptic Operations per second (GSOP/s) with an energy consumption of 36.5 pJ per synaptic operation. The accelerator can exploit the sparsity of spike-based computation by skipping unnecessary operations, resulting in total energy consumption in the order of a few hundred nanojoules per inference in different use cases involving biosignal analysis. Luca Martis, Gianluca Leone, Luigi Raffo, Paolo Meloni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Enabling SNN-Based Near-MEA Neural Decoding with Channel Selection: An Open-HW ApproachabstractAdvancements in CMOS microelectrode array sensors have significantly improved sensing area and resolution, paving the way to accurate Brain-Machine Interfaces (BMIs). However, near-sensor neural decoding on implantable computing devices is still an open problem. A promising solution is provided by Spiking Neural Networks (SNNs), which leverage event sparsity to improve energy consumption. However, given the typical data rates involved, the workload related to I/O acquisition and spike encoding is dominant and limits the benefits achievable with event-based processing. In this work, we present two power-efficient implementations, on FPGA and ASIC, of a dedicated processor for the decoding of intracortical action potentials from primary motor cortex. The processor leverages lightweight sparse SNNs to achieve state-of-the-art accuracy. To limit the impact of I/O transfers on energy efficiency, we introduced a channel selection scheme that reduced bandwidth requirements by 3x and power consumption by 2.3x and 1.6x on the FPGA and ASIC, respectively, enabling inference at 0.446 μJ and 1.04 μJ, with no significant loss in accuracy. To promote broad adoption in a specialized, research-intensive domain, we have based our implementations on open-source EDA tools, low-cost hardware, and an open PDK. Gianluca Leone, Luca Martis, Luigi Raffo, Paolo Meloni |
DATE | 3 |
| 2025 | SYNtzulu: A Tiny RISC-V-Controlled SNN Processor for Real-Time Sensor Data Analysis on Low-Power FPGAsabstractSpiking Neural Networks (SNNs) are energy- and performance-efficient tools that have been found to be very useful in AI applications at the edge. This paper introducesSYNtzulu, an SNN processing element designed to be used in low-cost and low-power FPGA devices for near-sensor data analysis. The system is equipped with a RISC-V subsystem responsible for controlling the input/output and setting runtime parameters, thus increasing its flexibility. We evaluated the system, which was implemented on a Lattice iCE40UP5K FPGA, in various use cases employing SNNs with accuracy comparable to the state-of-the-art.SYNtzuludissipates a maximum power of 12.05 mW when performing SNN inference, which can be reduced to an average of just 1.45 mW through the use of dynamic power management. Gianluca Leone, Matteo Antonio Scrugli, Lorenzo Badas, Luca Martis, Luigi Raffo, Paolo Meloni |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | A multithread AES accelerator for Cyber-Physical SystemsabstractComputing elements of CPSs must be flexible to ensure interoperability; and adaptive to cope with the evolving internal and external state, such as battery level and critical tasks. Cryptography is a common task needed in CPSs to guarantee private communication among different devices. In this work, we propose a reconfigurable FPGA accelerator for AES workloads with different key lengths. The accelerator architecture exploits tagged-dataflow models to support the concurrent execution of multiple threads on the same accelerator. This solution demonstrates to be more resource- and energy-efficient than a set of non-reconfigurable accelerators while keeping high performance and flexibility of execution. Francesco Ratto, Luigi Raffo, Francesca Palumbo |
CF | 2 |
| 2020 | Design and Usability Assessment of a Multi-Device SOA-Based Telecare Framework for the ElderlyabstractTelemonitoring is a branch of telehealth that aims at remotely monitoring vital signs, which is important for chronically ill patients and the elderly living alone. The available standalone devices and applications for the self-monitoring of health parameters largely suffer from interoperability problems; meanwhile, telemonitoring medical devices are expensive, self-contained, and are not integrated into user-friendly technological platforms for the end user. This paper presents the technical aspects and usability assessment of the telemonitoring features of the HEREiAM platform, which supports heterogeneous information technology systems. By exploiting a service-oriented architecture, the measured parameters collected by off-the-shelf Bluetooth medical devices are sent as XML documents to a private cloud that implements an interoperable health service infrastructure, which is compliant with the most recent healthcare standards and security protocols. This Android-based system is designed to be accessible both via TV and portable devices, and includes other utilities designed to support the elderly living alone. Four usability assessment sessions with quality-validated questionnaires were performed to accurately understand the ease of use, usefulness, acceptance, and quality of the proposed system. The results reveal that our system achieved very high usability scores even at its first use, and the scores did not significantly change over time during a field trial that lasted for four months, reinforcing the idea of an intuitive design. At the end of such a trial, the user-experience questionnaire achieved excellent scores in all aspects with respect to the benchmark. Good results were also reported by general practitioners who assessed the quality of their remote interfaces for telemonitoring. Silvia Macis, Daniela Loi, Andrea Ulgheri, Danilo Pani, Giuliana Solinas, Serena La Manna, Vincenzo Cestone, Davide Guerri, Luigi Raffo |
IEEE J. Biomed. Health Informatics | 9 |
| 2020 | NeuPow: A CAD Methodology for High-level Power Estimation Based on Machine LearningabstractIn this article, we present a new, simple, accurate, and fast power estimation technique that can be used to explore the power consumption of digital system designs at an early design stage. We exploit the machine learning techniques to aid the designers in exploring the design space of possible architectural solutions, and more specifically, their dynamic power consumption, which is application-, technology-, frequency-, and data-stimuli dependent. To model the power and the behavior of digital components, we adopt the Artificial Neural Networks (ANNs), while the final target technology is Application Specific Integrated Circuit (ASIC). The main characteristic of the proposed method, called NeuPow, is that it relies on propagating the signals throughout connected ANN models to predict the power consumption of a composite system. Besides a baseline version of the NeuPow methodology that works for a given predefined operating frequency, we also derive an upgraded version that is frequency-aware, where the same operating frequency is taken as additional input by the ANN models. To prove the effectiveness of the proposed methodology, we perform different assessments at different levels. Moreover, technology and scalability studies have been conducted, proving the NeuPow robustness in terms of these design parameters. Results show a very good estimation accuracy with less than 9% of relative error independently from the technology and the size/layers of the design. NeuPow is also delivering a speed-up factor of about 84× with respect to the classical power estimation flow. Yehya Nasser, Carlo Sau, Jean-Christophe Prévotet, Tiziana Fanni, Francesca Palumbo, Maryline Hélard, Luigi Raffo |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2019 | NeuPow: artificial neural networks for power and behavioral modeling of arithmetic components in 45nm ASICs technologyabstractIn this paper, we present a flexible, simple and accurate power modeling technique that can be used to estimate the power consumption of modern technology devices. We exploit Artificial Neural Networks for power and behavioral estimation in Application Specific Integrated Circuits. Our method, called NeuPow, relies on propagating the predictors between the connected neural models to estimate the dynamic power consumption of the individual components. As a first proof of concept, to study the effectiveness of NeuPow, we run both component level and system level tests on the Open GPDK 45 nm technology from Cadence, achieving errors below 1.5% and 9% respectively for component and system level. In addition, NeuPow demonstrated a speed up factor of 2490X. Yehya Nasser, Carlo Sau, Jean-Christophe Prévotet, Tiziana Fanni, Francesca Palumbo, Maryline Hélard, Luigi Raffo |
CF | 7 |
| 2019 | CERBERO: Cross-layer modEl-based fRamework for multi-oBjective dEsign of reconfigurable systems in unceRtain hybRid envirOnments: Invited paper: CERBERO teams from UniSS, UniCA, IBM Research, TASE, INSA-Rennes, UPM, USI, Abinsula, AmbieSense, TNO, S&T, CRFabstractCyber-Physical Systems (CPS) are embedded computational collaborating devices, capable of sensing and controlling physical elements and, often, responding to humans. Designing and managing systems able to respond to different, concurrent requirements during operation is not straightforward, and introduce the need of proper support at design-time and run-time. The Cross-layer modEl-based fRamework for multi-oBjective dEsign of Reconfigurable systems in unceRtain hybRid envirOnments (CERBERO) EU project has developed a design environment for adaptive CPS. CERBERO approach leverages on model-based methodologies including different technologies and tools developed to cover design and operation from user interactions down to low level computing layer implementation. Francesca Palumbo, Tiziana Fanni, Carlo Sau, Luca Pulina, Luigi Raffo, Michael Masin, Evgeny Shindin, Pablo Sanchez de Rojas, Karol Desnos, Maxime Pelcat, Alfonso Rodríguez 0002, Eduardo Juárez Martínez, Francesco Regazzoni 0001, Giuseppe Meloni, Maria Katiuscia Zedda, Hans I. Myrhaug, Leszek Kaliciak, Joost Adriaanse, Julio de Oliveira Filho, Antonella Toffetti |
CF | 5 |
| 2019 | A runtime-adaptive cognitive IoT node for healthcare monitoringabstractWearable and energy efficient processing nodes, allowing for continuous remote monitoring of patient vital parameters, are mainstream in modern health-care practice. Most recent approaches to the development of such systems combine near-sensor data processing with cognitive computing, to improve at the same time communication efficiency, responsiveness and accuracy of the analysis of the sensed data. In this paper, we present a hardware-software architecture for a connected sensor-processing node that allows the set of in-place processing tasks to be executed to be remotely controllable by an external user. The designed system is capable of dynamically adapting its operating point to the selected computational load, to minimize power consumption. The benefits of the proposed approach are tested on a use-case involving ECG monitoring, that, when selected, performs ECG classification using a lightweigth convolutional neural network. Experimental results show how the proposed approach can provide more than 50% power consumption reduction for common ECG activity, with less than 2% memory footprint overhead and reconfiguring the system in less than 1 ms. Matteo Antonio Scrugli, Daniela Loi, Luigi Raffo, Paolo Meloni |
CF | 3 |
| 2019 | An integrated hardware/software design methodology for signal processing systemsabstractThis paper presents a new methodology for design and implementation of signal processing systems on system-on-chip (SoC) platforms. The methodology is centered on the use of lightweight application programming interfaces for applying principles of dataflow design at different layers of abstraction. The development processes integrated in our approach are software implementation, hardware implementation, hardware-software co-design, and optimized application mapping. The proposed methodology facilitates development and integration of signal processing hardware and software modules that involve heterogeneous programming languages and platforms. As a demonstration of the proposed design framework, we present a dataflow-based deep neural network (DNN) implementation for vehicle classification that is streamlined for real-time operation on embedded SoC devices. Using the proposed methodology, we apply and integrate a variety of dataflow graph optimizations that are important for efficient mapping of the DNN system into a resource constrained implementation that involves cooperating multicore CPUs and field-programmable gate array subsystems. Through experiments, we demonstrate the flexibility and effectiveness with which different design transformations can be applied and integrated across multiple scales of the targeted computing system. Lin Li 0029, Carlo Sau, Tiziana Fanni, Jingui Li, Timo Viitanen, François Christophe, Francesca Palumbo, Luigi Raffo, Heikki Huttunen, Jarmo Takala, Shuvra S. Bhattacharyya |
J. Syst. Archit. | 8 |
| 2018 | NEURAghe: Exploiting CPU-FPGA Synergies for Efficient and Flexible CNN Inference Acceleration on Zynq SoCsabstractDeep convolutional neural networks (CNNs) obtain outstanding results in tasks that require human-level understanding of data, like image or speech recognition. However, their computational load is significant, motivating the development of CNN-specialized accelerators. This work presents NEURA ghe , a flexible and efficient hardware/software solution for the acceleration of CNNs on Zynq SoCs. NEURA ghe leverages the synergistic usage of Zynq ARM cores and of a powerful and flexible Convolution-Specific Processor deployed on the reconfigurable logic. The Convolution-Specific Processor embeds both a convolution engine and a programmable soft core, releasing the ARM processors from most of the supervision duties and allowing the accelerator to be controlled by software at an ultra-fine granularity. This methodology opens the way for cooperative heterogeneous computing: While the accelerator takes care of the bulk of the CNN workload, the ARM cores can seamlessly execute hard-to-accelerate parts of the computational graph, taking advantage of the NEON vector engines to further speed up computation. Through the companion NeuDNN SW stack, NEURA ghe supports end-to-end CNN-based classification with a peak performance of 169GOps/s, and an energy efficiency of 17GOps/W. Thanks to our heterogeneous computing model, our platform improves upon the state-of-the-art, achieving a frame rate of 5.5 frames per second (fps) on the end-to-end execution of VGG-16 and 6.6fps on ResNet-18. Paolo Meloni, Alessandro Capotondi, Gianfranco Deriu, Michele Brian, Francesco Conti 0001, Davide Rossi 0001, Luigi Raffo, Luca Benini |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2017 | Cross-layer design of reconfigurable cyber-physical systemsabstractIn the last few years, besides the concepts of embedded and interconnected systems, also the notion of Cyber-Physical Systems (CPS) has emerged: embedded computational collaborating devices, capable of sensing and controlling physical elements and, often, responding to humans. The continuous interaction between physical and computing layers makes their design and maintenance extremely complex. Uncertainty management and runtime reconfigurability, to mention the most relevant ones, are rarely tackled by available toolchains. In this context, the Cross-layer modEl-based fRamework for multi-oBjective dEsign of Reconfigurable systems in unceRtain hybRid envirOnments (CERBERO) EU project aims at developing a design environment for CPS based of two pillars: 1) a cross-layer model-based approach to describe, optimize, and analyze the system and all its different views concurrently and 2) an advanced adaptivity support based on a multi-layer autonomous engine. In this work, we describe the components and the required developments for seamless design of reusable and reconfigurable CPS and System of Systems in uncertain hybrid environments. Michael Masin, Francesca Palumbo, Hans I. Myrhaug, J. A. de Oliveira Filho, M. Pastena, Maxime Pelcat, Luigi Raffo, Francesco Regazzoni 0001, A. A. Sanchez, Antonella Toffetti, Eduardo de la Torre, Maria Katiuscia Zedda |
DATE | 7 |
| 2017 | Hardware design methodology using lightweight dataflow and its integration with low power techniques
Tiziana Fanni, Lin Li 0029, Timo Viitanen, Carlo Sau, Renjie Xie, Francesca Palumbo, Luigi Raffo, Heikki Huttunen, Jarmo Takala, Shuvra S. Bhattacharyya |
J. Syst. Archit. | 7 |
| 2017 | Real-Time neural signal decoding on heterogeneous MPSocs based on VLIW ASIPs
Paolo Meloni, Claudio Rubattu, Giuseppe Tuveri, Danilo Pani, Luigi Raffo, Francesca Palumbo |
J. Syst. Archit. | 5 |
| 2015 | The challenge of collaborative telerehabilitation: conception and evaluation of a telehealth system enhancement for home-therapy follow-upabstractSummary Telerehabilitation aims to solve problems like equitable access to the rehabilitation and cost reduction by providing rehabilitation services at a distance. The largest part of telerehabilitation systems implement a real‐time one‐to‐one process involving patient and therapist. Even though they can be successfully exploited in conditions such as post‐traumatic recovery, in complex scenarios, this simple model should be replaced by a more structured collaborative one envisioning a multidisciplinary team. This paper presents the design and evaluation of a patient‐centric collaborative telerehabilitation framework aimed at supporting a multidisciplinary team in the follow‐up of domiciliary patients. The proposed framework follows the experience of a clinical trial that exploited a novel telerehabilitation device not conceived to support collaborative scenarios. Compared with the original system, the proposed extension allows the hierarchical division of the responsibility within the medical team, promoting a collaborative management of the rehabilitation. Proactive and decisional behaviors, as well as consulting practices on shared data within the medical team, are fostered by the system. Semi‐structured interviews have been administered to a panel of experts to evaluate the proposed approach. The collected feedback can be exploited to finely tune the system in view of a new clinical trial including new functionalities. Copyright © 2014 John Wiley & Sons, Ltd. Danilo Pani, Gianluca Barabino, Alessia Dessì, Selene Uras, Luigi Raffo |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Computing Swarms for Self-Adaptiveness and Self-Organization in Floating-Point Array ProcessingabstractAdvancements in CMOS technology enable the integration of a huge number of resources on the same system-on-chip. Managing the consequent growing complexity, including fault tolerance issues in deep submicron technologies, is a hard challenge for hardware designers. Self-organization may represent a viable path toward the development of massively parallel architectures in current and future technologies. This approach is progressively more studied in multiprocessor architectures where, however, a further mind-set shift in terms of programming paradigm is required. In this article, self-organization and self-adaptiveness are exploited for the design of a coprocessing unit for array computations, supporting floating-point arithmetic. From the experience of previous explorations, an architecture embodying some principle of swarm intelligence to pursue adaptability, scalability, and fault tolerance is proposed. The architecture realizes a loosely structured collection of hardware agents implementing fixed behavioral rules aimed at the best exploitation of the available resources in whatever kind of context without any hardware reconfiguration. Comparisons with off-the-shelf very long instruction word (VLIW) digital signal processors (DSPs) on specific tasks reveal similar performance thus not paying the improved robustness with performance. The multitasking capabilities, together with the intrinsic scalability, make this approach valuable for future extensions as well, especially in the field of neuronal networks simulators. Danilo Pani, Carlo Sau, Francesca Palumbo, Luigi Raffo |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2014 | Real-time blind audio source separation: performance assessment on an advanced digital signal processor
Danilo Pani, Alessandro Pani, Luigi Raffo |
J. Supercomput. | 3 |
| 2013 | Exploring hardware support for scaling irregular applications on multi-node multi-core architecturesabstractThe recent emergence of large-scale knowledge discovery, data mining and social network analysis, irregular applications have gained renewed interest. Cache-based architectures do not provide optimal performances with such workloads, mainly due to the low spatial and temporal locality of their control and memory access patterns. This paper presents a multi-node, multi-core, multi-threaded shared-memory system architecture designed for the execution of large-scale irregular applications, and built on top of three pillars that support these workloads. First, transparent hardware support for Partitioned Global Address Space (PGAS) provides a large globally-shared address space with no software library overhead. Second, multithreaded multi-core processing nodes achieve the necessary latency tolerance required when accessing physically distributed global memory. Third, hardware support is provided for inter-thread synchronization on the global address space. An analytical performance model that accounts for the main architecture and application characteristics is presented. The hardware design of the proposed custom architectural building blocks is then described. Finally, a multi-board FPGA prototype of the proposed system with typical irregular kernels and benchmarks is presented. The experimental evaluation demonstrates the architecture performance scalability for different configurations of the whole system. Simone Secchi, Marco Ceriani, Antonino Tumeo, Oreste Villa, Gianluca Palermo, Luigi Raffo |
ASAP | 6 |
| 2012 | Exploiting binary translation for fast ASIP design space exploration on FPGAsabstractComplex Application Specific Instruction-set Processors (ASIPs) expose to the designer a large number of degrees of freedom, posing the need for highly accurate and rapid simulation environments. FPGA-based emulators represent an alternative to software cycle-accurate simulators, preserving maximum accuracy and reasonable simulation times. The work presented in this paper aims at exploiting FPGA emulation within technology aware design space exploration of ASIPs. The potential speedup provided by reconfigurable logic is reduced by the overhead of RTL synthesis/implementation. This overhead can be mitigated by reducing the number of FPGA implementation processes, through the adoption of binary-level translation. Hereby we present a prototyping method that, given a set of candidate ASIP configurations, defines an overdimensioned ASIP architecture, capable of emulating all the design space points under evaluation. This approach is then evaluated with a design space exploration case study. Along with execution time, by coupling FPGA emulation with activity-based physical modeling, we can extract area/power/energy figures. Sebastiano Pomata, Paolo Meloni, Giuseppe Tuveri, Luigi Raffo, Menno Lindwer |
DATE | 4 |
| 2012 | ASAM: Automatic Architecture Synthesis and Application MappingabstractThis paper focuses on mastering the automatic architecture synthesis and application mapping for heterogeneous massively-parallel MPSoCs based on customizable application-specific instruction-set processors (ASIPs). It presents an over-view of the research being currently performed in the scope of the European project ASAM of the ARTEMIS program. The paper briefly presents the results of our analysis of the main problems to be solved and challenges to be faced in the design of such heterogeneous MPSoCs. It explains which system, design, and electronic design automation (EDA) concepts seem to be adequate to resolve the problems and address the challenges. Finally, it introduces and briefly discusses the ASAM design-flow and its main stages. Lech Józwiak, Menno Lindwer, Rosilde Corvino, Paolo Meloni, Laura Micconi, Jan Madsen, Erkan Diken, Deepak Gangadharan, Roel Jordans, Sebastiano Pomata, Paul Pop, Giuseppe Tuveri, Luigi Raffo |
DSD | 13 |
| 2012 | System Adaptivity and Fault-Tolerance in NoC-based MPSoCs: The MADNESS Project ApproachabstractModern embedded systems increasingly require adaptive run-time management. The system may adapt the mapping of the applications in order to accommodate the current workload conditions, to balance load for efficient resource utilization, to meet quality of service agreements, to avoid thermal hot-spots and to reduce power consumption. As the possibility of experiencing run-time faults becomes increasingly relevant with deep-sub-micron technology nodes, in the scope of the MADNESS project, we focus particularly on the problem of graceful degradation by dynamic remapping in presence of run-time faults. In this paper, we summarize the major results achieved in the MADNESS project until now regarding the system adaptivity and fault tolerant processing. We report the first results of the integration between platform level and middleware level support for adaptivity and fault tolerance. A case study demonstrates the survival ability of the system via a low-overhead process migration mechanism and a near-optimal online remapping heuristic. Paolo Meloni, Giuseppe Tuveri, Luigi Raffo, Emanuele Cannella, Todor P. Stefanov, Onur Derin, Leandro Fiorin, Mariagiovanna Sami |
DSD | 3 |
| 2012 | Multi-purpose systems: A novel dataflow-based generation and mapping strategyabstractThe manual creation of specialized hard-ware infrastructures for complex multi-purpose systems is error-prone and time-consuming. Moreover, lots of effort is required to define an optimized and heterogeneous components library. To tackle these issues, we propose a novel design flow based on the Dataflow Process Networks Model of Computation. In particular, we have combined the operation of two state of the art tools, the Multi-Dataflow Composer and the Open RVC-CAL Compiler, handling respectively the automatic mapping of a reconfigurable multi-purpose substrate and the high level synthesis of hardware components. Our approach guarantees runtime efficiency and on-chip area saving both on FPGAs and ASICs. Jean-François Nezan, Nicolas Siret, Matthieu Wipliez, Francesca Palumbo, Luigi Raffo |
ISCAS | 5 |
| 2010 | Exploiting FPGAs for technology-aware system-level evaluation of multi-core architecturesabstractThe hardware-software co-development of modern complex MPSoC computing platforms exposes to the designer a huge complexity, resulting from the combination of vastly different architectural possibilities with strict demands posed by the target applications. To handle this complexity, highly accurate but rapid prototyping/evaluation environments need to be developed, that would possibly be able to provide an effective measurement of the system under design as soon as possible, allowing to comply with current time-to-market. While software-based fully cycle-accurate simulators do not seem to represent anymore an adequate solution to solve this issue, the attention has been recently shifted to the adoption of hardware emulators in the early stages of the design flow. In this work, we present an emulation framework for library-based semiautomatic instantiation of complex multi-core platforms that exploits FPGA devices to provide detailed functional information on the platform under development, and at the same time using hardware execution traces with technology-related analytical models to extract, already at system-level, physical metrics on power consumption, maximum operating frequency and area occupation of a prospective ASIC implementation of the system. Two prospective use case scenarios are presented to validate the usefulness of the presented framework: the first one analyzes the mapping and the scalability of a highly parallel application over a 2D homogeneous mesh architecture for increasing number of processors, while the second one employs the emulation infrastructure inside a design space exploration flow for the configuration of some interconnection network parameters. Simone Secchi, Paolo Meloni, Luigi Raffo |
ISPASS | 3 |
| 2010 | Impact of Half-Duplex and Full-Duplex DMA Implementations on NoC PerformanceabstractNoCs performance are usually explored stand-alone, overlooking the impact of the higher communication levels in the ISO OSI micro network stack. Nevertheless, since CPUs have to be relieved of communication management, higher communication levels such as DMA engines necessarily influence the communication performance. In this paper, we investigate how two different DMA implementations, full-duplex and half-duplex, can bias the behavior of a NoC designed for MPP architectures. From our studies, it turned out that a full-duplex DMA is more effective in preventing possible deadlock situations. Moreover, a deep performance analysis of a state-of-the-art NoC, in terms of transactions completion time, queuing time and injection delay, confirms the impact of the DMA in NoC-based MPP platforms, showing the advantages of a full-duplex approach. Francesca Palumbo, Danilo Pani, Alessandro Pilia, Luigi Raffo |
NOCS | 4 |
| 2010 | Enabling fast Network-on-Chip topology selection: an FPGA-based runtime reconfigurable prototyperabstractThe complexity of modern interconnect architecture design requires highly accurate and rapid simulation environments. FPGA-based emulators have been proposed as an alternative to software cycle-accurate simulators, preserving maximum accuracy and reasonable simulation times. However, the potential speedup is reduced by the time overhead needed for RTL synthesis/implementation. This paper proposes runtime reconfiguration of the architecture to push the hardware emulation one step further, by reducing the number of FPGA implementation processes to be run. To this aim, this work presents an algorithm that synthesizes, for a set of candidate architectural configurations, a connection topology capable of reconfiguring itself via software to emulate all the design space points under evaluation. We present the actual reconfiguration algorithm, the CAD tools and the hardware mechanisms that implement it. The design capabilities provided by this approach are evaluated with a design space exploration case study. Paolo Meloni, Simone Secchi, Luigi Raffo |
VLSI-SoC | 3 |
| 2008 | A Network on Chip Architecture for Heterogeneous Traffic Support with Non-Exclusive Dual-Mode SwitchingabstractAs the multi-core processors era took place, several design concerns have risen. Interconnection layer efficiency has gained particular relevance as a crucial issue to be addressed in order to leverage the large amount of on-chip resources that today's VLSI technologies are able to provide. At the same time, as the architectural parallelism will continue to grow and become more fine-grained, the kind of traffic generated by the different multithreaded applications is turning out to be very wide-ranging in terms of size and burstiness. In order to adapt to this large variety of traffic to be supported, several models of dual-mode routers have been developed, implementing both packet switching and circuit switching techniques, thus supporting both best effort and guaranteed throughput services. This paper introduces an innovative model of non-exclusive dual-mode router, able to combine the aforementioned features in a non exclusive way (i.e.: in parallel inside the network on the same link). This feature makes this NoC architecture well-suited for multi-processor system on-chip (MPSoC) architectures with a high level of parallelism which have to deal with heterogeneous traffic conditions, such as massively parallel processors (MPPs) and processor arrays (PAs). Simone Secchi, Francesca Palumbo, Danilo Pani, Luigi Raffo |
DSD | 4 |
| 2007 | On the impact of serialization on the cache performances in Network-on-Chip based MPSoCsabstractNetwork on Chip architectures are proposed as a solution to overcome functional and physical scalability shown by shared bus based MPSoC architecture. Unfortunately to implement and efficient communication infrastructure, the designer has to set a lot of parameters. An exhaustive knowledge of how the chosen settings influence the overall behaviour of the designed system is then mandatory. Aim of this paper is to discuss the relationship between the performances of a NoC and its configuration parameters in the case of traffic generated by cache operations (block replacements). We paid special attention to investigate the impact of the serialization factor, that was already not clearly assessed in literature for this important case study. A numerical analysis, referring to an actual implementation of the NoC on a state-of-the-art 65 nm technological process was performed. The obtained results were used to report an energy and execution time exploration over the complete design space of interest. Paolo Meloni, Giovanni Busonera, Salvatore Carta, Luigi Raffo |
DSD | 4 |
| 2007 | NoC Design and Implementation in 65nm TechnologyabstractAs embedded computing evolves towards ever more powerful architectures, the challenge of properly interconnecting large numbers of on-chip computation blocks is becoming prominent. Networks-on-chip (NoCs) have been proposed as a scalable solution to both physical design issues and increasing bandwidth demands. However, this claim has not been fully validated yet, since the design properties and tradeoffs of NoCs have not been studied in detail below the 100 nm threshold. This work is aimed at shedding light on the opportunities and challenges, both expected and unexpected, of NoC design in nanometer CMOS. We present fully working 65 nm NoC designs, a complete NoC synthesis flow and detailed scalability analysis Antonio Pullini, Federico Angiolini, Paolo Meloni, David Atienza 0001, Srinivasan Murali, Luigi Raffo, Giovanni De Micheli, Luca Benini |
NOCS | 6 |
| 2007 | A Layout-Aware Analysis of Networks-on-Chip and Traditional Interconnects for MPSoCsabstractThe ever-shrinking lithographic technologies available to chip designers enable performance and functionality breakthroughs; yet, they bring new hard problems. For example, multiprocessor systems-on-chip featuring several processing elements can be conceived, but efficiently interconnecting them while keeping the design complexity manageable is a challenge. Traditional buses are easy to deploy, but cannot provide enough bandwidth for such complex systems. A departure from legacy architectures is therefore called for. One radical path is represented by packet-switching networks-on-chip, whereas a more conservative approach interleaves bandwidth-rich components (e.g., crossbars) within the preexisting fabrics. This paper is aimed at analyzing the strengths and weaknesses of these alternative approaches by performing a thorough analysis based on actual chip floorplans after the interconnection place&route stages and after a clock tree has been distributed across the layout. Performance, area, and power results will be discussed while keeping an eye on the scalability prospects in future technology nodes Federico Angiolini, Paolo Meloni, Salvatore Carta, Luigi Raffo, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Synthesis of Predictable Networks-on-Chip-Based Interconnect Architectures for Chip MultiprocessorsabstractToday, chip multiprocessors (CMPs) that accommodate multiple processor cores on the same chip have become a reality. As the communication complexity of such multicore systems is rapidly increasing, designing an interconnect architecture with predictable behavior is essential for proper system operation. In CMPs, general-purpose processor cores are used to run software tasks of different applications and the communication between the cores cannot be precharacterized. Designing an efficient network-on-chip (NoC)-based interconnect with predictable performance is thus a challenging task. In this paper, we address the important design issue of synthesizing the most power efficient NoC interconnect for CMPs, providing guaranteed optimum throughput and predictable performance for any application to be executed on the CMP. In our synthesis approach, we use accurate delay and power models for the network components (switches and links) that are obtained from layouts of the components using industry standard tools. The synthesis approach utilizes the floorplan knowledge of the NoC to detect timing violations on the NoC links early in the design cycle. This leads to a faster design cycle and quicker design convergence across the high-level synthesis approach and the physical implementation of the design. We validate the design flow predictability of our proposed approach by performing a layout of the NoC synthesized for a 25-core CMP. Our approach maintains the regular and predictable structure of the NoC and is applicable in practice to existing NoC architectures. Srinivasan Murali, David Atienza 0001, Paolo Meloni, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2006 | Contrasting a NoC and a traditional interconnect fabric with layout awarenessabstractIncreasing miniaturization is posing multiple challenges to electronic designers. In the context of multi-processor system-on-chips (MPSoCs), we focus on the problem of implementing efficient interconnect systems for devices which are ever more densely packed with parallel computing cores. Easily seen that traditional buses can not provide enough bandwidth, a revolutionary path to scalability is provided by packet-switched network-on-chips (NoCs), while a more conservative approach dictates the addition of bandwidth-rich components (e.g. crossbars) within the preexisting fabrics. While both alternatives have already been explored, a thorough contrastive analysis is still missing. In this paper, we bring crossbar and NoC designs to the chip layout level in order to highlight the respective strengths and weaknesses in terms of performance, area and power, keeping an eye on future scalability Federico Angiolini, Paolo Meloni, Salvatore Carta, Luca Benini, Luigi Raffo |
DATE | 5 |
| 2006 | Automatic Application Partitioning on FPGA/CPU Systems Based on Detailed Low-Level InformationabstractReconfigurable FPGA/CPU systems are widely described in literature as a viable processing solution for embedded and high end processing. One of the key issues of this kind of approach is the code partitioning between CPU and FPGA. The development of automatic partitioning tools allows to obtain optimized architecture without a specific knowledge of digital design. In this paper we present a framework which, starting from an ANSI C application code: (i) automatically identifies code fragments suitable for hardware implementation as specialized functional units (ii) for all these segments a synthesizable code is generated and sent to a synthesis tool, (iii) from the synthesis results, the segments to be implemented on FPGA are selected (iv) bit stream to configure the FPGA and modified C code to be executed on the CPU are generated. We applied this tool to standard benchmarks obtaining, with respect to state of the art, an improvement of up to 250% in the accuracy of performances estimation related to the selected segments of code. This leads to a more optimized code partitioning Giovanni Busonera, Salvatore Carta, Andrea Marongiu, Luigi Raffo |
DSD | 4 |
| 2006 | Designing application-specific networks on chips with floorplan informationabstractWith increasing communication demands of processor and memory cores in Systems on Chips (SoCs), scalable Networks on Chips (NoCs) are needed to interconnect the cores. For the use of NoCs to be feasible in today's industrial designs, a custom-tailored, application-specific NoC that satisfies the design objectives and constraints of the targeted application domain is required. In this work, we present a design methodology that automates the synthesis of such application-specific NoC architectures. We present a floorplan aware design method that considers the wiring complexity of the NoC during the topology synthesis process. This leads to detecting timing violations on the NoC links early in the design cycle and to have accurate power estimations of the interconnect. We incorporate mechanisms to prevent deadlocks during routing, which is critical for proper operation of NoCs. We integrate the NoC synthesis method with an existing design flow, automating NoC synthesis, generation, simulation and physical design processes. We also present ways to ensure design convergence across the levels. Experiments on several SoC benchmarks are presented, which show that the synthesized topologies provide a large reduction in network power consumption (2.78x on average) and improvement in performance (1.59x on average) over the best mesh and mesh-based custom topologies. An actual layout of a multimedia SoC with the NoC designed using our methodology is presented, which shows that the designed NoC supports the required frequency of operation (close to 900 MHz) without any timing violations. We could design the NoC from input specifications to layout in 4 hours, a process that usually takes several weeks. Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
ICCAD | 8 |
| 2006 | Designing Message-Dependent Deadlock Free Networks on Chips for Application-Specific Systems on ChipsabstractNetworks on chip (NoC) has emerged as the paradigm for designing scalable communication architecture for systems on chips (SoCs). Avoiding the conditions that can lead to deadlocks in the network is critical for using NoCs in real designs. Methods that can lead to deadlock-free operation with minimum power and area overhead are important for designing application-specific NoCs. A major class of deadlocks that occur in NoCs are due to the dependencies among the resources shared by different message types. In this work, we consider the problem of avoiding message-dependent deadlocks during the NoC topology synthesis phase. We show that by considering this issue during topology synthesis, we can obtain a significantly better NoC design than traditional methods, where the deadlock avoidance issue is dealt with separately. Our experiments on several SoC benchmarks show that our proposed scheme provides large reduction in NoC power consumption (an average of 38.5%) and NoC area (an average of 30.7%) when compared to traditional approaches Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
VLSI-SoC | 8 |
| 2006 | Stigmergic approaches applied to flexible fault-tolerant digital VLSI architectures
Danilo Pani, Luigi Raffo |
J. Parallel Distributed Comput. | 2 |
| 2005 | xpipes Lite: A Synthesis Oriented Design Library For Networks on ChipsabstractThe limited scalability of current bus topologies for systems on chips (SoCs) dictates the adoption of networks on chips (NoCs) as a scalable interconnection scheme. Current SoCs are highly heterogeneous in nature, denoting homogeneous, preconfigured NoCs as inefficient drop-in alternatives. While highly parametric, fully synthesizeable (soft) NoC building blocks appear as a good match for heterogeneous MPSoC architectures, the impact of instantiation-time flexibility on performance, power and silicon cost has not yet been quantified. The paper details /spl times/pipes Lite, a design flow for automatic generation of heterogeneous NoCs. /spl times/pipes Lite is based on highly customizable, high frequency and low latency NoC modules, that are fully synthesizeable. Synthesis results provide modules that are directly comparable, if not better, than the current published state-of-the-art NoCs in terms of area, power latency and target operating frequency measurements. Stergios Stergiou, Federico Angiolini, Salvatore Carta, Luigi Raffo, Davide Bertozzi, Giovanni De Micheli |
DATE | 4 |
| 2005 | Run-time Adaptive Resources Allocation and Balancing on Nanoprocessors ArraysabstractModern processor architectures try to exploit the different kind of parallelism that may be found even in general purpose applications. In this paper we present a new architecture based on an array of nanoprocessors that parallely and cooperatively support both Thread and Instruction level parallelism. A such architecture doesn't explicitly require any particular programming techniques since it has been developed to deal with standard sequential programs. Preliminary results on a model of the architecture show the feasibility of the proposed approach. Danilo Pani, Giuseppe Passino, Luigi Raffo |
DSD | 3 |
| 2004 | A Swarm Intelligence Based VLSI Multiplication-and-Add Scheme
Danilo Pani, Luigi Raffo |
PPSN | 2 |
| 2000 | Block-matching evaluation in digital architectures for motion estimationabstractIn this paper a comparison between different methods for block matching in image processing with respect to the efficacy of their digital VLSI implementation is presented. In this framework a new method based on limiting the role of mismatching pixels is proposed. The results obtained on different kinds of images show that the new method achieves the best trade-off between complexity and results. Luigi Raffo, Maria Paola Zizola |
ISCAS | 1 |
| 2000 | A micro-power mixed signal IC for battery-operated burglar alarm systemsabstractThe design of the standard CMOS IC core of a commercial wireless burglar alarm system is presented as an example of a very low-power analog VLSI design for battery-operated systems. The main constraint is battery life, which must be at least five years (with standard camera-battery). The chip is composed of a digital (decision) part and an analog interface with sensors. The entire chip absorbs 10μA. Measures on each single component and test on working environment show full functionality and complied with specifications. Even though the example is application specific, the design solutions and each single element can also be utilized in many other battery-operated low-frequency devices (e.g. environmental parameter monitoring). Silvio Bolliri, Paolo Porcu, Luigi Raffo |
ISLPED | 3 |
| 1998 | Analog computation for phase-based disparity estimation: continuous and discrete models
Bruno Crespi, Alex Cozzi, Luigi Raffo, Silvio P. Sabatini |
Mach. Vis. Appl. | 3 |
| 1998 | Analogue VLSI primitives for perceptual tasks in machine vision
Giacomo M. Bisio, Luigi Raffo, Silvio P. Sabatini |
Neural Comput. Appl. | 2 |
| 1998 | Analog VLSI circuits as physical structures for perception in early visual tasksabstractA variety of computational tasks in early vision can be formulated through lattice networks. The cooperative action of these networks depends on the topology of interconnections, both feedforward and recurrent ones. This paper shows that it is possible to consider a distinct general architectural solution for all recurrent computations of any given order. The Gabor-like impulse response of a second-order network is analyzed in detail, pointing out how a near-optimal filtering behavior in space and frequency domains can be achieved through excitatory/inhibitory interactions without impairing the stability of the system. These architectures can be mapped, very efficiently at transistor level, on very large scale integration (VLSI) structures operating as analog perceptual engines. The problem of hardware implementation of early vision tasks can, indeed, be tackled by combining these perceptual agents through suitable weighted sums. A 17-node analog current-mode VLSI circuit has been implemented on a CMOS 2 microm, NWELL, single-poly, and double-metal technology, to demonstrate the feasibility of the approach. Applications of the perceptual engine to various machine vision algorithms are proposed. Luigi Raffo, Silvio P. Sabatini, Gian Marco Bo, Giacomo M. Bisio |
IEEE Trans. Neural Networks | 1 |
| 1997 | An Analog VLSI Computational Engine for Early Vision Tasks
Giacomo M. Bisio, Gian Marco Bo, M. Confalone, Luigi Raffo, Silvio P. Sabatini, M. P. Zizola |
ICANN | 4 |
| 1997 | Functional Periodic Intracortical Couplings Induced by Structured Lateral Inhibition in a Linear Cortical NetworkabstractThe spatial organization of cortical axon and dendritic fields could be an interesting structural paradigm to obtain a functional specificity without postulating highly specific feedforward connections. In this article, we investigate the functional implications of recurrent intracortical inhibition when it occurs through clustered medium-range interconnection schemes (Wörgötter & Koch, 1991; Somogyi, 1989; Kritzer, Cowey, & Somogyi, 1992). Moreover, the interaction between the inhibitory schemes and visual orientation maps is explored. Assuming linearity, we show that clustered inhibitory mechanisms can trigger a propagation process that allows the development of extra (i.e., induced) interactions among the cortical sites involved in the recurrent loops. In addition, we point out how these interactions functionally modify the response of cortical simple cells and yield to highly structured Gabor-like receptive fields. This study should be considered not as a realistic biological model of the primary visual cortex but as an attempt to explain possible computational principles related to intracortical connectivity and to the underlying single-cell properties. Silvio P. Sabatini, Luigi Raffo, Giacomo M. Bisio |
Neural Comput. | 2 |
| 1997 | Design of an ASIP architecture for low-level visual elaborationsabstractWe consider the design process of VLSI systems dedicated to the real-time implementation of cooperative algorithms whose functionalities can be characterized by multilayer ensembles of simple elements which interact locally. These algorithms are related, even though not exclusively, to the implementation of various tasks in low-level machine vision. The starting point in the design process is the formulation of the sequential algorithm that computes the behavior of the system. Algorithmic transformations are performed to expose the parallelism originally present in the task. Given the description in terms of parallel loops, we partition the system and organize it as a set of processing units. The architectural structure of these units takes properly into account the algorithmic constraints on precision both in data representation and computation. The program flow implemented by our programmable architectural solution (ASIP) is an iterative sequence of multiply-and-accumulate operations performed in parallel. The programmability concerns both the structure/coefficients of the algorithm-depending on the specific application-and its computational parameters. The architecture's main blocks are described in VHDL and synthesized as a semi-custom chip, using standard tools. Following this procedure, we designed an ASIP core for performing real-time texture-based image segregation. Luigi Raffo, Silvio P. Sabatini, Mauro Mantelli, Alessandro De Gloria, Giacomo M. Bisio |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1996 | A recurrent neural architecture mimicking cortical preattentive vision systems
Giacomo Indiveri, Luigi Raffo, Silvio P. Sabatini, Giacomo M. Bisio |
Neurocomputing | 2 |
| 1995 | A neuromorphic architecture for cortical multilayer integration of early visual tasks
Giacomo Indiveri, Luigi Raffo, Silvio P. Sabatini, Giacomo M. Bisio |
Mach. Vis. Appl. | 2 |