Paolo Meloni

dblp:50/6206 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-8106-4641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021
YearPublicationVenuePosition
2026 SYNtzulA: Open Hardware for Near-Sensor SNN Inference
abstract
Spiking Neural Networks (SNNs) exploit event-driven processing to offer high energy efficiency when deploying Artificial Intelligence (AI) on wearable edge devices. However, specialized hardware is needed to fully take advantage of this potential, which, despite recent advances, remains expensive and not widely accessible. To address this, open-source Electronic Design Automation (EDA) tools and Process Design Kits (PDKs) offer a path to democratize the development of neuromorphic hardware. In this work, we present SYNtzulA, a system-on-chip designed for SNN acceleration, developed using the open-source IHP-SG13G2 130 nm PDK and the OpenROAD toolchain. The chip integrates a RISC-V softcore and a dedicated SNN accelerator, occupying approximately 6.8mm2including I/O pads. It operates at up to 125 MHz, reaching a throughput of 2 Giga Synaptic Operations per second (GSOP/s) with an energy consumption of 36.5 pJ per synaptic operation. The accelerator can exploit the sparsity of spike-based computation by skipping unnecessary operations, resulting in total energy consumption in the order of a few hundred nanojoules per inference in different use cases involving biosignal analysis.
Luca Martis, Gianluca Leone, Luigi Raffo, Paolo Meloni
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Enabling SNN-Based Near-MEA Neural Decoding with Channel Selection: An Open-HW Approach
abstract
Advancements in CMOS microelectrode array sensors have significantly improved sensing area and resolution, paving the way to accurate Brain-Machine Interfaces (BMIs). However, near-sensor neural decoding on implantable computing devices is still an open problem. A promising solution is provided by Spiking Neural Networks (SNNs), which leverage event sparsity to improve energy consumption. However, given the typical data rates involved, the workload related to I/O acquisition and spike encoding is dominant and limits the benefits achievable with event-based processing. In this work, we present two power-efficient implementations, on FPGA and ASIC, of a dedicated processor for the decoding of intracortical action potentials from primary motor cortex. The processor leverages lightweight sparse SNNs to achieve state-of-the-art accuracy. To limit the impact of I/O transfers on energy efficiency, we introduced a channel selection scheme that reduced bandwidth requirements by 3x and power consumption by 2.3x and 1.6x on the FPGA and ASIC, respectively, enabling inference at 0.446 μJ and 1.04 μJ, with no significant loss in accuracy. To promote broad adoption in a specialized, research-intensive domain, we have based our implementations on open-source EDA tools, low-cost hardware, and an open PDK.
Gianluca Leone, Luca Martis, Luigi Raffo, Paolo Meloni
DATE4
2025 SYNtzulu: A Tiny RISC-V-Controlled SNN Processor for Real-Time Sensor Data Analysis on Low-Power FPGAs
abstract
Spiking Neural Networks (SNNs) are energy- and performance-efficient tools that have been found to be very useful in AI applications at the edge. This paper introducesSYNtzulu, an SNN processing element designed to be used in low-cost and low-power FPGA devices for near-sensor data analysis. The system is equipped with a RISC-V subsystem responsible for controlling the input/output and setting runtime parameters, thus increasing its flexibility. We evaluated the system, which was implemented on a Lattice iCE40UP5K FPGA, in various use cases employing SNNs with accuracy comparable to the state-of-the-art.SYNtzuludissipates a maximum power of 12.05 mW when performing SNN inference, which can be reduced to an average of just 1.45 mW through the use of dynamic power management.
Gianluca Leone, Matteo Antonio Scrugli, Lorenzo Badas, Luca Martis, Luigi Raffo, Paolo Meloni
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 On-FPGA Spiking Neural Networks for Integrated Near-Sensor ECG Analysis
abstract
The identification of cardiac arrhythmias is a significant issue in modern healthcare and a major application for Artificial Intelligence (AI) systems based on artificial neural networks. This research introduces a real-time arrhythmia diagnosis system that uses a Spiking Neural Network (SNN) to classify heartbeats into five types of arrhythmias from a single-lead electrocardiogram (ECG) signal. The system is implemented on a custom SNN processor running on a low-power Lattice iCE40-UltraPlus FPGA. It was tested using the MIT-BIH dataset, and achieved accuracy results that are comparable to the most advanced SNN models, reaching 98.4% accuracy. The proposed modules take advantage of the energy efficiency of SNNs to reduce the average execution time to 4.32 ms and energy consumption to 50.98 uJ per classification.
Matteo Antonio Scrugli, Paola Busia, Gianluca Leone, Paolo Meloni
DATE4
2022 Integration of Energy Storage Systems within Modular Multilevel Converters for Medium-Voltage Distribution Networks
abstract
This paper presents two novel Modular Multilevel Converter (MMC) configurations for medium-voltage distribution networks, which are achieved by replacing some capacitive cells of each MMC branch with Supercapacitors (SCs) and Battery Packs (BPs), respectively. The first configuration (MMC-SC) is able to exploit SC energy content by preserving MMC branch voltage capability, thus ensuring a suitable dynamic decoupling between grid and DC-link power demands. Steady-state decoupling can be achieved by the second configuration (MMC-BP), which can charge/discharge BP as needed. MMC-SC and MMC-BP functionality is guaranteed by a multi-stage control system architecture, which has been developed in order to ensure proper MMC energy, current and voltage management at any operating conditions. The effectiveness of both MMC-SC and MMC-BP is verified through simulation studies, which regard step-changing active and reactive grid power profiles.
Paolo Meloni, Alessandro Serpi
IECON1
2020 Biosensing IoT Platform for Water Management in Vineyards
abstract
We present an IoT platform specifically developed to manage the use of water in the production of crops typical of Mediterranean geographic area. A large number of innovative sensors able to measure the water content and, thus, the hydration condition of grapes is deployed in the vineyard. The sensing mechanism is based on the transduction of relevant parameters acquired directly from the plant and not, as usually done, from the environment (e.g. soil). The sensor data are collected by a set of sensor nodes able to send data over long distances to a central hub which uploads the data to the cloud making the remote monitoring of the field possible. Trained operators will, thus, be able to implement irrigation strategies based on the actual state of the plant sampled with a high temporal and spatial density. This technological platform makes possible the implementation of deficit irrigation practices useful to save water and optimize production but also to properly drive the final characteristics of the grapes and of the wine. The platform requires minimal maintenance and its installation does not interfere with regular farming activities.
Silvia Loddo, M. Soccol, A. Perra, M. Ucchesu, Paolo Meloni, Massimo Barbaro, Mauro Lo Cascio, C. Sirca
ISCAS5
2019 Optimization and deployment of CNNs at the edge: the ALOHA experience
abstract
Deep learning (DL) algorithms have already proved their effectiveness on a wide variety of application domains, including speech recognition, natural language processing, and image classification. To foster their pervasive adoption in applications where low latency, privacy issues and data bandwidth are paramount, the current trend is to perform inference tasks at the edge. This requires deployment of DL algorithms on low-energy and resource-constrained computing nodes, often heterogenous and parallel, that are usually more complex to program and to manage without adequate support and experience. In this paper, we present ALOHA, an integrated tool flow that tries to facilitate the design of DL applications and their porting on embedded heterogenous architectures. The proposed tool flow aims at automating different design steps and reducing development costs. ALOHA considers hardware-related variables and security, power efficiency, and adaptivity aspects during the whole development process, from pre-training hyperparameter optimization and algorithm configuration to deployment.
Paolo Meloni, Daniela Loi, Paola Busia, Gianfranco Deriu, Andy D. Pimentel, Dolly Sapra, Todor P. Stefanov, Svetlana Minakova, Francesco Conti 0001, Luca Benini, Maura Pintor, Battista Biggio, Bernhard Moser 0001, Natalia Shepeleva, Nikos Fragoulis, Ilias Theodorakopoulos, Michael Masin, Francesca Palumbo
CF1
2019 A runtime-adaptive cognitive IoT node for healthcare monitoring
abstract
Wearable and energy efficient processing nodes, allowing for continuous remote monitoring of patient vital parameters, are mainstream in modern health-care practice. Most recent approaches to the development of such systems combine near-sensor data processing with cognitive computing, to improve at the same time communication efficiency, responsiveness and accuracy of the analysis of the sensed data. In this paper, we present a hardware-software architecture for a connected sensor-processing node that allows the set of in-place processing tasks to be executed to be remotely controllable by an external user. The designed system is capable of dynamically adapting its operating point to the selected computational load, to minimize power consumption. The benefits of the proposed approach are tested on a use-case involving ECG monitoring, that, when selected, performs ECG classification using a lightweigth convolutional neural network. Experimental results show how the proposed approach can provide more than 50% power consumption reduction for common ECG activity, with less than 2% memory footprint overhead and reconfiguring the system in less than 1 ms.
Matteo Antonio Scrugli, Daniela Loi, Luigi Raffo, Paolo Meloni
CF4
2018 NEURAghe: Exploiting CPU-FPGA Synergies for Efficient and Flexible CNN Inference Acceleration on Zynq SoCs
abstract
Deep convolutional neural networks (CNNs) obtain outstanding results in tasks that require human-level understanding of data, like image or speech recognition. However, their computational load is significant, motivating the development of CNN-specialized accelerators. This work presents NEURA ghe , a flexible and efficient hardware/software solution for the acceleration of CNNs on Zynq SoCs. NEURA ghe leverages the synergistic usage of Zynq ARM cores and of a powerful and flexible Convolution-Specific Processor deployed on the reconfigurable logic. The Convolution-Specific Processor embeds both a convolution engine and a programmable soft core, releasing the ARM processors from most of the supervision duties and allowing the accelerator to be controlled by software at an ultra-fine granularity. This methodology opens the way for cooperative heterogeneous computing: While the accelerator takes care of the bulk of the CNN workload, the ARM cores can seamlessly execute hard-to-accelerate parts of the computational graph, taking advantage of the NEON vector engines to further speed up computation. Through the companion NeuDNN SW stack, NEURA ghe supports end-to-end CNN-based classification with a peak performance of 169GOps/s, and an energy efficiency of 17GOps/W. Thanks to our heterogeneous computing model, our platform improves upon the state-of-the-art, achieving a frame rate of 5.5 frames per second (fps) on the end-to-end execution of VGG-16 and 6.6fps on ResNet-18.
Paolo Meloni, Alessandro Capotondi, Gianfranco Deriu, Michele Brian, Francesco Conti 0001, Davide Rossi 0001, Luigi Raffo, Luca Benini
ACM Trans. Reconfigurable Technol. Syst.1
2017 Real-Time neural signal decoding on heterogeneous MPSocs based on VLIW ASIPs
Paolo Meloni, Claudio Rubattu, Giuseppe Tuveri, Danilo Pani, Luigi Raffo, Francesca Palumbo
J. Syst. Archit.1
2012 Exploiting binary translation for fast ASIP design space exploration on FPGAs
abstract
Complex Application Specific Instruction-set Processors (ASIPs) expose to the designer a large number of degrees of freedom, posing the need for highly accurate and rapid simulation environments. FPGA-based emulators represent an alternative to software cycle-accurate simulators, preserving maximum accuracy and reasonable simulation times. The work presented in this paper aims at exploiting FPGA emulation within technology aware design space exploration of ASIPs. The potential speedup provided by reconfigurable logic is reduced by the overhead of RTL synthesis/implementation. This overhead can be mitigated by reducing the number of FPGA implementation processes, through the adoption of binary-level translation. Hereby we present a prototyping method that, given a set of candidate ASIP configurations, defines an overdimensioned ASIP architecture, capable of emulating all the design space points under evaluation. This approach is then evaluated with a design space exploration case study. Along with execution time, by coupling FPGA emulation with activity-based physical modeling, we can extract area/power/energy figures.
Sebastiano Pomata, Paolo Meloni, Giuseppe Tuveri, Luigi Raffo, Menno Lindwer
DATE2
2012 ASAM: Automatic Architecture Synthesis and Application Mapping
abstract
This paper focuses on mastering the automatic architecture synthesis and application mapping for heterogeneous massively-parallel MPSoCs based on customizable application-specific instruction-set processors (ASIPs). It presents an over-view of the research being currently performed in the scope of the European project ASAM of the ARTEMIS program. The paper briefly presents the results of our analysis of the main problems to be solved and challenges to be faced in the design of such heterogeneous MPSoCs. It explains which system, design, and electronic design automation (EDA) concepts seem to be adequate to resolve the problems and address the challenges. Finally, it introduces and briefly discusses the ASAM design-flow and its main stages.
Lech Józwiak, Menno Lindwer, Rosilde Corvino, Paolo Meloni, Laura Micconi, Jan Madsen, Erkan Diken, Deepak Gangadharan, Roel Jordans, Sebastiano Pomata, Paul Pop, Giuseppe Tuveri, Luigi Raffo
DSD4
2012 System Adaptivity and Fault-Tolerance in NoC-based MPSoCs: The MADNESS Project Approach
abstract
Modern embedded systems increasingly require adaptive run-time management. The system may adapt the mapping of the applications in order to accommodate the current workload conditions, to balance load for efficient resource utilization, to meet quality of service agreements, to avoid thermal hot-spots and to reduce power consumption. As the possibility of experiencing run-time faults becomes increasingly relevant with deep-sub-micron technology nodes, in the scope of the MADNESS project, we focus particularly on the problem of graceful degradation by dynamic remapping in presence of run-time faults. In this paper, we summarize the major results achieved in the MADNESS project until now regarding the system adaptivity and fault tolerant processing. We report the first results of the integration between platform level and middleware level support for adaptivity and fault tolerance. A case study demonstrates the survival ability of the system via a low-overhead process migration mechanism and a near-optimal online remapping heuristic.
Paolo Meloni, Giuseppe Tuveri, Luigi Raffo, Emanuele Cannella, Todor P. Stefanov, Onur Derin, Leandro Fiorin, Mariagiovanna Sami
DSD1
2010 Exploiting FPGAs for technology-aware system-level evaluation of multi-core architectures
abstract
The hardware-software co-development of modern complex MPSoC computing platforms exposes to the designer a huge complexity, resulting from the combination of vastly different architectural possibilities with strict demands posed by the target applications. To handle this complexity, highly accurate but rapid prototyping/evaluation environments need to be developed, that would possibly be able to provide an effective measurement of the system under design as soon as possible, allowing to comply with current time-to-market. While software-based fully cycle-accurate simulators do not seem to represent anymore an adequate solution to solve this issue, the attention has been recently shifted to the adoption of hardware emulators in the early stages of the design flow. In this work, we present an emulation framework for library-based semiautomatic instantiation of complex multi-core platforms that exploits FPGA devices to provide detailed functional information on the platform under development, and at the same time using hardware execution traces with technology-related analytical models to extract, already at system-level, physical metrics on power consumption, maximum operating frequency and area occupation of a prospective ASIC implementation of the system. Two prospective use case scenarios are presented to validate the usefulness of the presented framework: the first one analyzes the mapping and the scalability of a highly parallel application over a 2D homogeneous mesh architecture for increasing number of processors, while the second one employs the emulation infrastructure inside a design space exploration flow for the configuration of some interconnection network parameters.
Simone Secchi, Paolo Meloni, Luigi Raffo
ISPASS2
2010 Enabling fast Network-on-Chip topology selection: an FPGA-based runtime reconfigurable prototyper
abstract
The complexity of modern interconnect architecture design requires highly accurate and rapid simulation environments. FPGA-based emulators have been proposed as an alternative to software cycle-accurate simulators, preserving maximum accuracy and reasonable simulation times. However, the potential speedup is reduced by the time overhead needed for RTL synthesis/implementation. This paper proposes runtime reconfiguration of the architecture to push the hardware emulation one step further, by reducing the number of FPGA implementation processes to be run. To this aim, this work presents an algorithm that synthesizes, for a set of candidate architectural configurations, a connection topology capable of reconfiguring itself via software to emulate all the design space points under evaluation. We present the actual reconfiguration algorithm, the CAD tools and the hardware mechanisms that implement it. The design capabilities provided by this approach are evaluated with a design space exploration case study.
Paolo Meloni, Simone Secchi, Luigi Raffo
VLSI-SoC1
2007 On the impact of serialization on the cache performances in Network-on-Chip based MPSoCs
abstract
Network on Chip architectures are proposed as a solution to overcome functional and physical scalability shown by shared bus based MPSoC architecture. Unfortunately to implement and efficient communication infrastructure, the designer has to set a lot of parameters. An exhaustive knowledge of how the chosen settings influence the overall behaviour of the designed system is then mandatory. Aim of this paper is to discuss the relationship between the performances of a NoC and its configuration parameters in the case of traffic generated by cache operations (block replacements). We paid special attention to investigate the impact of the serialization factor, that was already not clearly assessed in literature for this important case study. A numerical analysis, referring to an actual implementation of the NoC on a state-of-the-art 65 nm technological process was performed. The obtained results were used to report an energy and execution time exploration over the complete design space of interest.
Paolo Meloni, Giovanni Busonera, Salvatore Carta, Luigi Raffo
DSD1
2007 NoC Design and Implementation in 65nm Technology
abstract
As embedded computing evolves towards ever more powerful architectures, the challenge of properly interconnecting large numbers of on-chip computation blocks is becoming prominent. Networks-on-chip (NoCs) have been proposed as a scalable solution to both physical design issues and increasing bandwidth demands. However, this claim has not been fully validated yet, since the design properties and tradeoffs of NoCs have not been studied in detail below the 100 nm threshold. This work is aimed at shedding light on the opportunities and challenges, both expected and unexpected, of NoC design in nanometer CMOS. We present fully working 65 nm NoC designs, a complete NoC synthesis flow and detailed scalability analysis
Antonio Pullini, Federico Angiolini, Paolo Meloni, David Atienza 0001, Srinivasan Murali, Luigi Raffo, Giovanni De Micheli, Luca Benini
NOCS3
2007 A Layout-Aware Analysis of Networks-on-Chip and Traditional Interconnects for MPSoCs
abstract
The ever-shrinking lithographic technologies available to chip designers enable performance and functionality breakthroughs; yet, they bring new hard problems. For example, multiprocessor systems-on-chip featuring several processing elements can be conceived, but efficiently interconnecting them while keeping the design complexity manageable is a challenge. Traditional buses are easy to deploy, but cannot provide enough bandwidth for such complex systems. A departure from legacy architectures is therefore called for. One radical path is represented by packet-switching networks-on-chip, whereas a more conservative approach interleaves bandwidth-rich components (e.g., crossbars) within the preexisting fabrics. This paper is aimed at analyzing the strengths and weaknesses of these alternative approaches by performing a thorough analysis based on actual chip floorplans after the interconnection place&route stages and after a clock tree has been distributed across the layout. Performance, area, and power results will be discussed while keeping an eye on the scalability prospects in future technology nodes
Federico Angiolini, Paolo Meloni, Salvatore Carta, Luigi Raffo, Luca Benini
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Synthesis of Predictable Networks-on-Chip-Based Interconnect Architectures for Chip Multiprocessors
abstract
Today, chip multiprocessors (CMPs) that accommodate multiple processor cores on the same chip have become a reality. As the communication complexity of such multicore systems is rapidly increasing, designing an interconnect architecture with predictable behavior is essential for proper system operation. In CMPs, general-purpose processor cores are used to run software tasks of different applications and the communication between the cores cannot be precharacterized. Designing an efficient network-on-chip (NoC)-based interconnect with predictable performance is thus a challenging task. In this paper, we address the important design issue of synthesizing the most power efficient NoC interconnect for CMPs, providing guaranteed optimum throughput and predictable performance for any application to be executed on the CMP. In our synthesis approach, we use accurate delay and power models for the network components (switches and links) that are obtained from layouts of the components using industry standard tools. The synthesis approach utilizes the floorplan knowledge of the NoC to detect timing violations on the NoC links early in the design cycle. This leads to a faster design cycle and quicker design convergence across the high-level synthesis approach and the physical implementation of the design. We validate the design flow predictability of our proposed approach by performing a layout of the NoC synthesized for a 25-core CMP. Our approach maintains the regular and predictable structure of the NoC and is applicable in practice to existing NoC architectures.
Srinivasan Murali, David Atienza 0001, Paolo Meloni, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo
IEEE Trans. Very Large Scale Integr. Syst.3
2006 Contrasting a NoC and a traditional interconnect fabric with layout awareness
abstract
Increasing miniaturization is posing multiple challenges to electronic designers. In the context of multi-processor system-on-chips (MPSoCs), we focus on the problem of implementing efficient interconnect systems for devices which are ever more densely packed with parallel computing cores. Easily seen that traditional buses can not provide enough bandwidth, a revolutionary path to scalability is provided by packet-switched network-on-chips (NoCs), while a more conservative approach dictates the addition of bandwidth-rich components (e.g. crossbars) within the preexisting fabrics. While both alternatives have already been explored, a thorough contrastive analysis is still missing. In this paper, we bring crossbar and NoC designs to the chip layout level in order to highlight the respective strengths and weaknesses in terms of performance, area and power, keeping an eye on future scalability
Federico Angiolini, Paolo Meloni, Salvatore Carta, Luca Benini, Luigi Raffo
DATE2
2006 Designing application-specific networks on chips with floorplan information
abstract
With increasing communication demands of processor and memory cores in Systems on Chips (SoCs), scalable Networks on Chips (NoCs) are needed to interconnect the cores. For the use of NoCs to be feasible in today's industrial designs, a custom-tailored, application-specific NoC that satisfies the design objectives and constraints of the targeted application domain is required. In this work, we present a design methodology that automates the synthesis of such application-specific NoC architectures. We present a floorplan aware design method that considers the wiring complexity of the NoC during the topology synthesis process. This leads to detecting timing violations on the NoC links early in the design cycle and to have accurate power estimations of the interconnect. We incorporate mechanisms to prevent deadlocks during routing, which is critical for proper operation of NoCs. We integrate the NoC synthesis method with an existing design flow, automating NoC synthesis, generation, simulation and physical design processes. We also present ways to ensure design convergence across the levels. Experiments on several SoC benchmarks are presented, which show that the synthesized topologies provide a large reduction in network power consumption (2.78x on average) and improvement in performance (1.59x on average) over the best mesh and mesh-based custom topologies. An actual layout of a multimedia SoC with the NoC designed using our methodology is presented, which shows that the designed NoC supports the required frequency of operation (close to 900 MHz) without any timing violations. We could design the NoC from input specifications to layout in 4 hours, a process that usually takes several weeks.
Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo
ICCAD2
2006 Designing Message-Dependent Deadlock Free Networks on Chips for Application-Specific Systems on Chips
abstract
Networks on chip (NoC) has emerged as the paradigm for designing scalable communication architecture for systems on chips (SoCs). Avoiding the conditions that can lead to deadlocks in the network is critical for using NoCs in real designs. Methods that can lead to deadlock-free operation with minimum power and area overhead are important for designing application-specific NoCs. A major class of deadlocks that occur in NoCs are due to the dependencies among the resources shared by different message types. In this work, we consider the problem of avoiding message-dependent deadlocks during the NoC topology synthesis phase. We show that by considering this issue during topology synthesis, we can obtain a significantly better NoC design than traditional methods, where the deadlock avoidance issue is dealt with separately. Our experiments on several SoC benchmarks show that our proposed scheme provides large reduction in NoC power consumption (an average of 38.5%) and NoC area (an average of 30.7%) when compared to traditional approaches
Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo
VLSI-SoC2