Francesco Barchi

dblp:173/1134 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-5155-6883ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Graph Reasoning into Lightweight CNNs for Near-Sensor Point Cloud Corruption Detection
abstract
Real-world point cloud corruption on automotive LiDAR lenses can significantly degrade the reliability of down-stream perception, particularly object detection models trained on clean data, which may yield overconfident false positives. To address this, we propose a near-sensor gating module that classifies incoming point clouds as either clean or contaminated using a teacher–student knowledge distillation pipeline. A Graph Attention Network (GAT), trained directly on raw point clouds, serves as the teacher. On real-world contaminated LiDAR data, the distilled student achieves an average F1-score of 0.83, closely matching the GAT teacher’s 0.88, and consistently outperforming other supervised baselines across diverse test environments. Importantly, the student’s 2D-CNN architecture reduces preprocessing complexity from O(n log n) of graph construction to O(n), enabling faster and more efficient point cloud handling. The student model is quantized to 16-bit fixed-point and deployed on a GAP8 (RISC-V) platform. It achieves an inference latency of 26 milliseconds, consumes only 210µJ per inference, and fits within 84KB of L2 memory. This makes the proposed solution a practical and resource-efficient near-sensor gating module for robust, contaminant-aware perception in embedded automotive systems. The implementation will be available at https://gitlab.com/ecs-lab/distilling-pointcloud-corruption
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva
DATE3
2026 Targeting language models for compile-time computing resource optimization: A novel approach based on masked graph autoencoders
Federico Cichetti, Emanuele Parisi, Andrea Acquaviva, Francesco Barchi
Data Knowl. Eng.4
2026 A transformer-based approach for source code classification for heterogeneous device mapping
abstract
The optimization of code allocation for heterogeneous architectures, such as Central Processing Units (CPUs) and Graphics Processing Units (GPUs), remains challenging due to the limitations of traditional compiler heuristics and existing machine learning approaches. This paper presents a systematic evaluation of Large Language Models (LLMs) for classifying source code execution targets in heterogeneous device mapping. We fine-tune and compare six models: Distilled Bidirectional Encoder Representations from Transformers (DistilBERT), Code Bidirectional Encoder Representations from Transformers (CodeBERT), Code Bidirectional Encoder Representations from Transformers with RoBERTa (Robustly Optimized BERT Pretraining Approach) architecture (CodeBERTa), CodeT5, jTrans, and Deep Learning Low Level Virtual Machine (DeepLLVM), trained on Open Computing Language (OpenCL) kernels. Results show that general-purpose LLMs achieve up to 92.8% accuracy, matching or surpassing code-specific models, and outperform the previous state of the art (DeepLLVM) by up to 5%. Our findings indicate that LLMs pre-trained on general text are not necessarily inferior to code-specialized models, with tokenizer design and pre-training objectives impacting performance more than domain specialization. These results demonstrate the effectiveness of Transformer-based LLMs as a state-of-the-art approach for source code classification in heterogeneous computing contexts. • LLMs achieve SOTA in source code classification for heterogeneous device mapping. • Code-specific LLMs don’t always beat general ones–their edge depends on the dataset. • Transformer LLMs beat DeepLLVM in code-to-architecture mapping; CodeBERTa leads overall.
Marco Siino, Emanuele Parisi, Francesco Barchi, Andrea Acquaviva, Andrea Bartolini
Eng. Appl. Artif. Intell.3
2026 An OpenSSL Engine for Secure and Transparent Offloading of Cryptographic Operations to OP-TEE in Embedded IoT Platforms
abstract
Embedded IoT platforms are frequently deployed in hostile or physically exposed environments, where compromise of the operating system is a realistic threat. In conventional deployments, OpenSSL executes entirely in user space, leaving cryptographic keys and intermediate material potentially exposed in the presence of a compromised OS or privileged malware. This work presents a portable OpenSSL engine that enables secure and transparent offloading of cryptographic operations to OP-TEE, a software stack leveraging a Trusted Execution Environment (TEE) to confine cryptographic key material and security-critical computations within secure-world memory without modifying existing applications or altering the OpenSSL EVP interface. The proposed engine uses a GlobalPlatform Client API implementation as an underlying communication layer and supports both Copy-Based (CB) data transfer and a Shared-Memory (SM) mode, enabling performance tuning under the resource constraints typical of embedded IoT platforms. An experimental evaluation on an NXP i.MX7 industrial gateway shows that secure-world execution introduces bounded overhead dominated by REE–TEE transitions and memory-management operations. For bandwidth-intensive primitives such as SHA-256, SM mode improves throughput by approximately 10–15% over CB transfers. For symmetric encryption algorithms such as AES-256-CBC, SM communication provides a substantial throughput improvement of approximately 40%, indicating that data-movement overhead plays a significant role on embedded platforms. In contrast, asymmetric primitives (e.g. RSA-1024) remain largely insensitive to transfer optimisations because they operate on small, fixed-size operands, making the overall execution time dominated by invocation overheads and internal big-number computations rather than data movement.
Tina Sayarmoafi, Francesco Barchi, Lorenzo Bottaccioli, Andrea Acquaviva, Edoardo Patti, Paolo Montuschi, Luca Barbierato
IEEE Internet Things J.2
2025 ANZIL: Attention-based network for zero-risk inspection of LiDAR point cloud in self-driving cars
abstract
LiDAR is widely used in autonomous vehicle perception, and its performance relies on algorithmic confidence. However, sensor contamination can lead to catastrophic mistakes in downstream tasks, such as incorrect object detections occurring with high confidence. This underscores the need for point cloud contaminant detection to identify data reliability before downstream processing. To overcome this, we propose a model-agnostic approach that integrates contaminant detection with downstream tasks, using graph representations and attention networks. Trained on real contaminated LiDAR, the contaminant detector complements any existing clean-trained downstream model. Point clouds passing through the contaminant detector are discarded if contamination is detected. We propose a novel cost-benefit methodology, evaluating contaminant detectors in downstream processing. The benefit measures the proportion of discarded frames that would have caused high-confidence errors in the downstream task. The cost includes the model misalignment cost, representing frames wrongly discarded despite being correctly processed, and the true cost, which reflects uncontaminated frames falsely identified as contaminated. The contaminant detector was tested on 78,000 frames with water, dust, mud, salt, oil, and foam contamination from tunnel and outdoor locations. It achieved an F1-score of 0.884 and a recall of 0.957 on a static dataset, and a 0.975 F1-score in unseen real-world environments, demonstrating high sensitivity. In model-agnostic evaluation, the method is applicable to any object detector and reduces catastrophic mistakes by at least 97%, successfully identifying 4,046 failure cases. Deployed on the NVIDIA Jetson AGX Xavier, it achieved inference in under 5 ms and point cloud transformation in 200 ms, making it suitable for edge processing. This enhances sensor reliability, enables automatic cleaning, and improves vehicle safety.
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva
Expert Syst. Appl.3
2025 Unleashing OpenTitan's Potential: a Silicon-Ready Embedded Secure Element for Root of Trust and Cryptographic Offloading
abstract
The rapid advancement and exploration of open-hardware RISC-V platforms are catalyzing substantial changes across critical sectors, including autonomous vehicles, smart-city infrastructure, and medical devices. Within this technological evolution, OpenTitan emerges as a groundbreaking open-source RISC-V design, renowned for its comprehensive security toolkit and role as a stand-alone system-on-chip (SoC). OpenTitan encompasses different SoC implementations such as Earl Grey, 1 fully implemented and silicon proven, and Darjeeling, 2 announced but not yet fully implemented. The former targets a stand-alone SoC implementation; the latter is oriented towards an integrable implementation. Therefore, the literature currently lacks a silicon-ready embedded implementation of an open-source Root of Trust despite the effort made by lowRISC on the Darjeeling implementation of OpenTitan. We address the limitations of existing implementations, focusing on optimizing data transfer latency between memory and cryptographic accelerators to prevent under-utilization and ensure efficient task acceleration. Our contributions include a comprehensive methodology for integrating custom extensions and intellectual properties (IPs) into the Earl Grey architecture, architectural enhancements for system-level integration, support for varied boot modes, and improved data movement across the platform. These advancements facilitate the deployment of OpenTitan in broader SoCs, even in scenarios lacking specific technology-dependent IPs, providing a deployment-ready research vehicle for the community. We integrated the extended Earl Grey architecture into a reference architecture in 22-nm FDX technology node. Then, we benchmarked the enhanced architecture’s performance, analyzing the latency introduced by the external memory hierarchic levels, presenting significant improvements in cryptographic processing speed, achieving up to 2.7 x speedup for SHA-256/HMAC and 1.6 x for AES accelerators compared with baseline Earl Grey architecture.
Maicol Ciani, Emanuele Parisi, Alberto Musa, Francesco Barchi, Andrea Bartolini, Ari Kulmala, Rafail Psiakis, Angelo Garofalo, Andrea Acquaviva, Davide Rossi 0001
ACM Trans. Embed. Comput. Syst.4
2024 TinyLid: a RISC-V accelerated Neural Network For LiDAR Contaminant Classification in Autonomous Vehicle
abstract
LiDAR plays a critical role in autonomous car perception. Hence, the robustness of LiDAR data is imperative. However, malfunctions resulting from sensor cover contaminants are unavoidable and can lead to erroneous data that slowly degrade performance.
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva
CF3
2024 Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-Chips
abstract
RISC-V open-source systems are emerging in deployment scenarios where safety and security are critical. OpenTitan is an open-source silicon root-of-trust designed to be deployed in a wide range of systems, from high-end to deeply embedded secure environments. Despite the availability of various cryptographic hardware accelerators that make OpenTitan suitable for offloading cryptographic workloads from the main processor, there has been no accurate and quantitative establishment of the benefits derived from using OpenTitan as a secure accelerator. This paper addresses this gap by thoroughly analysing strengths and inefficiencies when offloading cryptographic workloads to OpenTitan. The focus is on three key IPs --- HMAC, AES, and OpenTitan Big Number accelerator (OTBN) --- which can accelerate four security workloads: Secure Hash Functions, Message Authentication Codes, Symmetric cryptography, and Asymmetric cryptography. For every workload, we develop a bare-metal driver for the OpenTitan accelerator and analyze its efficiency when computation is offloaded from a RISC-V application core within a System-on-Chip designed for secure Cyber-Physical Systems applications. Finally, we assess it against a software implementation on the application core. The characterization was conducted on a cycle-accurate RTL simulator of the System-on-Chip (SoC). Our study demonstrates that OpenTitan significantly outperforms software implementations, with speedups ranging from 4.3x to 12.5x. However, there is potential for even greater gains as the current OpenTitan utilizes a fraction of the accelerator bandwidths, which ranges from 16% to 61%, depending on the memory being accessed and the accelerator used. Our results open the way to the optimization of OpenTitan-based secure platforms, providing design guidelines to unlock the full potential of its accelerators in secure applications.
Emanuele Parisi, Alberto Musa, Maicol Ciani, Francesco Barchi, Davide Rossi 0001, Andrea Bartolini, Andrea Acquaviva
CF4
2024 TitanCFI: Toward Enforcing Control-Flow Integrity in the Root -of- Trust
abstract
Modern RISC-V platforms control and monitor security-critical systems such as industrial controllers and autonomous vehicles. While these platforms feature a Root-of-Trust (RoT) to store authentication secrets and enable secure boot technologies, they often lack Control-Flow Integrity (CFI) enforcement and are vulnerable to cyber-attacks which divert the control flow of an application to trigger malicious behaviours. Recent techniques to enforce CFI in RISC-V systems include ISA modifications or custom hardware IPs, all requiring ad-hoc binary toolchains or design of CFI primitives in hardware. This paper proposes TitanCFI, a novel approach to enforce CFI in the RoT. TitanCFI modifies the commit stage of the protected core to stream control flow instructions to the RoT and it integrates the CFI enforcement policy in the RoT firmware. Our approach enables maximum reuse of the hardware resource present in the System-on-Chip (SoC), and it avoids the design of custom IPs and the modification of the compilation toolchain, while exploiting the RoT tamper-proof storage and cryptographic accelerators to secure CFI metadata. We implemented the proposed architecture on a modern RISC-V SoC along with a return address protection policy in the RoT, and benchmarked area and runtime overhead. Experimental results show that TitanCFI achieves overhead comparable to SoA hardware CFI solutions for most benchmarks, with lower area overhead, resulting in 1 % of additional area occupation.
Emanuele Parisi, Alberto Musa, Simone Manoni, Maicol Ciani, Davide Rossi 0001, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva
DATE6
2024 DeepCodeGraph: A Language Model for Compile-Time Resource Optimization Using Masked Graph Autoencoders
Federico Cichetti, Emanuele Parisi, Andrea Acquaviva, Francesco Barchi
NLDB (1)4
2024 TitanSSL: Towards Accelerating OpenSSL in a Full RISC-V Architecture Using OpenTitan Root-of-Trust
Alberto Musa, Franco Volante, Emanuele Parisi, Luca Barbierato, Edoardo Patti, Andrea Bartolini, Andrea Acquaviva, Francesco Barchi
SAFECOMP8
2024 Enhancing workplace safety: A flexible approach for personal protective equipment monitoring
abstract
Workplace safety is a prominent concern, motivating researchers across diverse disciplines to investigate valuable ways to address its challenges. However, creating an efficient system to address this issue remains a significant challenge. Since many accidents happen due to improper usage or complete removal of Personal Protective Equipment (PPE), one straightforward method for enhancing workplace security involves monitoring their usage This paper introduces an Operator Area Network (OAN) system which improves the existing solutions by increasing portability across different users and environments, non-intrusiveness and privacy. To enhance robustness in detecting the situations in which PPEs are not used correctly, we take advantage of Machine Learning to analyse the received signal strength indicator (RSSI) between PPEs in the same OAN The novelty of this work is that it does not exploit RSSI as a proxy of the distance but instead recognises a signature of the correct wearing of the PPE By employing this system, employers can effectively ensure the proper usage of PPE devices at their worksites while also minimising any adverse effects on workers’ comfort and reducing the setup burden for employers. The system runs a Support Vector Machine (SVM) model several times per second and employs a post-processing algorithm to enhance its initial accuracy further As a result, the system effectively reduces false positives by about 80% and swiftly detects instances of improper usage of the worker’s PPE, raising the alarm in less than seven seconds. Moreover, the post-processing algorithm can be customised to meet the specific needs of different use cases, allowing for a flexible trade-off between the detection time interval and the overall accuracy of the detection system.
Alessia Pisu, Nicola Elia, Livio Pompianu, Francesco Barchi, Andrea Acquaviva, Salvatore Carta
Expert Syst. Appl.4
2024 Energy efficient and low-latency spiking neural networks on embedded microcontrollers through spiking activity tuning
abstract
Abstract In this work, we target the efficient implementation of spiking neural networks (SNNs) for low-power and low-latency applications. In particular, we propose a methodology for tuning SNN spiking activity with the objective of reducing computation cycles and energy consumption. We performed an analysis to devise key hyper-parameters, and then we show the results of tuning such parameters to obtain a low-latency and low-energy embedded LSNN (eLSNN) implementation. We demonstrate that it is possible to adapt the firing rate so that the samples belonging to the most frequent class are processed with less spikes. We implemented the eLSNN on a microcontroller-based sensor node and we evaluated its performance and energy consumption using a structural health monitoring application processing a stream of vibrations for damage detection (i.e. binary classification). We obtained a cycle count reduction of 25% and an energy reduction of 22% with respect to a baseline implementation. We also demonstrate that our methodology is applicable to a multi-class scenario, showing that we can reduce spiking activity between 68 and 85% at iso-accuracy.
Francesco Barchi, Emanuele Parisi, Luca Zanatta, Andrea Bartolini, Andrea Acquaviva
Neural Comput. Appl.1
2023 Directly-trained Spiking Neural Networks for Deep Reinforcement Learning: Energy efficient implementation of event-based obstacle avoidance on a neuromorphic accelerator
abstract
Spiking Neural Networks (SNN) promise extremely low-power and low-latency inference on neuromorphic hardware. Recent studies demonstrate the competitive performance of SNNs compared with Artificial Neural Networks (ANN) in conventional classification tasks. In this work, we present an energy-efficient implementation of a Reinforcement Learning (RL) algorithm using SNNs to solve an obstacle avoidance task performed by an Unmanned Aerial Vehicle (UAV), taking a Dynamic Vision Sensor (DVS) as event-based input. We train the SNN directly, improving upon state-of-art implementations based on hybrid (not directly trained) SNNs. For this purpose, we devise an adaptation of the Spatio-Temporal Backpropagation algorithm (STBP) for RL. We then compare the SNN with a state-of-art Convolutional Neural Network (CNN) designed to solve the same task. To this aim, we train both networks by exploiting a photorealistic training pipeline based on AirSim. To achieve a realistic latency and throughput assessment for embedded deployment, we designed and trained three different embedded SNN versions to be executed on state-of-art neuromorphic hardware, targeting state-of-the-art. We compared SNN and CNN in terms of obstacle avoidance performance showing that the SNN algorithm achieves better results than the CNN with a factor of 6× less energy. We also characterise the different SNN hardware implementations in terms of energy and spiking activity.
Luca Zanatta, Alfio Di Mauro, Francesco Barchi, Andrea Bartolini, Luca Benini, Andrea Acquaviva
Neurocomputing3
2022 Smart Contracts for Certified and Sustainable Safety-Critical Continuous Monitoring Applications
Nicola Elia, Francesco Barchi, Emanuele Parisi, Livio Pompianu, Salvatore Carta, Andrea Bartolini, Andrea Acquaviva
ADBIS2
2022 Meet Monte Cimone: exploring RISC-V high performance compute clusters
abstract
The new open and royalty-free RISC-V ISA is attracting interest across the whole computing continuum, from microcontrollers to supercomputers. High-performance RISC-V processors and accelerators have been announced, but RISC-V-based HPC systems will need a holistic co-design effort, spanning memory, storage hierarchy interconnects and full software stack. In this paper, we describe Monte Cimone, a fully-operational multi-blade computer prototype and hardware-software test-bed based on U740, a double precision capable multi-core, 64 bit RISC-V SoC. Monte Cimone does not aim to achieve strong floating point performance, but it was built with the purpose of "priming the pipe" and exploring the challenges of integrating a multi-node RISC-V cluster capable of providing an HPC production stack including interconnect, storage and power monitoring infrastructure on RISC-V hardware. We present the results of our hardware/software integration effort, which demonstrate a remarkable level of software and hardware readiness and maturity - showing that the first-generation of RISC-V HPC machines may not be so far in the future.
Federico Ficarelli, Andrea Bartolini, Emanuele Parisi, Francesco Beneventi, Francesco Barchi, Daniele Gregori, Fabrizio Magugliani, Marco Cicala, Cosimo Gianfreda, Daniele Cesarini, Andrea Acquaviva, Luca Benini
CF5
2022 Artificial versus spiking neural networks for reinforcement learning in UAV obstacle avoidance
abstract
Spiking Neural Networks (SNN) are gaining more interest from the scientific community thanks to the promise of greater energy-efficient and greater computational power. This poses several challenges as today's SNN training for RL is based on Artificial Neural Network (ANN) training and then conversion from ANN to SNN, which does not leverage SNN event-based processing inherent capabilities. The present work compares an ANN and an SNN in an event-camera-based obstacle avoidance task, trained with Reinforcement Learning (RL) using the Deep Q-Learning (DQL) algorithm. We create an experimental setup composed of Unreal Engine 4, AirSim, and an event camera that simulates a real-world obstacle avoidance environment. Additionally, we train an SNN with a gradient-based training method enabling the use of all their expressiveness even in the training phase, showing comparable performance between the ANN and the SNN. To the best of our knowledge, we are the first that implements an entire realistic pipeline with a photo-realistic simulator (Airsim) and train an SNN without converting it from a pre-trained ANN.
Luca Zanatta, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva
CF2
2022 Making the Most of Scarce Input Data in Deep Learning-Based Source Code Classification for Heterogeneous Device Mapping
abstract
Despite its relatively recent history, deep learning (DL)-based source code analysis is already a cornerstone in machine learning for compiler optimization. When applied to the classification of pieces of code to identify the best computational unit in a heterogeneous Systems-on-Chip, it can be effective in supporting decisions that a programmer has otherwise to take manually. Several techniques have been proposed exploiting different networks and input information, prominently sequence-based and graph-based representations, complemented by auxiliary information typically related to payload and device configuration. While the accuracy of DL methods strongly depends on the training and test datasets, so far no exhaustive and statistically meaningful analysis has been done on its impact on the results and on how to effectively extract the available information. This is relevant also considering the scarce availability of source code datasets that can be labeled by profiling on heterogeneous compute units. In this article, we first present such a study, which leads us to devise the contribution of code sequences and auxiliary inputs separately. Starting from this analysis, we then demonstrate that by using the normalization of auxiliary information, it is possible to improve state-of-the-art results in terms of accuracy. Finally, we propose a novel approach exploiting Siamese networks that further improve mapping accuracy by increasing the cardinality of the dataset, thus compensating for its relatively small size.
Emanuele Parisi, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Source Code Classification for Energy Efficiency in Parallel Ultra Low-Power Microcontrollers
abstract
The analysis of source code through machine learning techniques is an increasingly explored research topic aiming at increasing smartness in the software toolchain to exploit modern architectures in the best possible way. In the case of low-power, parallel embedded architectures, this means finding the configuration, for instance in terms of the number of cores, leading to minimum energy consumption. Depending on the kernel to be executed, the energy optimal scaling configuration is not trivial. While recent work has focused on general-purpose systems to learn and predict the best execution target in terms of the execution time of a snippet of code or kernel (e.g. offload OpenCL kernel on multicore CPU or GPU), in this work we focus on static compile-time features to assess if they can be successfully used to predict the minimum energy configuration on PULP, an ultra-low-power architecture featuring an on-chip cluster of RISC-V processors. Experiments show that using machine learning models on the source code to select the best energy scaling configuration automatically is viable and has the potential to be used in the context of automatic system configuration for energy minimisation.
Emanuele Parisi, Francesco Barchi, Andrea Bartolini, Giuseppe Tagliavini, Andrea Acquaviva
DATE2
2021 Exploration of Convolutional Neural Network models for source code classification
Francesco Barchi, Emanuele Parisi, Gianvito Urgese, Elisa Ficarra, Andrea Acquaviva
Eng. Appl. Artif. Intell.1
2019 Code Mapping in Heterogeneous Platforms Using Deep Learning and LLVM-IR
abstract
Modern heterogeneous platforms require compilers capable of choosing the appropriate device for the execution of program portions. This paper presents a machine learning method designed for supporting mapping decisions through the analysis of the program source code represented in LLVM assembly language (IR) for exploiting the advantages offered by this generalised and optimised representation. To evaluate our solution, we trained an LSTM neural network on OpenCL kernels compiled in LLVM-IR and processed with our tokenizer capable of filtering less-informative tokens. We tested the network that reaches an accuracy of 85% in distinguishing the best computational unit.
Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva
DAC1
2018 Impact of graph partitioning on SNN placement for a multi-core neuromorphic architecture: work-in-progress
Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva
CASES1
2018 Multiple alignment of packet sequences for efficient communication in a many-core neuromorphic system: work-in-progress
Gianvito Urgese, Luca Peres, Francesco Barchi, Enrico Macii, Andrea Acquaviva
CASES3
2018 Directed Graph Placement for SNN Simulation into a multi-core GALS Architecture
abstract
In this paper, we present a methodology for effi-ciently mapping neural networks over a neuromorphic computing architecture. The target architecture is a globally asynchronous locally synchronous (GALS) multi-core designed for simulating spiking neural networks (SNN) in real-time, that is spike timings should be the same as in the human brain. The SNN is implemented as a set of concurrent tasks modelling the behaviour of biological neurons, which are executed on the processing cores and communicate through spikes travelling on a network-on-chip. The problem of neuron-to-core mapping is relevant as a non-efficient allocation may impact real-time and reliability of the neural network execution. We designed a task placement pipeline capable of analysing the network of neurons and producing a placement configuration that enables a reduction of communication between computational nodes. The neuron-to-core mapping problem has been formalised as a problem of minimisation of synaptic elongation. Intuitively, this metric represents the cumulative distance that spikes generated by neurons running on a specific core have to travel to reach their destination core. The proposed placement methodology allows using different techniques to solve the problem. In this work Spectral Analysis, Multilevel Static Mapping, and Simulated Annealing were compared evaluating the overall post-placement synaptic elongation. Results point out that mapping solutions taking into account the directionality of the SNN provide a better placement and quantify this impact. Between all techniques considered only the Simulated Annealing was able to overcome an improvement of 25% compared to a random placement.
Francesco Barchi, Gianvito Urgese, Andrea Acquaviva, Enrico Macii
VLSI-SoC1
2016 Toolchain integration of runtime variability and aging awareness in multicore platforms
abstract
The impact of process variations and wear-out mechanisms in current and next generation technology nodes is becoming relevant and cannot be compensated at the device or architectural level. Intra-die process variations raising at the core level and platform level makes parallel multicore platforms intrinsically heterogeneous, because the various cores are clocked at different operational frequencies. Power consumption becomes heterogeneous too, both considering dynamic and leakage consumption. Wear-out processes add further uncertainty over time. In this context, to fully exploit the computational capability of the platform parallelism, variability and wear-out aware task allocation strategies must be developed. In this work, we discuss techniques to perform task allocation and we show how they can be integrated in a software toolchain. We report results of the implementation of variability and wear-out awareness in state-of-art multicore platforms.
Ramakrishna Nittala, Francesco Barchi, Gianvito Urgese, Andrea Acquaviva
FDL2