Manuele Rusci

dblp:188/1110 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-7458-4019ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A System-Level Performance Analysis of On-Device Learning on an Ultra-Low-Power Edge System
abstract
Training Deep Neural Networks (DNNs) on ultra-low-power processing systems is often hindered by the memory-intensive nature of backpropagation. This paper examines the trade-offs between memory and computational costs when mapping DNN training onto an ultra-low-power RISC-V-based System-on-Chip (SoC) with an 8-core accelerator, 1.5 MB on-chip memory, and a 32 MB off-chip memory. To this end, we develop a framework that automates the deployment of complete training graphs across heterogeneous SoCs. The generated code interleaves memory transfer operations with layer-wise processing functions, which are executed either on the host CPU or, for convolutional layers, on the multi-core accelerator. Overall, our work is the first to provide a full system-level analysis of end-to-end DNN training on OS-less, memory-constrained (32 MB of off-chip memory) edge systems, and we show that computation is the primary bottleneck by accounting for up to 91% of the total energy consumption.
Manuele Rusci, Mohamed Amine Hamdi, Daniele Jahier Pagliari, Francesco Conti 0001, Alessio Burrello
CF1
2026 Efficient On-Device Domain Learning for Keyword Spotting on Ultra-Low-Power Platforms
abstract
Deep Neural Network-based Keyword Spotting accuracy degrades in noisy environments. On-site adaptation to previously unseen noise is crucial to recover accuracy loss, and on-device learning is required in scenarios where adaptation has to happen in the field. In this work, we propose a fully on-device domain adaptation system, enabling edge devices to achieve noise-robust keyword spotting. We achieve up to 14% accuracy gains over already-robust keyword spotting models, and up to 21% increments when evaluating our methodology on keyword datasets disjoint from the offline training set. In extreme edge scenarios where Keyword Spotting is critical, using as little as 10 kB of memory and only 100 labeled utterances, we enable on-device learning and demonstrate accuracy recovery of up to 5% after adapting to complex, non-stationary speech noise. We show that domain adaptation can be achieved on ultra-low-power microcontrollers with as little as 357 mJ within 14 seconds on always-on, battery-operated devices. This work is the first to demonstrate an end-to-end on-device domain adaptation system for noise robust keyword spotting models on ultra-low-power, extreme edge platforms.
Cristian Cioflan, Lukas Cavigelli, Manuele Rusci, Miguel de Prado, Luca Benini
IEEE Internet Things J.3
2026 Multimodal On-Device Learning for Monocular Depth Estimation on Ultralow-Power MCUs
abstract
Monocular depth estimation (MDE) plays a crucial role in enabling spatially-aware applications in Ultra-low-power (ULP) Internet-of-Things (IoT) platforms. However, the limited number of parameters of Deep Neural Networks for the MDE task, designed for IoT nodes, results in severe accuracy drops when the sensor data observed in the field shifts significantly from the training dataset. To address this domain shift problem, we present a multi-modal On-Device Learning (ODL) technique, deployed on an IoT device integrating a Greenwaves GAP9 MicroController Unit (MCU), a 80 mW monocular camera and a 8 × 8 pixel depth sensor, consuming ∼300mW. In its normal operation, this setup feeds a tiny 107 k-parameter μPyD-Net model with monocular images for inference. The depth sensor, usually deactivated to minimize energy consumption, is only activated alongside the camera to collect pseudo-labels when the system is placed in a new environment. Then, the fine-tuning task is performed entirely on the MCU, using the new data. To optimize our backpropagation-based on-device training, we introduce a novel memory-driven sparse update scheme, which minimizes the fine-tuning memory to 1.2 MB, 2.2× less than a full update, while preserving accuracy (i.e., only 2% and 1.5% drops on the KITTI and NYUv2 datasets). Our in-field tests demonstrate, for the first time, that ODL for MDE can be performed in 17.8 minutes on the IoT node, reducing the root mean squared error from 4.9 to 0.6m with only 3 k self-labeled samples, collected in a real-life deployment scenario.
Davide Nadalini, Manuele Rusci, Elia Cereda, Luca Benini, Francesco Conti 0001, Daniele Palossi
IEEE Internet Things J.2
2025 Self-Incremental Training for Personalized Voice Command Recognition in a Wireless Audio Sensor Network
abstract
This paper studies self-incremental training in the context of personalized Deep Neural Networks (DNNs) for voice command recognition tailored for resource-constrained sensor nodes. The learning task runs when new unsupervised data becomes available within a Wireless Audio Sensor Network (WASN). After collecting a new multi-sensor dataset of voice commands, we experimentally investigate network-level policies to assign pseudo-labels to the new data. Our baseline analysis shows an accuracy improvement of up to +15% with respect to models pretrained on a large keyword corpus dataset. The multi-sensor labeling strategy closely approximates the performance achieved in a single-sensor scenario providing a clean signal, while we observe +4.7% compared to other sensors with degraded signal quality.
Manuele Rusci, Hugo Van hamme, Tinne Tuytelaars
ICASSP1
2025 Self-Learning for Personalized Keyword Spotting on Ultralow-Power Audio Sensors
abstract
This article proposes a self-learning method to incrementally train (fine-tune) a personalized keyword spotting (KWS) model after the deployment on ultralow power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5 M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% versus the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training$10\times $lower than the labeling energy if sampling a new utterance every 6.1 or 18.8 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.
Manuele Rusci, Francesco Paci, Marco Fariselli, Eric Flamand, Tinne Tuytelaars
IEEE Internet Things J.1
2024 On-device Self-supervised Learning of Visual Perception Tasks aboard Hardware-limited Nano-quadrotors
abstract
Sub-50g nano-drones are gaining momentum in both academia and industry. Their most compelling applications rely on onboard deep learning models for perception despite severe hardware constraints (i.e., sub-100mW processor). When deployed in unknown environments not represented in the training data, these models often underperform due to domain shift. To cope with this fundamental problem, we propose, for the first time, on-device learning aboard nano-drones, where the first part of the in-field mission is dedicated to self-supervised finetuning of a pre-trained convolutional neural network (CNN). Leveraging a real-world vision-based regression task, we thoroughly explore performance-cost trade-offs of the fine-tuning phase along three axes: i) dataset size (more data increases the regression performance but requires more memory and longer computation); ii) methodologies (e.g., fine-tuning all model parameters vs. only a subset); and iii) self-supervision strategy. Our approach demonstrates an improvement in mean absolute error up to 30% compared to the pre-trained baseline, requiring only 22s fine-tuning on an ultra-low-power GWT GAP9 System-on-Chip. Addressing the domain shift problem via on-device learning aboard nano-drones not only marks a novel result for hardware-limited robots but lays the ground for more general advancements for the entire robotics community.
Elia Cereda, Manuele Rusci, Alessandro Giusti, Daniele Palossi
ICRA2
2023 Bio-inspired Autonomous Exploration Policies with CNN-based Object Detection on Nano-drones
abstract
Nano-sized drones, with palm-sized form factor, are gaining relevance in the Internet-of-Things ecosystem. Achieving a high degree of autonomy for complex multi-objective missions (e.g., safe flight, exploration, object detection) is extremely challenging for the onboard chip-set due to tight size, payload (2unknown room in a 3 minutes flight. By combining the detection CNN and the exploration policy, we show an average detection rate of 90 % on six target objects in a never-seen-before environment.
Lorenzo Lamberti, Luca Bompani, Victor Kartsch, Manuele Rusci, Daniele Palossi, Luca Benini
DATE4
2023 Few-Shot Open-Set Learning for On-Device Customization of KeyWord Spotting Systems
abstract
sponsorship: This work is partly supported by the European Horizon Europe program under grant agreement 101067475. (European Horizon Europe program|101067475, Marie Curie Actions (MSCA)|101067475)
Manuele Rusci, Tinne Tuytelaars
INTERSPEECH1
2023 Reduced precision floating-point optimization for Deep Neural Network On-Device Learning on microcontrollers
Davide Nadalini, Manuele Rusci, Luca Benini, Francesco Conti 0001
Future Gener. Comput. Syst.2
2021 Low-Power License Plate Detection and Recognition on a RISC-V Multi-Core MCU-Based Vision System
abstract
In this paper, we present the first (to the best of our knowledge) demonstration of a low-power MCU-based edge device for Automatic License Plate Recognition (ALPR). The design leverages on a 9-core RISC-V processor, GAP8, coupled with a QVGA ultra-low-power greyscale imager. The proposed visual processing pipeline uses a multi-model inference approach based on SSDlite-MobilenetV2 for license plate detection and LPRNet for optical character recognition, reaching a 38.9% mAP score for the first task and a recognition rate of >99.13% for the latter on public datasets. On real-world data, the pipeline recognizes registration numbers when the size of LP crops is as small as 30×5 pixels. Thanks to the applied compression and optimization strategies, the multi-model inference (687 MMAC) achieves a throughput of 1.09 FPS at a power cost of 117 mW when running on GAP8. Our solution is the first MCU-class device embedding such a level of network complexity, resulting to be 73× more energy-efficient w.r.t. precedent mobile-class ALPR system featuring a Raspberry Pi3. The proposed design does not resort to any hardwired acceleration engines, thus retaining full flexibility for future algorithmic improvements.
Lorenzo Lamberti, Manuele Rusci, Marco Fariselli, Francesco Paci, Luca Benini
ISCAS2
2021 Robustifying the Deployment of tinyML Models for Autonomous Mini-Vehicles
abstract
Standard-size autonomous navigation vehicles have rapidly improved thanks to the breakthroughs of deep learning. However, scaling autonomous driving to low-power systems deployed on dynamic environments poses several challenges that prevent their adoption. To address them, we propose a closed- loop learning flow for autonomous driving mini-vehicles that includes the target environment in-the-loop. We leverage a family of compact and high-throughput tinyCNNs to control the mini- vehicle, which learn in the target environment by imitating a computer vision algorithm, i.e., the expert. Thus, the tinyCNNs, having only access to an on-board fast-rate linear camera, gain robustness to lighting conditions and improve over time. Further, we leverage GAP8, a parallel ultra-low-power RISC-V SoC, to meet the inference requirements. When running the family of CNNs, our GAP8's solution outperforms any other implementation on the STM32L4 and NXP k64f (Cortex-M4), reducing the latency by over 13x and the energy consummation by 92%.
Miguel de Prado, Manuele Rusci, Romain Donze, Alessandro Capotondi, Serge Monnerat, Luca Benini, Nuria Pazos
ISCAS2
2018 Always-ON visual node with a hardware-software event-based binarized neural network inference engine
abstract
This work introduces an ultra-low-power visual sensor node coupling event-based binary acquisition with Binarized Neural Networks (BNNs) to deal with the stringent power requirements of always-on vision systems for IoT applications. By exploiting in-sensor mixed-signal processing, an ultra-low-power imager generates a sparse visual signal of binary spatial-gradient features. The sensor output, packed as a stream of events corresponding to the asserted gradient binary values, is transferred to a 4-core processor when the amount of data detected after frame difference surpasses a given threshold. Then, a BNN trained with binary gradients as input runs on the parallel processor if a meaningful activity is detected in a pre-processing stage. During the BNN computation, the proposed Event-based Binarized Neural Network model achieves a system energy saving of 17.8% with respect to a baseline system including a low-power RGB imager and a Binarized Neural Network, while paying a classification performance drop of only 3% for a real-life 3-classes classification scenario. The energy reduction increases up to 8x when considering a long-term always-on monitoring scenario, thanks to the event-driven behavior of the processing sub-system.
Manuele Rusci, Davide Rossi 0001, Eric Flamand, Massimo Gottardi, Elisabetta Farella, Luca Benini
CF1
2018 Design Automation for Binarized Neural Networks: A Quantum Leap Opportunity?
abstract
Design automation in general, and in particular logic synthesis, can play a key role in enabling the design of application-specific Binarized Neural Networks (BNN). This paper presents the hardware design and synthesis of a purely combinational BNN for ultra-low power near-sensor processing. We leverage the major opportunities raised by BNN models, which consist mostly of logical bit-wise operations and integer counting and comparisons, for pushing ultra-low power deep learning circuits close to the sensor and coupling them with binarized mixed-signal image sensor data. We analyze area, power and energy metrics of BNNs synthesized as combinational networks. Our synthesis results in GlobalFoundries 22 nm SOI technology shows a silicon area of 2.61 mm2for implementing a combinational BNN with 32×32 binary input sensor receptive field and weight parameters fixed at design time. This is 2.2× smaller than a synthesized network with re-configurable parameters. With respect to other comparable techniques for deep learning near-sensor processing, our approach features a 10× higher energy efficiency.
Manuele Rusci, Lukas Cavigelli, Luca Benini
ISCAS1
2017 A Sub-mW IoT-Endnode for Always-On Visual Monitoring and Smart Triggering
abstract
This paper presents a fully programmable Internet of Things visual sensing node that targets sub-mW power consumption in always-on monitoring scenarios. The system features a spatial-contrast 128 × 64 binary pixel imager with focal-plane processing. The sensor, when working at its lowest power mode (10 μW at 10 frames/s), provides as output the number of changed pixels. Based on this information, a dedicated camera interface, implemented on a low-power field-programmable gate array, wakes up an ultralow-power parallel processing unit to extract context-aware visual information. We evaluate the smart sensor on three always-on visual triggering application scenarios. Triggering accuracy comparable to RGB image sensors is achieved at nominal lighting conditions, while consuming an average power between 193 and 277 μW, depending on context activity. The digital subsystem is extremely flexible, thanks to a fully programmable digital signal processing engine, but still achieves 19× lower power consumption compared to MCU-based cameras with significantly lower on-board computing capabilities.
Manuele Rusci, Davide Rossi 0001, Elisabetta Farella, Luca Benini
IEEE Internet Things J.1