EDBT 2026 Demo / reviewers in the wild / expert
Manuele Rusci
dblp:188/1110
· DBLP profile ↗
14ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-7458-4019ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A System-Level Performance Analysis of On-Device Learning on an Ultra-Low-Power Edge SystemabstractTraining Deep Neural Networks (DNNs) on ultra-low-power processing systems is often hindered by the memory-intensive nature of backpropagation. This paper examines the trade-offs between memory and computational costs when mapping DNN training onto an ultra-low-power RISC-V-based System-on-Chip (SoC) with an 8-core accelerator, 1.5 MB on-chip memory, and a 32 MB off-chip memory. To this end, we develop a framework that automates the deployment of complete training graphs across heterogeneous SoCs. The generated code interleaves memory transfer operations with layer-wise processing functions, which are executed either on the host CPU or, for convolutional layers, on the multi-core accelerator. Overall, our work is the first to provide a full system-level analysis of end-to-end DNN training on OS-less, memory-constrained (32 MB of off-chip memory) edge systems, and we show that computation is the primary bottleneck by accounting for up to 91% of the total energy consumption. Manuele Rusci, Mohamed Amine Hamdi, Daniele Jahier Pagliari, Francesco Conti 0001, Alessio Burrello |
CF | 1 |
| 2026 | Efficient On-Device Domain Learning for Keyword Spotting on Ultra-Low-Power PlatformsabstractDeep Neural Network-based Keyword Spotting accuracy degrades in noisy environments. On-site adaptation to previously unseen noise is crucial to recover accuracy loss, and on-device learning is required in scenarios where adaptation has to happen in the field. In this work, we propose a fully on-device domain adaptation system, enabling edge devices to achieve noise-robust keyword spotting. We achieve up to 14% accuracy gains over already-robust keyword spotting models, and up to 21% increments when evaluating our methodology on keyword datasets disjoint from the offline training set. In extreme edge scenarios where Keyword Spotting is critical, using as little as 10 kB of memory and only 100 labeled utterances, we enable on-device learning and demonstrate accuracy recovery of up to 5% after adapting to complex, non-stationary speech noise. We show that domain adaptation can be achieved on ultra-low-power microcontrollers with as little as 357 mJ within 14 seconds on always-on, battery-operated devices. This work is the first to demonstrate an end-to-end on-device domain adaptation system for noise robust keyword spotting models on ultra-low-power, extreme edge platforms. Cristian Cioflan, Lukas Cavigelli, Manuele Rusci, Miguel de Prado, Luca Benini |
IEEE Internet Things J. | 3 |
| 2026 | Multimodal On-Device Learning for Monocular Depth Estimation on Ultralow-Power MCUsabstractMonocular depth estimation (MDE) plays a crucial role in enabling spatially-aware applications in Ultra-low-power (ULP) Internet-of-Things (IoT) platforms. However, the limited number of parameters of Deep Neural Networks for the MDE task, designed for IoT nodes, results in severe accuracy drops when the sensor data observed in the field shifts significantly from the training dataset. To address this domain shift problem, we present a multi-modal On-Device Learning (ODL) technique, deployed on an IoT device integrating a Greenwaves GAP9 MicroController Unit (MCU), a 80 mW monocular camera and a 8 × 8 pixel depth sensor, consuming ∼300mW. In its normal operation, this setup feeds a tiny 107 k-parameter μPyD-Net model with monocular images for inference. The depth sensor, usually deactivated to minimize energy consumption, is only activated alongside the camera to collect pseudo-labels when the system is placed in a new environment. Then, the fine-tuning task is performed entirely on the MCU, using the new data. To optimize our backpropagation-based on-device training, we introduce a novel memory-driven sparse update scheme, which minimizes the fine-tuning memory to 1.2 MB, 2.2× less than a full update, while preserving accuracy (i.e., only 2% and 1.5% drops on the KITTI and NYUv2 datasets). Our in-field tests demonstrate, for the first time, that ODL for MDE can be performed in 17.8 minutes on the IoT node, reducing the root mean squared error from 4.9 to 0.6m with only 3 k self-labeled samples, collected in a real-life deployment scenario. Davide Nadalini, Manuele Rusci, Elia Cereda, Luca Benini, Francesco Conti 0001, Daniele Palossi |
IEEE Internet Things J. | 2 |
| 2025 | Self-Incremental Training for Personalized Voice Command Recognition in a Wireless Audio Sensor NetworkabstractThis paper studies self-incremental training in the context of personalized Deep Neural Networks (DNNs) for voice command recognition tailored for resource-constrained sensor nodes. The learning task runs when new unsupervised data becomes available within a Wireless Audio Sensor Network (WASN). After collecting a new multi-sensor dataset of voice commands, we experimentally investigate network-level policies to assign pseudo-labels to the new data. Our baseline analysis shows an accuracy improvement of up to +15% with respect to models pretrained on a large keyword corpus dataset. The multi-sensor labeling strategy closely approximates the performance achieved in a single-sensor scenario providing a clean signal, while we observe +4.7% compared to other sensors with degraded signal quality. Manuele Rusci, Hugo Van hamme, Tinne Tuytelaars |
ICASSP | 1 |
| 2025 | Self-Learning for Personalized Keyword Spotting on Ultralow-Power Audio SensorsabstractThis article proposes a self-learning method to incrementally train (fine-tune) a personalized keyword spotting (KWS) model after the deployment on ultralow power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5 M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% versus the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training$10\times $lower than the labeling energy if sampling a new utterance every 6.1 or 18.8 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge. Manuele Rusci, Francesco Paci, Marco Fariselli, Eric Flamand, Tinne Tuytelaars |
IEEE Internet Things J. | 1 |
| 2024 | On-device Self-supervised Learning of Visual Perception Tasks aboard Hardware-limited Nano-quadrotorsabstractSub-50g nano-drones are gaining momentum in both academia and industry. Their most compelling applications rely on onboard deep learning models for perception despite severe hardware constraints (i.e., sub-100mW processor). When deployed in unknown environments not represented in the training data, these models often underperform due to domain shift. To cope with this fundamental problem, we propose, for the first time, on-device learning aboard nano-drones, where the first part of the in-field mission is dedicated to self-supervised finetuning of a pre-trained convolutional neural network (CNN). Leveraging a real-world vision-based regression task, we thoroughly explore performance-cost trade-offs of the fine-tuning phase along three axes: i) dataset size (more data increases the regression performance but requires more memory and longer computation); ii) methodologies (e.g., fine-tuning all model parameters vs. only a subset); and iii) self-supervision strategy. Our approach demonstrates an improvement in mean absolute error up to 30% compared to the pre-trained baseline, requiring only 22s fine-tuning on an ultra-low-power GWT GAP9 System-on-Chip. Addressing the domain shift problem via on-device learning aboard nano-drones not only marks a novel result for hardware-limited robots but lays the ground for more general advancements for the entire robotics community. Elia Cereda, Manuele Rusci, Alessandro Giusti, Daniele Palossi |
ICRA | 2 |
| 2023 | Bio-inspired Autonomous Exploration Policies with CNN-based Object Detection on Nano-dronesabstractNano-sized drones, with palm-sized form factor, are gaining relevance in the Internet-of-Things ecosystem. Achieving a high degree of autonomy for complex multi-objective missions (e.g., safe flight, exploration, object detection) is extremely challenging for the onboard chip-set due to tight size, payload (2unknown room in a 3 minutes flight. By combining the detection CNN and the exploration policy, we show an average detection rate of 90 % on six target objects in a never-seen-before environment. Lorenzo Lamberti, Luca Bompani, Victor Kartsch, Manuele Rusci, Daniele Palossi, Luca Benini |
DATE | 4 |
| 2023 | Few-Shot Open-Set Learning for On-Device Customization of KeyWord Spotting Systemsabstractsponsorship: This work is partly supported by the European Horizon Europe program under grant agreement 101067475. (European Horizon Europe program|101067475, Marie Curie Actions (MSCA)|101067475) Manuele Rusci, Tinne Tuytelaars |
INTERSPEECH | 1 |
| 2023 | Reduced precision floating-point optimization for Deep Neural Network On-Device Learning on microcontrollers
Davide Nadalini, Manuele Rusci, Luca Benini, Francesco Conti 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | Low-Power License Plate Detection and Recognition on a RISC-V Multi-Core MCU-Based Vision SystemabstractIn this paper, we present the first (to the best of our knowledge) demonstration of a low-power MCU-based edge device for Automatic License Plate Recognition (ALPR). The design leverages on a 9-core RISC-V processor, GAP8, coupled with a QVGA ultra-low-power greyscale imager. The proposed visual processing pipeline uses a multi-model inference approach based on SSDlite-MobilenetV2 for license plate detection and LPRNet for optical character recognition, reaching a 38.9% mAP score for the first task and a recognition rate of >99.13% for the latter on public datasets. On real-world data, the pipeline recognizes registration numbers when the size of LP crops is as small as 30×5 pixels. Thanks to the applied compression and optimization strategies, the multi-model inference (687 MMAC) achieves a throughput of 1.09 FPS at a power cost of 117 mW when running on GAP8. Our solution is the first MCU-class device embedding such a level of network complexity, resulting to be 73× more energy-efficient w.r.t. precedent mobile-class ALPR system featuring a Raspberry Pi3. The proposed design does not resort to any hardwired acceleration engines, thus retaining full flexibility for future algorithmic improvements. Lorenzo Lamberti, Manuele Rusci, Marco Fariselli, Francesco Paci, Luca Benini |
ISCAS | 2 |
| 2021 | Robustifying the Deployment of tinyML Models for Autonomous Mini-VehiclesabstractStandard-size autonomous navigation vehicles have rapidly improved thanks to the breakthroughs of deep learning. However, scaling autonomous driving to low-power systems deployed on dynamic environments poses several challenges that prevent their adoption. To address them, we propose a closed- loop learning flow for autonomous driving mini-vehicles that includes the target environment in-the-loop. We leverage a family of compact and high-throughput tinyCNNs to control the mini- vehicle, which learn in the target environment by imitating a computer vision algorithm, i.e., the expert. Thus, the tinyCNNs, having only access to an on-board fast-rate linear camera, gain robustness to lighting conditions and improve over time. Further, we leverage GAP8, a parallel ultra-low-power RISC-V SoC, to meet the inference requirements. When running the family of CNNs, our GAP8's solution outperforms any other implementation on the STM32L4 and NXP k64f (Cortex-M4), reducing the latency by over 13x and the energy consummation by 92%. Miguel de Prado, Manuele Rusci, Romain Donze, Alessandro Capotondi, Serge Monnerat, Luca Benini, Nuria Pazos |
ISCAS | 2 |
| 2018 | Always-ON visual node with a hardware-software event-based binarized neural network inference engineabstractThis work introduces an ultra-low-power visual sensor node coupling event-based binary acquisition with Binarized Neural Networks (BNNs) to deal with the stringent power requirements of always-on vision systems for IoT applications. By exploiting in-sensor mixed-signal processing, an ultra-low-power imager generates a sparse visual signal of binary spatial-gradient features. The sensor output, packed as a stream of events corresponding to the asserted gradient binary values, is transferred to a 4-core processor when the amount of data detected after frame difference surpasses a given threshold. Then, a BNN trained with binary gradients as input runs on the parallel processor if a meaningful activity is detected in a pre-processing stage. During the BNN computation, the proposed Event-based Binarized Neural Network model achieves a system energy saving of 17.8% with respect to a baseline system including a low-power RGB imager and a Binarized Neural Network, while paying a classification performance drop of only 3% for a real-life 3-classes classification scenario. The energy reduction increases up to 8x when considering a long-term always-on monitoring scenario, thanks to the event-driven behavior of the processing sub-system. Manuele Rusci, Davide Rossi 0001, Eric Flamand, Massimo Gottardi, Elisabetta Farella, Luca Benini |
CF | 1 |
| 2018 | Design Automation for Binarized Neural Networks: A Quantum Leap Opportunity?abstractDesign automation in general, and in particular logic synthesis, can play a key role in enabling the design of application-specific Binarized Neural Networks (BNN). This paper presents the hardware design and synthesis of a purely combinational BNN for ultra-low power near-sensor processing. We leverage the major opportunities raised by BNN models, which consist mostly of logical bit-wise operations and integer counting and comparisons, for pushing ultra-low power deep learning circuits close to the sensor and coupling them with binarized mixed-signal image sensor data. We analyze area, power and energy metrics of BNNs synthesized as combinational networks. Our synthesis results in GlobalFoundries 22 nm SOI technology shows a silicon area of 2.61 mm2for implementing a combinational BNN with 32×32 binary input sensor receptive field and weight parameters fixed at design time. This is 2.2× smaller than a synthesized network with re-configurable parameters. With respect to other comparable techniques for deep learning near-sensor processing, our approach features a 10× higher energy efficiency. Manuele Rusci, Lukas Cavigelli, Luca Benini |
ISCAS | 1 |
| 2017 | A Sub-mW IoT-Endnode for Always-On Visual Monitoring and Smart TriggeringabstractThis paper presents a fully programmable Internet of Things visual sensing node that targets sub-mW power consumption in always-on monitoring scenarios. The system features a spatial-contrast 128 × 64 binary pixel imager with focal-plane processing. The sensor, when working at its lowest power mode (10 μW at 10 frames/s), provides as output the number of changed pixels. Based on this information, a dedicated camera interface, implemented on a low-power field-programmable gate array, wakes up an ultralow-power parallel processing unit to extract context-aware visual information. We evaluate the smart sensor on three always-on visual triggering application scenarios. Triggering accuracy comparable to RGB image sensors is achieved at nominal lighting conditions, while consuming an average power between 193 and 277 μW, depending on context activity. The digital subsystem is extremely flexible, thanks to a fully programmable digital signal processing engine, but still achieves 19× lower power consumption compared to MCU-based cameras with significantly lower on-board computing capabilities. Manuele Rusci, Davide Rossi 0001, Elisabetta Farella, Luca Benini |
IEEE Internet Things J. | 1 |