Valentino Peluso

dblp:191/6827 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0003-0527-8602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 8 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Runtime Feature Compression for Adaptive Keyword Spotting on Embedded Systems
abstract
Voice user interfaces rely on keyword spotting (KWS) to detect wake-word commands, enabling low-power devices to switch from drowsy to active states and initiate more complex tasks. In embedded systems, KWS combines handcrafted acoustic features extraction with lightweight neural network classifiers to achieve accurate detection within strict resource constraints. Adapting KWS to time-varying energy budgets requires optimization strategies that operate at runtime. Most existing approaches adjust the complexity of the neural model but overlook that a substantial amount of latency, and thus energy consumption, is due to feature extraction, which remains unaffected by model scaling. This work introducesRuntime Feature Compression(RFC), a dynamic rescaling strategy that modulates the workload of the entire KWS pipeline. RFC promotes thehop-lengthparameter of the Short-Time Fourier Transform as a runtime control knob to adjust the number of time frames in speech features, allowing a single model to operate across multiple latency modes. To support this flexibility, we introduce two training-time techniques:HopAugment, a data augmentation scheme that exposes the model to variable hop lengths during training, andMasked Layers, which preserve consistent activation statistics during training and inference under compressed feature settings. Evaluations on four KWS datasets using the TC-ResNet model family show that RFC outperforms model scaling techniques, offering a wider range of latency-accuracy trade-offs. RFC achieves up to 31.8% lower latency without accuracy degradation, or up to 0.30% higher accuracy within equivalent latency bounds. That proves RFC improves adaptability in energy-constrained IoT speech interfaces. A set of ablation studies further demonstrates the robustness of RFC by evaluating the role of its training components, batching strategies, ability to preserve accuracy with a shared weight set, scalability across operating modes, and applicability to different model architectures.
Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi
IEEE Internet Things J.1
2026 FedAGF: Adaptive Concurrency via Gradient Feedback for Mitigating Extreme Label Skew in Budget-Constrained Federated Learning
abstract
Cross-device Federated Learning (FL) enables large fleets of distributed edge devices to collaboratively train a global classification model without sharing their data. The resulting training quality is strongly limited byextreme label skew, a condition where each device holds samples from a subset of the target classes. In such cases, local updates become biased toward local distributions, slowing down convergence and degrading the global model accuracy. These effects get critical when devices operate under constrained energy budgets that restrict their participation to a limited number of synchronization rounds, further reducing the achievable accuracy. To overcome these limitations, we introduce Federated Learning with Adaptive Concurrency via Gradient Feedback (FedAGF), a control policy that dynamically adjusts the number of devices selected for synchronization throughout training. FedAGF adapts to training dynamics by monitoring global model updates and elevating participation whenever progress slows. This adaptive mechanism balances accuracy improvement with efficient use of devices’ energy budgets, allocating resources when they provide the greatest benefit to convergence. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that FedAGF effectively mitigates extreme label skew, achieving up to 14.14% higher accuracy than state-of-the-art FL methods and enabling efficient and scalable training, even under skewed label distributions and resource constraints.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Privacy-Preserving Federated Learning for Household Characteristic Identification
abstract
This work presents a privacy-preserving training framework for household characteristic identification from electricity consumption data. The proposed framework integrates two main components: (i) a synthetic data generation pipeline capable of replicating realistic energy traces from diverse family compositions, capturing fine-grained sociodemographic attributes such as household size, employment status, age groups, and home occupancy patterns; (ii) a training strategy based on Federated Learning (FL) secured with homomorphic encryption, enabling collaborative model training while preserving data ownership. Our synthetic dataset enables the performance assessment of different training scenarios, including siloed model training by individual energy utilities and secure collaboration via FL. Experimental results show that siloed training leads to inconsistent and suboptimal performance, while privacy-preserving FL achieves accuracy comparable to conventional centralized training—an ideal yet not viable option due to data regulation constraints. Our findings highlight the effectiveness of FL as a secure solution for collaborative sociodemographic profiling in smart grids.
Erich Malan, Claudia De Vizia, Marco Castangia, Valentino Peluso, Andrea Calimera, Enrico Macii
COMPSAC4
2025 Gradient-Aware Participation for Energy Reduction in Federated Learning with Extreme Label Skew
abstract
Federated Learning (FL) enables distributed clients to train a global classification model collaboratively while preserving data privacy. A major challenge in FL is ensuring efficient training with limited computing and communication resources, especially when clients’ datasets contain samples from a restricted subset of target classes, a problem known as extreme label skew. Under such a condition, model updates from clients are biased toward their local data distributions, resulting in slow convergence and increased energy consumption due to the need for additional training rounds. This paper introduces FL with Gradient-Aware Participation (FedGAP), a novel strategy aimed at reducing energy consumption while preserving model accuracy even with extreme label skew. FedGAP dynamically adjusts the cohort size, i.e., the number of participating clients per training round, based on the evolution of the global model’s pseudo-gradient. By detecting stagnant phases where progress toward convergence stalls, FedGAP increases the cohort size to escape suboptimal regions and accelerate learning, thereby minimizing the waste of resources. Experiments on CIFAR-10 and CIFAR-100 demonstrate that FedGAP achieves up to 2.74× greater energy efficiency compared to state-of-the-art methods without compromising accuracy.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
IJCNN2
2024 Private Tensor Freezing for an Efficient Federated Learning with Homomorphic Encryption
abstract
Federated Learning (FL) is a privacy-preserving machine learning strategy where distributed clients share updates of locally trained models with a central server. The server aggregates those updates to refine a global version of the model without accessing the clients' data. Even if the raw data never leaves clients, adversarial attacks on the server side can still extract sensitive information from the transmitted model updates. Homomorphic Encryption (HE) offers a robust solution to this privacy concern: clients send encrypted model updates to the server; the server operates the aggregation of the received updates without having to decrypt. Unfortunately, HE leads to a substantial computational and communication overhead on both the client and the server side, preventing the adoption in practical, real-life applications. In this work, we introduce a training option to mitigate this problem by promoting Private Tensor Freezing (PTF), a progressive and secure gating scheme by which the number of model tensors involved in the training and synchronization stages gradually reduces over time, alleviating (i) the pressure of HE encryption/decryption on the client side, (ii) the communication volumes from/to the server, and (iii) the computing complexity of the aggregation stage on the server side. Experiments on four image classification benchmarks trained within a state-of-the-art FL framework secured with CKKS encryption reveal the effectiveness of PTF: up to 37.4% in data volume reduction, 35.5% less compute time on the client side, and 36.4% less compute time on the server side.
Valentino Peluso, Erich Malan, Andrea Calimera, Enrico Macii
ICCD1
2024 Automatic Layer Freezing for Communication Efficiency in Cross-Device Federated Learning
abstract
Federated learning (FL) is a collaborative machine learning paradigm where network-edge clients train a global model under the orchestration of a central server. Unlike traditional distributed learning, each participating client keeps its data locally, ensuring privacy protection by default. However, state-of-the-art FL implementations suffer from massive information exchange between clients and the server. This issue prevents the adoption in constrained environments, typical of the Internet of Things domain, where the communication bandwidth and the energy budget are severely limited. To achieve higher efficiency at scale, the future of FL calls for additional optimizations to reach high-quality learning capability with lower communication pressure. To address this challenge, we propose automatic layer freezing (ALF), an embedded mechanism that gradually drops a growing portion of the model out of the training and synchronization phases of the learning loop, reducing the volume of exchanged data with the central server. ALF monitors the evolution of model updates and identifies layers that have reached a stable representation, where further weight updates would have minimal impact on accuracy. By freezing these layers, ALF achieves substantial savings in communication bandwidth and energy consumption. The proposed implementation of the ALF mechanism is compatible with any FL strategy, requiring minimal effort and without interfering with existing optimizations. The extensive experiments conducted using a representative set of FL strategies applied to two image classification tasks show that ALF improves the communication efficiency of the baseline FL implementations, ensuring up to 83.91% of data volume savings with no or marginal losses of accuracy.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi
IEEE Internet Things J.2
2023 Enabling DVFS Side-Channel Attacks for Neural Network Fingerprinting in Edge Inference Services
abstract
The Inference-as-a-Service (IaaS) delivery model provides users access to pre-trained deep neural networks while safeguarding network code and weights. However, IaaS is not immune to security threats, like side-channel attacks (SCAs), that exploit unintended information leakage from the physical characteristics of the target device. Exposure to such threats grows when IaaS is deployed on distributed computing nodes at the edge. This work identifies a potential vulnerability of low-power CPUs that facilitates stealing the deep neural network architecture without physical access to the hardware or interference with the execution flow. Our approach relies on a Dynamic Voltage and Frequency Scaling (DVFS) side-channel attack, which monitors the CPU frequency state during the inference stages. Specifically, we introduce a dedicated load-testing methodology that imprints distinguishable signatures of the network on the frequency traces. A machine learning classifier is then used to infer the victim architecture. Experimental results on two commercial ARM Cortex-A CPUs, the A72 and A57, demonstrate the attack can identify the target architecture from a pool of 12 convolutional neural networks with an average accuracy of 98.7% and 92.4%
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
ISLPED2
2022 Energy-Quality Scalable Monocular Depth Estimation on Low-Power CPUs
abstract
The recent advancements in deep learning have demonstrated that inferring high-quality depth maps from a single image has become feasible and accurate, thanks to convolutional neural networks (CNNs), but how to process such compute- and memory-intensive models on portable and low-power devices remains a concern. Dynamic energy-quality scaling is an interesting yet less explored option in this field. It can improve efficiency through opportunistic computing policies where performances are boosted only when needed, achieving on average substantial energy savings. Implementing such a computing paradigm encompasses the availability of a scalable inference model, which is the target of this work. Specifically, we describe and characterize the design of an energy-quality scalable pyramidal network (EQPyD-Net), a lightweight CNN capable of modulating at runtime the computational effort with minimal memory resources. We describe the architecture of the network and the optimization flow, covering the important aspects that enable the dynamic scaling, namely, the optimized training procedures, the compression stage via fixed-point quantization, and the code optimization for the deployment on commercial low-power CPUs adopted in the edge segment. To assess the effect of the proposed design knobs, we evaluated the prediction quality on the standard KITTI data set and the energy and memory resources on the ARM Cortex-A53 CPU. The collected results demonstrate the flexibility of the proposed network and its energy efficiency. EQPyD-Net can be shifted across five operating points, ranging from a maximum accuracy of 82.2% with 0.4 Frame/J and up to 92.6% of energy savings with 6.1% of accuracy loss, still keeping a compact memory footprint of 5.2 MB for the weights and 38.3 MB (in the worst case) for the processing.
Antonio Cipolletta, Valentino Peluso, Andrea Calimera, Matteo Poggi, Fabio Tosi, Filippo Aleotti, Stefano Mattoccia
IEEE Internet Things J.2
2022 Monocular Depth Perception on Microcontrollers for Edge Applications
abstract
Depth estimation is crucial in several computer vision applications, and a recent trend in this field aims at inferring such a cue from a single camera. Unfortunately, despite the compelling results achieved, state-of-the-art monocular depth estimation methods are computationally demanding, thus precluding their practical deployment in several application contexts characterized by low-power constraints. Therefore, in this paper, we propose a lightweight Convolutional Neural Network based on a shallow pyramidal architecture, referred to as$\mu $PyD-Net, enabling monocular depth estimation on microcontrollers. The network is trained in a peculiar self-supervised manner leveraging proxy labels obtained through a traditional stereo algorithm. Moreover, we propose optimization strategies aimed at performing computations with quantized 8-bit data and map the high-level description of the network to low-level layers optimized for the target microcontroller architecture. Exhaustive experimental results on standard datasets and an in-depth evaluation with a device belonging to the popular Arm Cortex-M family confirm that obtaining sufficiently accurate monocular depth estimation on microcontrollers is feasible. To the best of our knowledge, our proposal is the first one enabling such remarkable achievement, paving the way for the deployment of monocular depth cues onto the tiny end-nodes of distributed sensor networks.
Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Matteo Poggi, Fabio Tosi, Filippo Aleotti, Stefano Mattoccia
IEEE Trans. Circuits Syst. Video Technol.1
2021 AdapTTA: Adaptive Test-Time Augmentation for Reliable Embedded ConvNets
abstract
Convolutional Neural Networks (ConvNets) are trained offline using the few available data and may therefore suffer from substantial accuracy loss when ported on the field, where unseen input patterns received under unpredictable external conditions can mislead the model. Test-Time Augmentation (TTA) techniques aim to alleviate such common side effect at inference-time, first running multiple feed-forward passes on a set of altered versions of the same input sample, and then computing the main outcome through a consensus of the aggregated predictions. Unfortunately, the implementation of TTA on embedded CPUs introduces latency penalties that limit its adoption on edge applications. To tackle this issue, we propose AdapTTA, an adaptive implementation of TTA that controls the number of feed-forward passes dynamically, depending on the complexity of the input. Experimental results on state-of-the-art ConvNets for image classification deployed on a commercial ARM Cortex-A CPU demonstrate AdapTTA reaches remarkable latency savings, from $1.40 \times$ to $2.21 \times$, and hence a higher frame rate compared to static TTA, still preserving the same accuracy gain.
Luca Mocerino, Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Enrico Macii
VLSI-SoC3
2020 Optimization Tools for ConvNets on the Edge
abstract
The shift of Convolutional Neural Networks (ConvNets) into low-power devices with limited compute and memory resources calls for cross-layer strategies spanning from hardware to software optimization. This work answers to this need, presenting a collection of tools for efficient deployment of ConvNets on the edge,
Valentino Peluso, Enrico Macii, Andrea Calimera
VLSI-SOC1
2019 Enabling Energy-Efficient Unsupervised Monocular Depth Estimation on ARMv7-Based Platforms
abstract
This work deals with the implementation of energy-efficient monocular depth estimation using a low-cost CPU for low-power embedded systems. It first describes the PyD-Net depth estimation network, which consists of a lightweight CNN able to approach state-of-the-art accuracy with ultra-low resource usage. Then it proposes an accuracy-driven complexity reduction strategy based on a hardware-friendly fixed-point quantization. Finally, it introduces the low-level optimization enabling effective use of integer neural kernels. The objective is threefold: (i) prove the efficiency of the new quantization flow on a depth estimation network, that is, the capability to retaining the accuracy reached by floating-point arithmetic using 16- and 8-bit integers, (ii) demonstrate the portability of the quantized model into a general-purpose 32-bit RISC architecture of the ARM Cortex family, (iii) quantify the accuracy-energy tradeoff of unsupervised monocular estimation to establish its use in the embedded domain. The experiments have been run on a Raspberry PI board powered by a Broadcom BCM2837 chipset. A parametric analysis conducted over the KITTI date-set shows marginal accuracy loss with 16-bit (8-bit) integers and energy savings up to 6.55× (9.23×) w.r.t. floating-point. Compared to high-end CPU and GPU the proposed solution improves scalability.
Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Matteo Poggi, Fabio Tosi, Stefano Mattoccia
DATE1
2019 Arbitrary-Precision Convolutional Neural Networks on Low-Power IoT Processors
abstract
The deployment of Convolutional Neural Networks (CNNs) on resource-constrained IoT devices calls for accurate model re-sizing and optimization. Among the proposed compression strategies, n-ary fixed-point quantization has proven effective in reducing both computational effort and memory footprint with no (or limited) accuracy loss. However, its use requires custom components and special memory allocation strategies which are not available and burdensome to implement on low-power/low-cost cores. In order to bridge this gap, this work introduces Virtual Quantization (VQ), a hardware-friendly compression method which allows to implement equivalent n ary CNNs on general purpose instruction-set architectures. The proposed VQ framework is validated for the IoT family of ARM MCUs (ARM Cortex-M) and tested with three different real-life applications (i.e. Image Classification, Keyword Spotting, Facial Expression Recognition).
Valentino Peluso, Matteo Grimaldi, Andrea Calimera
VLSI-SoC1
2018 All-digital embedded meters for on-line power estimation
abstract
Modern low power designs use multiple knobs for concurrent dynamic and leakage power optimization; supply voltage and threshold voltage are the most adopted. An efficient control of these knobs needs management policies aware of the power breakdown. This implies the availability of smart on-chip strategies for dynamic and leakage power estimation at runtime. In this paper, we address this issue proposing the implementation of embedded dynamic/static power meters that use an optimized regression model fed with data collected from in-situ activity monitors. The number of sensors, their bitwidth and optimal placement are obtained through an automated design flow. The methodology works for general logic and applies not just to processor cores, but also to application-specific designs. We apply our solution to a representative class of benchmarks, showing that it can achieve an average estimation error smaller than 3%, with limited area and power overheads.
Daniele Jahier Pagliari, Valentino Peluso, Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE2
2018 Energy-performance design exploration of a low-power microprogrammed deep-learning accelerator
abstract
This paper presents the design space exploration of a novel microprogrammable accelerator in which PEs are connected with a Network-on-Chip and benefit from low-power features enabled through a practical implementation of a Dual-Vddassignment scheme. An analytical model, fitted with postlayout data obtained with a 28nm FDSOI design kit, returns implementations with optimal energy-performance tradeoff by taking into consideration all the key design-space variables. The obtained Pareto analysis helps us infer optimization rules aimed at improving quality of design.
Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto
DATE3
2018 Scalable-effort ConvNets for multilevel classification
abstract
This work introduces the concept of scalable-effort Convolutional Neural Networks (ConvNets), an effort-accuracy scalable model for classification of data at multilevel abstraction. Scalable-effort ConvNets are able to adapt at run-time to the complexity of the classification problem, i.e. the level of abstraction defined by the application (or context), and reach a given classification accuracy with minimal computational effort. The mechanism is implemented using a single-weight scalable-precision model rather than an ensemble of quantized weight models; this makes the proposed strategy highly flexible and particularly suited for embedded architectures with limited resource availability. The paper describes (i) a hardware/software vertical implementation of scalable-precision multiply&accumulate arithmetic, (ii) an accuracy-constrained heuristic that delivers near-optimal layer-by-layer precision mapping at a predefined level of abstraction. It also reports the validation for three state-of-the-art nets, i.e. AlexNet, SqueezeNet and MobileNet, trained and tested with ImageNet. Collected results show scalable-effort ConvNets guarantee flexibility and substantial savings: 47.07% computational effort reduction at minimum accuracy, or 30.6% accuracy improvement at maximum effort w.r.t. standard flat ConvNets (average over the three benchmarks for high-level classification).
Valentino Peluso, Andrea Calimera
ICCAD1
2018 Weak-MAC: Arithmetic Relaxation for Dynamic Energy-Accuracy Scaling in ConvNets
abstract
This work introduces Weak-MAC (WeMAC), a precision scaling strategy for fixed point Multiply&Accumulate operations. By leveraging an algorithmic relaxation of multiplication, it improves the efficiency of deep convolutional neural networks. WeMAC can be run on a software-programmable 8-bit MAC unit and is retraining free, namely, it makes ConvNet achieving reasonable inference accuracy even without retraining. Simulation results conducted on three ConvNets trained over well recognized data-set (SVHN, CIFAR-10 and CIFAR-100) demonstrate WeMAC enables new Pareto points in the energy-accuracy tradeoff of HW accelerators. With 25% more energy efficiency w.r.t. full precision, WeMAC represents a middle way between the low energy consumption of 8-bit arithmetic and the high accuracy of 16-bit arithmetic.
Valentino Peluso, Andrea Calimera
ISCAS1
2018 Design-Space Exploration of Pareto-Optimal Architectures for Deep Learning with DVFS
abstract
Specialized computing engines are required to accelerate the execution of Deep Learning (DL) algorithms in an energy-efficient way. To adapt the processing throughput of these accelerators to the workload requirements while saving power, Dynamic Voltage and Frequency Scaling (DVFS) seems the natural solution. However, DL workloads need to frequently access the off-chip memory, which tends to make the performance of these accelerators memory-bound rather than computation-bound, hence reducing the effectiveness of DVFS. In this work we use a performance-power analytical model fitted on a parametrized implementation of a DL accelerator in a 28-nm FDSOI technology to explore a large design space and to obtain the Pareto points that maximize the effectiveness of DVFS in the sub-space of throughput and energy efficiency. In our model we consider the impact on performance and power of the off-chip memory using real data of a commercial low-power DRAM.
Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto
ISCAS3
2018 Energy-Driven Precision Scaling for Fixed-Point ConvNets
abstract
Data precision scaling is a well-known technique for power/energy minimization in error-resilient applications. It has proven particularly suited for embedded Convolutional Neural Networks (ConvNets) made run on fixed-point arithmetic coprocessors. The key observation is that methods that only account for accuracy during the precision assignment process may lead to sub-optimal energy minimization. This work introduces an energy-driven optimization that delivers per-layer quantization under a user-defined accuracy constraint. The tool is conceived for accelerators that dynamically adapt their energy and accuracy through software-programmable multiprecision Multiply&Accumulate (MAC) units. Simulation results collected on different ConvNets trained with public data-set show substantial energy savings and improved energy-accuracy tradeoffs w.r.t. conventional fixed-point methods.
Valentino Peluso, Andrea Calimera
VLSI-SoC1
2017 Early bird sampling: A short-paths free error detection-correction strategy for data-driven VOS
abstract
Razor is a milestone in the field of Error Detection&Correction strategies for low-power operation. Despite the impressive level of maturity, its application on circuits other than pipelined processors still remains an open issue. Firstly, the error detection mechanism relies on special flip-flops (FFs), the Razor-FFs, whose use imposes heavy hold-time fixing and large circuit area/power overheads; secondly, the error correction is performed through instruction replay, a practice that is not available (or very expensive to implement) in generic circuits. This work introduces Early Bird Sampling (EBS), a Razor variant that applies to low-power sequential circuits. The EBS allows to (i) solve the problem of short-path races bypassing tedious holdtime fixing design stages, (ii) reduce design overhead exploiting a local logic-masking mechanism for error correction. As a key feature, EBS enables Data-Driven Voltage Over-Scaling (DD-VOS), an aggressive dynamic voltage scaling strategy particularly suited for ultra-low power error-resilient applications. Simulation runs on a representative set of circuits provide a fair comparison with a standard Razor strategy. The collected results show EBS reduces area overheads (3.6% against 71.6% for Razor) and improves the voltage scaling profile achieving lower energy-per-operation (savings w.r.t. Razor range from 19.1% to 53.1%).
Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Jun Zhou 0017, Xin Liu 0015
VLSI-SoC2
2016 Ultra-Fine Grain Vdd-Hopping for energy-efficient Multi-Processor SoCs
abstract
This paper introduces Ultra-Fine Grain Vdd-Hopping (FINE-VH), an extension of Dynamic Voltage-Frequency Scaling (DVFS) for energy efficient Multi-Processor SoCs (MPSoCs). The proposed technique leverages the working principle of Vdd-Hopping applied at ultra-fine granularity, i.e., within the core, by means of a layout-assisted, level-shifter free, dynamic dual-Vdd control strategy where leakage currents are minimized through an optimal timing-driven poly-bias assignment procedure.
Valentino Peluso, Andrea Calimera, Enrico Macii, Massimo Alioto
VLSI-SoC1