Barbara De Salvo

dblp:119/8340 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-0810-9903ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 H4H: Hybrid Convolution-Transformer Architecture Search for NPU-CIM Heterogeneous Systems for AR/VR Applications
abstract
Low-latency and low-power edge AI is crucial for Augmented/Virtual Reality applications. Recent advances demonstrate that hybrid models, combining convolution layers (CNN) and transformers (ViT), often achieve a superior accuracy/performance tradeoff on various computer vision and machine learning (ML) tasks. However, hybrid ML models can present system challenges for latency and energy efficiency due to their diverse nature in dataflow and memory access patterns. In this work, we leverage architecture heterogeneity from Neural Processing Units (NPU) and Compute-In-Memory (CIM) and explore diverse execution schemas for efficient hybrid model executions. We introduce H4H-NAS, a two-stage Neural Architecture Search (NAS) framework to automate the design of hybrid CNN/ViT models for heterogeneous edge systems featuring both NPU and CIM. We propose a two-phase incremental supernet training in our NAS to resolve gradient conflicts between sampled subnets caused by different block types in a hybrid model search space. Our H4H-NAS approach is also powered by a performance estimator built with NPU performance results measured on real silicon, and CIM performance based on industry IPs. H4H-NAS searches hybrid CNN-ViT models with fine granularity and achieves significant (up to 1.34%) top-1 accuracy improvement on ImageNet-1k. Moreover, results from our algorithm/hardware co-design reveal up to 56.08% overall latency and 41.72% energy improvements by introducing heterogeneous computing over baseline solutions. Overall, our framework guides the design of hybrid network architectures and system architectures for NPU+CIM heterogeneous systems.
Yiwei Zhao 0001, Sai Qian Zhang, Syed Shakib Sarwar, Kleber Stangherlin, Jorge Gomez 0001, Jae-sun Seo, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li 0001
ASP-DAC8
2025 DFT Gaze: Distilled and Fine-Tuned Gaze Estimation for Personalization on Tiny Devices
abstract
Real-time personalized gaze estimation on AR/VR devices requires both accuracy and efficiency, especially when adapting to individual users with limited personal data. This task is challenging due to low-latency requirements, the presence of dataset biases from dominant gaze directions, and risk of catastrophic forgetting during adaptation. We present Distilled and Fine-Tuned (DFT) Gaze, a lightweight model for personalized gaze estimation. Distilled from a larger teacher model, DFT Gaze reduces model size while retaining essential visual features through knowledge distillation, without relying on gaze-specific supervision. During fine-tuning, it integrates gaze-specific supervision with Adapters, reaching 281K parameters for efficient adaptation and online updates on edge devices. To mitigate dataset biases and reduce catastrophic forgetting, we introduce a clustering-based sampling that balances gaze distribution for better generalization and improves adaptation to individual gaze patterns, even with only 5 personal images. DFT Gaze outperforms state-of-the-art methods on the MPIIFaceGaze dataset for personalized gaze estimation. Despite having the smallest model size at 281K parameters, it maintains low gaze errors across other datasets, including MPIIGaze, OpenEDS2020, and AEA. At 10× smaller than its teacher model, DFT Gaze achieves fast inference, a low parameter count, and effective adaptation, making it well-suited for real-time applications in resource-constrained environments.
He-Yen Hsieh, Ziyun Li 0001, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung 0001
ICIP6
2025 Exploring MRAM for On-Chip Texture Storage in Rendering Applications
abstract
In recent years, Magnetoresistive Random-Access Memory (MRAM) has attracted considerable attention as a high-density, non-volatile alternative to conventional embedded memory technologies. While MRAM has been recently adopted for storing neural network weights, its application in rendering workloads remains unexplored. In this study, we investigate the potential of MRAM for on-chip texture storage within a tile-based rasterization workload. Leveraging Siracusa, a RISC-V-based System-on-Chip (SoC) that integrates both MRAM and SRAM at the same memory hierarchy level, we conduct a comparative evaluation focusing on latency and energy consumption across varying frame rates. The results suggest that MRAM achieves substantial energy savings at lower frame rates due to its ability to enter deep-sleep mode between rendering cycles. However, this benefit diminishes as frame rates increase, with SRAM becoming more energy-efficient beyond a threshold of 43 frames per second. These findings demonstrate that MRAM is particularly wellsuited to read-intensive, energy-constrained rendering tasks.
Nicolás Villegas, Stefano Romanini, Moritz Scherer 0001, Warren Hunt, Syed Shakib Sarwar, Barbara De Salvo, Chiao Liu, Francesco Conti 0001, Davide Rossi 0001, Luca Benini, Jorge Gomez 0001
VLSI-SoC6
2024 Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical Framework
abstract
Augmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration.
Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong
ACM Trans. Design Autom. Electr. Syst.12
2024 Thermally Constrained Codesign of Heterogeneous 3-D Integration of Compute-in-Memory, Digital ML Accelerator, and RISC-V Cores for Mixed ML and Non-ML Workloads
abstract
Heterogeneous 3-D (H3D) integration not only reduces the chip form factor and fabrication cost but also allows the merging of diverse compute paradigms that suit different applications. This is especially attractive when modern algorithms, such as the augmented reality/virtual reality (AR/VR) workloads, consist of mixed machine learning (ML) and non-ML workloads. To date, codesign that considers the thermal, latency, and power constraints of H3D hardware is largely unexplored. In this work, a thermally aware framework for H3D hardware design is developed to evaluate the thermal, latency, and power trade-offs for a heterogeneous system with compute-in-memory (CIM), digital ML cores, and RISC-V cores. The framework solves for runtime tunable operating points described as the optimal speedup factor, the number of activated RISC-V cores, the cooling coefficient, and the activity rate based on user-defined criteria, achieving up to 135 TOPS and 215 TOPS/W under$74~^{\circ }$C for the AR/VR workloads.
Yuan-Chun Luo, Anni Lu, Janak Sharda, Moritz Scherer 0001, Jorge Gomez 0001, Syed Shakib Sarwar, Ziyun Li 0001, Reid Frederick Pinkham, Barbara De Salvo, Shimeng Yu
IEEE Trans. Very Large Scale Integr. Syst.9
2023 ANSA: Adaptive Near-Sensor Architecture for Dynamic DNN Processing in Compact Form Factors
abstract
Advanced edge sensing/computing devices, such as AR/VR devices, have a uniquely challenging adaptive baseline workload and camera sensor structure. These devices must process images in real-time from multiple sensors, placing a large burden on a typical centralized mobile SoC processor. Augmenting the sensors with a package-integrated near-sensor processor can improve the device’s processing performance as well as reduce energy consumption. This near-sensor processor must adapt to the dynamic workloads, fit within a limited silicon footprint and energy envelope, and satisfy the real-time requirement. In this work, we present ANSA, a near-sensor processor architecture supporting flexible processing schemes and dataflows to maintain high efficiency for dynamic CNN workloads. ANSA is scalable to sub-mm2 sizes to match the footprint of advanced image sensors. ANSA supports module-level power gating to adapt the compute capacity to dynamic workloads. Finally, ANSA leverages recent advancements in high-density non-volatile memory and 3D packaging to support weight storage within the area constraints of an image sensor. Overall, ANSA achieves inference energy consumption up to$30\times $lower than a standard SIMD baseline. Additionally, our design’s scalability allows it to achieve up to$2.76\times $lower average inference energy at$4.5\times $lower silicon area compared to competing edge accelerator designs.
Reid Pinkham, Jack Erhardt, Barbara De Salvo, Andrew Berkovich, Zhengya Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 SplitNets: Designing Neural Architectures for Efficient Distributed Computing on Head-Mounted Systems
abstract
We design deep neural networks (DNNs) and corresponding networks' splittings to distribute DNNs' workload to camera sensors and a centralized aggregator on head mounted devices to meet system performance targets in inference accuracy and latency under the given hardware resource constraints. To achieve an optimal balance among computation, communication, and performance, a split-aware neural architecture search framework, SplitNets, is introduced to conduct model designing, splitting, and communication reduction simultaneously. We further extend the framework to multi-view systems for learning to fuse inputs from multiple camera sensors with optimal performance and systemic efficiency. We validate SplitNets for single-view system on ImageNet as well as multi-view system on 3D classification, and show that the SplitNets framework achieves state-of-the-art (SOTA) performance and system latency compared with existing approaches.
Xin Dong 0009, Barbara De Salvo, Meng Li 0004, Chiao Liu, Zhongnan Qu, H. T. Kung 0001, Ziyun Li 0001
CVPR2
2019 Hybrid CMOS-RRAM Neurons with Intrinsic Plasticity
abstract
Brain-inspired architectures in neuromorphic hardware are currently subject to intensive research as an alternative to the limits of traditional computer organisation. The remarkable computing performance and efficiency of biological nervous systems are widely attributed to the co-localisation of memory and computation spatially throughout the structure. Moreover, it appears that a number of local self-organising neural mechanisms play their part in efficient biological computation. An example is neuronal intrinsic plasticity, where a neuron adapts its parameters to maximise its information capacity based on the statistical properties of its input while minimising the power it consumes. CMOS circuits implementing neuron models have been proposed but require their parameters to be set by biases originating from a centralised memory. In this work, we propose a hybrid CMOS-RRAM circuit that addresses this problem through storing neuron parameters within programmable nonvolatile resistive memories incorporated into the CMOS neuron. Additional circuits exploit the stochastic switching properties of resisitive memories to map a local intrinsic plasticity algorithm onto the proposed neuron. We demonstrate the computational advantages of this algorithm through simulation, calibrated on experimental data, whereby the neuron maximises its information capacity while minimising its power consumption, as is the case for biological neurons.
Thomas Dalgaty, Melika Payvand, Barbara De Salvo, Jerome Casas, Giusy Lama, Etienne Nowak, Giacomo Indiveri, Elisa Vianello
ISCAS3
2017 Bioinspired Programming of Resistive Memory Devices for Implementing Spiking Neural Networks
abstract
In this work, we will focus on the role that non-volatile resistive memory technologies (RRAM) can play for modeling key features of biological synapses. We will present an architecture and a reading/programming strategy to emulate both Short and Long Term Plasticity (STP, LTP) rules using non-volatile OxRAM arrays. A visual-pattern extraction application is discussed using spiking neural networks. We demonstrated that Long-Term plasticity allows the neural networks to learn patterns and the Short Term plasticity allows to improve accuracy (reduction of the false positive events generated by white noise in the input data) in presence of significant background noise in the input data.
Elisa Vianello, Thilo Werner, Alessandro Grossi, Etienne Nowak, Barbara De Salvo, Luca Perniola, Olivier Bichler, Blaise Yvert
ACM Great Lakes Symposium on VLSI5
2016 Real-time decoding of brain activity by embedded Spiking Neural Networks using OxRAM synapses
abstract
An innovative approach for decoding of brain signals based on Spiking Neural Networks is presented in this paper. Synapses are implemented by BEOL compatible oxide resistive RAM (OxRAM) devices providing low programming voltages (<;2.5V) and currents (~30μA). Spike-timing-dependent plasticity enables the network for autonomous online spike sorting of measured biological signals. Ultra-low synaptic power consumption in the range of 10nW, recognition rates around 90% and real-time functionality bear high potential for future healthcare applications.
Thilo Werner, Daniele Garbin, Elisa Vianello, Olivier Bichler, Daniel Cattaert, Blaise Yvert, Barbara De Salvo, Luca Perniola
ISCAS7
2015 Emerging resistive memories for low power embedded applications and neuromorphic systems
abstract
In this work, we will focus on the role that new nonvolatile resistive memory technologies can play in emerging fields of application, such as non-volatile logic circuits or neuromorphic circuits, to save energy and increase performance. Concerning the introduction of non-volatile functionalities at the logic level, we will demonstrate hybrid CMOS logic plus ReRAM (specifically CBRAM and OXRAM) circuits for ultra low power FPGA and fixed-logic IC design, as Non Volatile Flip-Flops. Concerning neuromorphic circuits, we will focus on the emulation of synaptic plasticity effects with resistive memory synapses. We will present large-scale energy efficient neuromorphic systems based on ReRAM as stochastic-binary synapses. Prototype applications such as complex visual- and auditory-pattern extraction will be also discussed using feedforward spiking neural networks.
Barbara De Salvo, Elisa Vianello, Olivier Thomas, Fabien Clermidy, Olivier Bichler, Christian Gamrat, Luca Perniola
ISCAS1
2011 Phase change memory for synaptic plasticity application in neuromorphic systems
abstract
In this paper, we show that Phase Change Memory (PCM) can be used to emulate specific functions of a biological synapse similar to Long Term Potentiation (LTP) and Long Term Depression (LTD) plasticity effects. The dependence of synaptic weight on programming pulse width and pulse amplitude is shown experimentally for the PCM devices. Different combinations of consecutive LTD and LTP events have been experimentally demonstrated and analyzed for the PCM synapse.
Manan Suri, Veronique Sousa, Luca Perniola, Dominique Vuillaume, Barbara De Salvo
IJCNN5