Elena Ferro

dblp:329/6441 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-8618-8643ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CIM-FLEX: An Integer-Only Flexible Periphery for Distributed Compute-In-Memory Architectures
abstract
Distributed compute-in-memory (CIM) architectures are emerging as a path towards energy-efficient and highly parallel edge-class deep neural network (DNN) inference. However, the lack of a flexible, unified digital periphery limits the ability of distributed CIM systems to support integer-quantized DNN inference schemes, activation functions, and scalable tile-to-tile dataflow. We present CIM-FLEX, an integer-only periphery architecture that enables each CIM tile to autonomously perform partial-sum rescaling, asymmetric activation processing, and lightweight programmable activation functions. CIM-FLEX also supports mixed-precision integer multiply-and-accumulate (MAC) by decomposing higher-precision operations into uniform-precision partial MACs, without modifying existing CIM primitives. CIM-FLEX enables flexible tile-to-tile dataflow across distributed CIM tiles, exposing a tradeoff between parallel output generation and energy efficiency, where deeper accumulation chains reduce concurrent periphery utilization and improve overall efficiency. Numerical studies and TFLite-based LUT activation experiments validate a high signal-to-noise ratio (∼30 dB) for programmable activation functions, and a 14 nm CMOS implementation shows a 0.012 mm2 area for the CIM-FLEX.
Irem Sanli, Michele Rossi, Elena Ferro, Andreas Burg, Surinder Pal Singh, Thomas Boesch, Irem Boybat
ISLPED3
2026 HARP: Heterogeneous Analog-Digital Resource-Aware Performance and Scheduling Framework for Transformer Acceleration
Elena Ferro, Hadjer Benmeziane, Irem Boybat
IEEE Trans. Parallel Distributed Syst.1
2025 Multi-Mode Borderguard Controllers for Efficient On-Chip Communication in Heterogeneous Digital/Analog Neural Processing Units
abstract
Driven by the growing demand for data-intensive parallel computation, particularly for Matrix-Vector Multiplications (MVMs), and the pursuit of high energy efficiency, Analog In-Memory Computing (AIMC) has garnered significant attention. AIMC addresses the data movement bottleneck by performing MVMs directly within memory, significantly reducing latency and enhancing energy efficiency. Integrating AIMC with digital units for non-MVM operations yields heterogeneous Neural Processing Units (NPUs) that can be combined in a tiled architecture to deliver promising solutions for end-to-end AI inference. Besides powerful heterogeneous NPUs, an efficient on-chip communication infrastructure is also pivotal for inter-node data transmission and efficient AI model execution. This paper introduces the Borderguard Controller (BG-CTRL), a multi-mode, path-through routing controller designed to support three distinct operating modes-time-scheduling, data-driven, and time-sliced data-driven (TSDD)-each offering varying levels of routing flexibility and energy efficiency depending on the data flow patterns and AI model complexity. To demonstrate the design, BG-CTRLs are integrated into a 9-node system of heterogeneous NPUs, arranged in a 3x3 grid and connected using a 2D mesh topology. The system is synthesized using STM 28nm FD-SOI technology. Experimental results show that the BG-CTRL cluster achieves an aggregate throughput of 983 Gb/s, with an energy efficiency of up to 0.41 pJ/B/hop at 0.64 GHz, and a minimal area overhead of 204 kGE.
Hong Pang, Carmine Cappetta, Riccardo Massa, Athanasios Vasilopoulos, Elena Ferro, Gamze Islamoglu, Angelo Garofalo, Francesco Conti 0001, Luca Benini, Irem Boybat, Thomas Boesch
DATE5
2025 NIMA: Near In-Memory High-Precision Accumulation Unit for Heterogeneous Analog/Digital Deep Learning Acceleration
abstract
Analog In-Memory Computing (AIMC) crossbars often face size limitations that hinder mapping entire neural network layers onto a single AIMC tile. To overcome this, tall layers are typically distributed across multiple tiles or within a single tile by multiplexing different sets of weights at various time intervals. However, this generates partial vector-matrix multiplication (VMM) results that need to be accumulated to produce the final output, underscoring the need for integrated accumulation capabilities in AIMC systems. In this work, we propose a Near In-Memory High-Precision Accumulation Unit (NIMA) with built-in internal and external tile accumulation functionality, positioned at the periphery of the AIMC crossbar. This unit leverages the affine correction capabilities of existing systems and ensures high precision during partial VMM accumulation. We have physically implemented NIMA in a 14nm CMOS technology, providing comprehensive performance and area evaluations and comparisons with prior art. We investigate and compare different layer mappings and accumulation schemes supported by the proposed unit, evaluating both latency and Mean Absolute Error (MAE). We further assess the precision of the proposed unit on large layers of ResNet9 and ResNet32 for image classification on CIFAR10/CIFAR100 datasets.
Irem Sanli, Elena Ferro, Athanasios Vasilopoulos, Thomas Boesch, Abu Sebastian, Irem Boybat
ISCAS2
2025 CiMBA: Accelerating Genome Sequencing Through On-Device Basecalling via Compute-in-Memory
abstract
As genome sequencing is finding utility in a wide variety of domains beyond the confines of traditional medical settings, its computational pipeline faces two significant challenges. First, the creation of up to 0.5 GB of data per minute imposes substantial communication and storage overheads. Second, the sequencing pipeline is bottlenecked at the basecalling step, consuming >40% of genome analysis time. A range of proposals have attempted to address these challenges, with limited success. We propose to address these challenges with a Compute-in-Memory Basecalling Accelerator (CiMBA), the first embedded ($\sim 25$mm$^{2}$) accelerator capable of real-time, on-device basecalling, coupled with AnaLog (AL)-Dorado, a new family of analog focused basecalling DNNs. Our resulting hardware/software co-design greatly reduces data communication overhead, is capable of a throughput of 4.77 million bases per second, 24× that required for real-time operation, and achieves 17 × /27× power/area efficiency over the best prior basecalling embedded accelerator while maintaining a high accuracy comparable to state-of-the-art software basecallers.
William Andrew Simon, Irem Boybat, Riselda Kodra, Elena Ferro, Gagandeep Singh 0002, Mohammed Alser, Shubham Jain 0004, Hsinyu Tsai, Geoffrey W. Burr, Onur Mutlu, Abu Sebastian
IEEE Trans. Parallel Distributed Syst.4
2024 A Precision-Optimized Fixed-Point Near-Memory Digital Processing Unit for Analog In-Memory Computing
abstract
Analog In-Memory Computing (AIMC) is an emerging technology for fast and energy-efficient Deep Learning (DL) inference. However, a certain amount of digital post-processing is required to deal with circuit mismatches and non-idealities associated with the memory devices. Efficient near-memory digital logic is critical to retain the high area/energy efficiency and low latency of AIMC. Existing systems adopt Floating Point 16 (FP16) arithmetic with limited parallelization capability and high latency. To overcome these limitations, we propose a Near-Memory digital Processing Unit (NMPU) based on fixed-point arithmetic. It achieves competitive accuracy and higher computing throughput than previous approaches while minimizing the area overhead. Moreover, the NMPU supports standard DL activation steps, such as ReLU and Batch Normalization. We perform a physical implementation of the NMPU design in a 14 nm CMOS technology and provide detailed performance, power, and area assessments. We validate the efficacy of the NMPU by using data from an AIMC chip and demonstrate that a simulated AIMC system with the proposed NMPU outperforms existing FP16- based implementations, providing 139 × speed-up, 7.8 × smaller area, and a competitive power consumption. Additionally, our approach achieves an inference accuracy of 86.65 %/65.06 %, with an accuracy drop of just 0.12 %/0.4 % compared to the FP16 baseline when benchmarked with ResNet9/ResNet32 networks trained on the CIFAR10/CIFAR100 datasets, respectively.
Elena Ferro, Athanasios Vasilopoulos, Corey Lammie, Manuel Le Gallo, Luca Benini, Irem Boybat, Abu Sebastian
ISCAS1