EDBT 2026 Demo / reviewers in the wild / expert
Giacomo Pedretti
dblp:217/7001
· DBLP profile ↗
18ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-4501-8672ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 11 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On-Access Error Correction in Certain Types of Content-Addressable Memories
Ron M. Roth, Giacomo Pedretti |
IEEE Trans. Inf. Theory | 2 |
| 2025 | Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient AccelerationabstractThe ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FCCM | 6 |
| 2025 | Enhancing FPGAs with Analog In-Memory Computing MacrosabstractWhile the AI computing needs are ever-increasing and the innovation in models generates tens of new architectures yearly, the performance gain from improvements in chip manufacturing has slowed down. Within this context, FPGA-based accelerators play a fundamental role. FPGAs are the backbone of specialized architectures, their reconfigurability being the key differentiation that enables an effective design space exploration. At the same time, to overcome the limitations induced by the memory bottleneck, the computing architectures community has proposed the in-memory computing paradigm: storage and computations are both performed in non-volatile memory devices. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Rand Jean, Xia Sheng, Giacomo Pedretti, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FPGA | 7 |
| 2025 | KLIMA: Low-latency mixed-signal In-Memory Computing accelerator for solving arbitrary-order Boolean Satisfiability
Tinish Bhattacharya, Dongseok Kwon, George Higgins Hutchinson, Xiangyi Zhang, Giacomo Pedretti, Fabian Böhm, John Paul Strachan, Thomas Van Vaerenbergh, Raymond G. Beausoleil, Ignacio Rozada, Dmitri B. Strukov |
HCS | 5 |
| 2025 | RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer AccelerationabstractTransformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Vector-Matrix Multiplications (VMMs) with high computational parallelism and minimal data movement, employing it for other crucial DNN operators remains a formidable task. This challenge is exacerbated by the extensive use of complex activation functions, Softmax, and data-dependent matrix multiplications (DMMuls) within Transformer models. To address this challenge, we introduce a Reconfigurable Analog Computing Engine (RACE) by enhancing Analog Content Addressable Memories (ACAMs) to support broader operations. Based on the RACE, we propose the RACE-IT accelerator (meaning RACE for In-memory Transformers) to enable efficient analog-domain execution of all core operations of Transformer models. Given the flexibility of our proposed RACE in supporting arbitrary computations, RACE-IT is well-suited for adapting to emerging and non-traditional DNN architectures without requiring hardware modifications. We compare RACE-IT with various accelerators. Results show that RACE-IT increases performance by 453× and 15×, and reduces energy by 354× and 122× over the state-of-the-art GPUs and existing Transformer-specific IMC accelerators, respectively. Aishwarya Natarajan, Luca Buonanno, Archit Gajjar, Ron M. Roth, Sergey Serebryakov, John Moon, Omar Eldash, Jim Ignowski, Giacomo Pedretti |
ICCD | 10 |
| 2025 | Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label PropagationabstractTask Scheduling is a popular technique for exploiting parallelism in modern computing systems.In particular, HW-accelerated Task Scheduling has been shown to be effective at improving the performance of fine-grained workloads by dynamically assigning tasks to cores based on their data dependencies with minimal overhead, allowing the handling of tasks with execution times in the order of thousands of cycles.However, the performance of applications assisted by accelerated Task Scheduling is limited by the fact that once a task has all its dependencies fulfilled, it is typically executed on the first available core, which might not be locality-optimal.We thus propose a novel approach to Task Scheduling that leverages HW-accelerated Label Propagation (LP), a graph clustering algorithm, to group tasks with intersecting data patterns such that they are executed on the same core.We show that our approach can significantly improve the performance of task-based applications, improving overall program execution times by up to 1.50× while simultaneously reducing average task sizes by up to 1.81×, augmenting both synthetic benchmarks and real-world applications running on a 24-core RISC-V processor mapped to the Alveo U55C FPGA.These gains rely heavily on the low-latency nature of our proposed label propagation accelerator, which will typically cluster dynamic task graphs in under 300 cycles, up to 581× faster than an equivalent software implementation.Furthermore, by ensuring that ideal placement predictions are used as a hint rather than a hard constraint, we allow the system to benefit from improved data locality for memory-intensive applications while also maintaining high core utilization in compute-bound scenarios.Our results hence demonstrate the potential of HW-accelerated label propagation to improve the performance of Task Scheduling systems with low-latency, dynamic data locality optimization. Lucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo, Giacomo Pedretti, Jim Ignowski, Michael Frank 0008, Xavier Martorell, Daniel Jiménez-González, Carlos Álvarez 0001 |
MICRO | 5 |
| 2024 | CAMSHAP: Accelerating Machine Learning Model Explainability with Analog CAMabstractThe recent success of machine learning (ML) models has led to increasing demands for model explanations - why a result was given - along with model predictions. Tree-based ML models are considered more explainable than deep neural networks and higher performers in several domains. However, algorithms computing model explanations are irregular and scale poorly with model size. While many custom accelerators for training and inference have been proposed, little attention has been paid to accelerating model explanations. This lack of explanatory capability has limited the use of these models for real-time decision-making systems in critical fields such as healthcare, autonomous operation and cybersecurity. John Moon, Giacomo Pedretti, Pedro Bruel, Sergey Serebryakov, Omar Eldash, Luca Buonanno, Catherine Graves, Paolo Faraboschi, Jim Ignowski |
ICCAD | 2 |
| 2024 | Memristive Quaternary Content-Addressable Memories for Implementing Boolean FunctionsabstractIn-memory computing is, in current literature, the most common paradigm used to counteract the Von-Neumann bottleneck, proposing the use of memory elements to define complex input-output relations of the computing kernels. While in classical CMOS computing a similar paradigm can be implemented with look-up tables (LUT), this solution is power and area-hungry. This paper presents the use of Quaternary Content-Addressable Memories (QCAMs), a generalization of the Ternary Content-Addressable Memories (TCAMs), for implementing boolean functions. Content-Addressable Memories can be used as a building block for in-memory processing, using the states of the cells to define a ${\mathbb{B}^{\text{N}}} \to {\mathbb{B}^{\text{M}}}$ function which projects the search word into a new string of bits. The quaternary alphabet allows to represent a more complex function space with respect to the TCAMs while using the same number of cells, enhancing area, power consumption and latency performances achieved when representing arbitrary functions with the CAM hardware. For comparison, it can be demonstrated that QCAMs represent arbitrary Boolean functions with half the number of cells than that would be needed in a standard TCAM implementation, and a ×10 smaller area with respect to SRAM-based LUTs. Along with the table of states and a toy example where the QCAM states are used to define the product among two 2-bit precision real values, this paper presents multiple circuit schemes and encoding schemes for memristor-based QCAMs. Luca Buonanno, Giacomo Pedretti, Aishwarya Natarajan, Todd Richmond, John Moon, Rand Jean, Xia Sheng, Ron M. Roth, Jim Ignowski |
ISCAS | 2 |
| 2024 | On-Access Error Correction in Certain Types of Content-Addressable MemoriesabstractA summing content-addressable memory ($\Sigma$-CAM) is a device which consists of an$\ell\times n$array of cells, where each cell can be programmed to one of two states, 0 or 1. The input to the array is a binary$\ell$-vector, and the output is an integer n-vector whose entries are the Hamming distances between the input vector and the contents of the columns. A summing ternary CAM ($\Sigma$-TCAM) is a variant of a$\Sigma$-CAM where the state/input alphabet is endowed with a third symbol (“don't care”) whose Hamming distance from any symbol is defined to be 0. The purpose of this work is to present coding schemes for on-access correction of errors in both$\Sigma$-CAMs and$\Sigma$-TCAMs. In the case of$\Sigma$-CAMs, our scheme builds upon schemes that have been proposed for discrete vector-matrix multipliers. The adaptation of such schemes to$\Sigma$-TCAMs, however, is more involved and requires a special type of positional binary representation of integer pairs where the representations of the two integers in a pair do not share a 1 in the same position. Ron M. Roth, Giacomo Pedretti |
ISIT | 2 |
| 2024 | RD-FAXID: Ransomware Detection with FPGA-Accelerated XGBoostabstractOver the last decade, there has been a rise in cyberattacks, particularly ransomware, causing significant disruption and financial repercussions across public and private sectors. Tremendous efforts have been spent on developing techniques to detect ransomware to, ideally, protect data or have as minimum data loss as possible. Ransomware attacks are becoming more frequent and sophisticated as there is a constant tussle between attackers and cybersecurity defenders. Machine Learning (ML) approaches have proven more effective in detecting ransomware than classical signature-based detection. In particular, tree-based algorithms such as Decision Trees (DT), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) spike up interest among cybersecurity researchers. However, due to the nature of the problem, traditional CPUs and GPUs fail to keep up with the desired performance, especially for large data workloads. Thus, the problem demands a customized solution to detect the ransomware. Here, we propose an FPGA accelerated tree-based ML model for multi-dataset ransomware detection. We show the capability of the proposed prototype to address the problem from more than one set of features, reducing false positive and negative rates to have robust predictions by looking at Hardware Performance Counters (HPCs), Operating System (OS) calls, and network traffic information simultaneously. With 1,000 samples per batch, the FPGA prototype has 65.8 \({\times}\) and 4.1 \({\times}\) lower latency over the CPU and GPU, respectively. Moreover, the FPGA design is up to 11.3 \({\times}\) cost-effective and 643 \({\times}\) energy-efficient compared to the CPU and 3 \({\times}\) cost-effective and 16.8 \({\times}\) energy-efficient over the GPU. Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Chris Cheng, Giacomo Pedretti, Jim Ignowski |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2023 | Accelerating massive MIMO in 6G communications by analog in-memory computing circuitsabstractWireless communication systems are the backbone of today's digital society. To achieve unprecedented throughput and efficiency, 5G and 6G networks will leverage parallel communication over multiple spatial channels under the massive MIMO (multiple-input multiple-output) framework. One of the main limitations of MIMO is its heavy load of matrix operations, a task for which von Neumann-based computers are deeply unoptimized. This bottleneck can be solved by in-memory computing (IMC), where computation is performed directly within the memory, thus eliminating the constant data shuttling between the memory and the processing unit. Here, we provide an experimental validation of a closed-loop IMC (CL-IMC) system on a 90 nm CMOS integrated circuit. Zero-forcing and regularized-zero-forcing decoding are executed with$14\times 7$MIMO link. The hardware demonstration shows 99.91% accuracy, which is close to a floating-point precision decoder. These results support CL-IMC as a promising candidate for data processing in massive MIMO for next-generation cellular networks. Piergiulio Mannocci, Enrico Melacarne, Giacomo Pedretti, Corrado Villa, Flavio Sancandi, Umberto Spagnolini, Daniele Ielmini |
ISCAS | 3 |
| 2022 | A general tree-based machine learning accelerator with memristive analog CAMabstractDeep learning models have reached high accuracy in multiple classification tasks. However these models lack explainability, namely the capability of understanding why a certain class is chosen along with the class predicted. On the other hand, tree-based models are top performers in several applications, particularly when the training set is limited, while also being more explainable. However, tree-based models are difficult to accelerate with conventional digital hardware due to irregular memory access patterns. Here we show a tree-based ML accelerator based on a novel analog content addressable memory with memristor devices, capable of handling multiple types of bagging and boosting techniques common in tree-based algorithms. Our results show a large improvement of $\sim 60 \times $ lower latency and $160 \times $ reduced energy consumption compared to the state of the art, demonstrating the promise of our accelerator approach. Giacomo Pedretti, Sergey Serebryakov, John Paul Strachan, Catherine Graves |
ISCAS | 1 |
| 2021 | A Universal, Analog, In-Memory Computing Primitive for Linear Algebra Using MemristorsabstractThe increasing demand for data-intensive computing applications, such as artificial intelligence (AI) and more specifically machine learning (ML), raises the need for novel computing hardware architectures capable of massive parallelism in performing core algebraic operations. Among the new paradigms, in-memory computing (IMC) with analogue devices is attracting significant interest for its large-scale integration potential, together with unrivaled speed and energy performance. Here, we present a fully-analogue, universal primitive capable of executing linear algebra operations such as regression, generalized least-square minimization and linear system solution with and without preconditioning. We study the impact of the main circuit parameters on accuracy and bandwidth with analytical closed-form expressions and SPICE simulations. Scaling challenges due to parasitic resistance/capacitance and their impact on key parameters such as bandwidth and accuracy are discussed. Finally, a comparison with existing solvers belonging to the same IMC framework is made to assess advantages and disadvantages of the proposed circuit. Piergiulio Mannocci, Giacomo Pedretti, Elisabetta Giannone, Enrico Melacarne, Zhong Sun, Daniele Ielmini |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | A Bio-Inspired Recurrent Neural Network with Self-Adaptive Neurons and PCM Synapses for Solving Reinforcement Learning TasksabstractOne of the main challenges in artificial intelligence is the realization of systems capable of learning from their own experience and adapting themselves to a constantly changing environment. In nature, neurobiological systems modify the morphology of the synaptic connections in response to the past experience in order to optimize the interactions with the surrounding world. The introduction of experience-driven mechanisms in artificial systems would thus enable resilience and reinforcement learning in neural networks. Here, we present a novel brain-inspired recurrent neural network (RNN) with PCM synapses capable of advanced tasks such as maze navigation by reinforcement learning. We experimentally demonstrate that the multilevel synaptic capability provided by PCM devices mimics biology and allows a self-optimization of the navigation task. The autonomous agent can rely on PCM-based plasticity and neuronal spike-frequency adaptation to explore the environment and become its own expert teacher via a penalty/reward scheme. From these results, PCM-based local edge computing appears a key concept to enable learning and autonomous navigation in agents such as robots and cars. Stefano Bianchi, I. Muñoz-Martín, Shahin Hashemkhani, Giacomo Pedretti, Daniele Ielmini |
ISCAS | 4 |
| 2020 | Hardware Implementation of PCM-Based Neurons with Self-Regulating Threshold for Homeostatic Scaling in Unsupervised LearningabstractBrain-inspired neuromorphic engineering aims at designing networks capable of learning from their own experience, in terms of both plasticity and stability. In biology, homeostatic scaling can regulate the frequency of neural processing in the brain and enable efficient synaptic learning activity. Implementing homeostatic regulation into hardware neural networks can thus enable stable, energy-efficient learning. Here, we present a novel artificial neuron based on phase change memory (PCM) devices capable of homeostatic regulation and power saving via self-adaptive threshold control. We experimentally show that this mechanism optimizes multi-pattern learning of the Fashion-MNIST dataset with asynchronous spike-timing-dependent plasticity (STDP). The PCM-based adaptive threshold is shown to act as a spike-frequency modulator of the whole neural network, giving robustness to the system against external perturbations. This work highlights the suitability of PCM devices for the optimization of synaptic dynamics and the implementation of brain-inspired neuromorphic circuits for cognitive agents and edge computing. I. Muñoz-Martín, Stefano Bianchi, Shahin Hashemkhani, Giacomo Pedretti, Daniele Ielmini |
ISCAS | 4 |
| 2020 | A Spiking Recurrent Neural Network with Phase Change Memory Synapses for Decision MakingabstractNeuronal activity of recurrent neural networks (RNNs) experimentally observed in the hippocampus is widely believed to play a key role for mammalian ability to associate concepts and make decisions. For this reason, RNNs have rapidly gained strong interest as computational enabler of brain-inspired cognitive functions in hardware. From the technology viewpoint, nonvolatile memory devices such as phase change memory (PCM) and resistive switching memory (RRAM) have become a key asset to allow for high synaptic density and biorealistic cognitive functionality. In this work, we demonstrate for the first time associative learning and decision making in a hardware Hopfield RNN with 6 spiking neurons and PCM synapses via storage, recall and competition of attractor states. We also experimentally demonstrate the solution of a constraint satisfaction problem (CSP) namely a Sudoku with size 2×2 in hardware and 9×9 in simulation. These results support spiking RNNs with PCM devices for the implementation of decision making capabilities in hardware neuromorphic systems. Giacomo Pedretti, Valerio Milo, Shahin Hashemkhani, Piergiulio Mannocci, Octavian Melnic, Elisabetta Chicca, Daniele Ielmini |
ISCAS | 1 |
| 2018 | Resistive switching synapses for unsupervised learning in feed-forward and recurrent neural networksabstractEmerging memory devices such as resistive switching memory (RRAM) and phase change memory (PCM) are gaining interest as future synapses for smart neuromorphic systems, capable of learning and inference similar to the human brain. Developing neuromorphic systems with emerging memory technologies requires accurate co-design of devices, synapses, and neural networks, aiming at the replication of the fundamental learning processes in the human brain, such as spike-timing dependent plasticity (STDP) and spike-rate dependent plasticity (SRDP). This work addresses the development of RRAM synapses for unsupervised learning via STDP. This learning scheme is implemented in a simple one-transistor/one-resistor (1T1R) structure capable of long term potentiation and depression with standard memory-grade RRAM devices. 1T1R synapses are implemented in a spiking neural network (SNN) with feedforward architecture, allowing for the hardware demonstration of unsupervised learning. Recurrent SNNs employing the same fundamental STDP rule are then addressed by simulation of associative learning, pattern reconstruction, and recall of spatiotemporal sequences. Valerio Milo, Giacomo Pedretti, Mario Laudato, Alessandro Bricalli, Elia Ambrosi, Stefano Bianchi, Elisabetta Chicca, Daniele Ielmini |
ISCAS | 2 |
| 2018 | A 4-Transistors/1-Resistor Hybrid Synapse Based on Resistive Switching Memory (RRAM) Capable of Spike-Rate-Dependent Plasticity (SRDP)abstractMimicking the cognitive functions of the brain in hardware is a primary challenge for several fields, including device physics, neuromorphic engineering, and biological neuroscience. A key element in cognitive hardware systems is the ability to learn via biorealistic plasticity rules, combined with the area scaling capability to enable integration of high-density neuron/synapse networks. To this purpose, resistive switching memory (RRAM) devices have recently attracted a strong interest as potential synaptic elements. Here, we present a novel hybrid 4-transistors/1-resistor synapse capable of spike-rate-dependent plasticity. The frequency-dependent learning behavior of the synapse is shown by experiments on HfO2 RRAM devices. Unsupervised learning, update, and recognition of one or more visual patterns in sequence is demonstrated at the level of neural network, thus, supporting the feasibility of hybrid CMOS/RRAM integrated circuits matching the learning capability in the human brain. Valerio Milo, Giacomo Pedretti, Roberto Carboni, Alessandro Calderoni, Nirmal Ramaswamy, Stefano Ambrogio, Daniele Ielmini |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |