EDBT 2026 Demo / reviewers in the wild / expert
Daniele Ielmini
dblp:149/4853
· DBLP profile ↗
19ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-1853-1614ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 4 |
| 2025 | Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 ProjectabstractRecent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini |
DATE | 4 |
| 2025 | In-memory reconstruction of compressively-sampled signals by nonlinear closed-loop analog circuitsabstractTo reduce power consumption associated with data transmission and monitoring in next-generation wireless body sensor networks (WSBNs), compressive sensing (CS), a signal-processing technique for signal reconstruction from a limited number of measurements, was proposed. In CS, knowledge of the signal sparsity under some basis is used to cast the reconstruction problem as the solution of an ℓ1-penalized underdetermined linear system, a data-intensive task in energy-hungry conventional digital solvers. By eliminating the energy and latency overheads associated with continuous data shuttling between the memory and processing units, in-memory computing (IMC) may be a key enabler for highly energy-efficient next-generation WSBNs.Here, we show a novel closed-loop IMC (CL-IMC) circuit for signal reconstruction in the CS framework exploiting nonlinearities of analog operational amplifiers. We derive closed-form equations to describe the circuit operation and characterize its solution time and accuracy. We benchmark the circuit performance in electrocardiogram signal reconstruction against a floating-point 64-bit (FP64) digital solver, obtaining up to 1120 × energy consumption reduction. These results support the position of CL-IMC as a promising candidate for energy-efficient data processing in edge devices for biometric applications. Piergiulio Mannocci, Giuseppe Falcone, Daniele Ielmini |
ISCAS | 3 |
| 2023 | NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated ChipsabstractThe NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2. Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov |
DATE | 20 |
| 2023 | Thermal-Induced Multi-State Memristors for Neuromorphic EngineeringabstractWith the rapidly evolving internet of things (IoT) era, the ever-rising demand for data transfer and storage has put a knotty problem on conventional computers, known as the von Neumann bottleneck and memory wall problem. Slow scaling of CMOS transistors due to physical and economical limitations further exacerbates the situation. It is only logical to mimic what has been known so far as the most energy-efficient system, the human brain. The brain-inspired neuromorphic computing systems compute and store the data locally, which dramatically reduces area and energy consumption. In this work, we demonstrate thermal-induced multi-state memristors for neuromorphic engineering applications. We show that in a neural network that uses a memristor-spintronic nano oscillator connection to implement the synapse-neuron pair, with increased temperature, the total power consumption could be reduced by more than 50 % without degrading the output power of a spintronic-based neuron. Sonal Shreya, Saverio Ricci, Davide Bridarolli, Daniele Ielmini, Hooman Farkhani, Farshad Moradi |
ISCAS | 5 |
| 2023 | Accelerating massive MIMO in 6G communications by analog in-memory computing circuitsabstractWireless communication systems are the backbone of today's digital society. To achieve unprecedented throughput and efficiency, 5G and 6G networks will leverage parallel communication over multiple spatial channels under the massive MIMO (multiple-input multiple-output) framework. One of the main limitations of MIMO is its heavy load of matrix operations, a task for which von Neumann-based computers are deeply unoptimized. This bottleneck can be solved by in-memory computing (IMC), where computation is performed directly within the memory, thus eliminating the constant data shuttling between the memory and the processing unit. Here, we provide an experimental validation of a closed-loop IMC (CL-IMC) system on a 90 nm CMOS integrated circuit. Zero-forcing and regularized-zero-forcing decoding are executed with$14\times 7$MIMO link. The hardware demonstration shows 99.91% accuracy, which is close to a floating-point precision decoder. These results support CL-IMC as a promising candidate for data processing in massive MIMO for next-generation cellular networks. Piergiulio Mannocci, Enrico Melacarne, Giacomo Pedretti, Corrado Villa, Flavio Sancandi, Umberto Spagnolini, Daniele Ielmini |
ISCAS | 7 |
| 2022 | A Hybrid Memristor/CMOS SNN for Implementing One-Shot Winner-Takes-All TrainingabstractThis paper presents a spiking neural network for pattern recognition. The network synapses are realized by resistive switching random access memory (ReRAM) cells, which are a stack of Au/Ti/C/Ti/HfO2/Pt. These cells are connected to an array of NMOS transistors (fabricated in a CMOS 180nm technology) to form a 4by4 1T1R crossbar between pre and postsynaptic circuitries. The pre-synaptic part contains conditioning circuits to reshape the inputs before applying them to the memristive crossbar. The post-synaptic section includes current attenuators that allowed the memristor domain currents to be mapped to neuron domain currents, as well as physiologically realistic neuron circuits fabricated in a CMOS 180nm technology. As a demonstrator, the network is trained with one-shot winner-takes-all method to differentiate four input patterns in its inference mode. Javad Ahmadi-Farsani, Saverio Ricci, Shahin Hashemkhani, Daniele Ielmini, Bernabé Linares-Barranco, Teresa Serrano-Gotarredona |
ISCAS | 4 |
| 2022 | Experimental verification and benchmark of in-memory principal component analysis by crosspoint arrays of resistive switching memoryabstractIn-memory computing (IMC) is gaining momentum as the most promising candidate for the upcoming non-von-Neumann, machine learning-optimized computing paradigm. Its intrinsic parallelism is well-suited to accelerate matrix-vector multiplications (MVM), which prove challenging for traditional architectures and are a fundamental operation in principal component analysis (PCA), one of the most renowned algorithms for data classification. Here, we show an experimental demonstration of a novel, IMC-based PCA algorithm by in-memory power iteration and deflation executed in a 4-kbit array of resistive random-access memory (RRAM). Our algorithm achieves 95.25% classification accuracy on the Wisconsin Diagnostic Breast Cancer dataset, matching closely results of a floating-point machine while providing a $250\times$ improvement in energy efficiency. Piergiulio Mannocci, Andrea Baroni, Enrico Melacarne, Cristian Zambelli, Piero Olivo, Christian Wenger, Daniele Ielmini |
ISCAS | 8 |
| 2022 | End-to-end modeling of variability-aware neural networks based on resistive-switching memory arraysabstractResistive-switching random access memory (RRAM) is a promising technology that enables advanced applications in the field of in-memory computing (IMC). By operating the memory array in the analogue domain, RRAM-based IMC architectures can dramatically improve the energy efficiency of deep neural networks (DNNs). However, achieving a high inference accuracy is challenged by significant variation of RRAM conductance levels, which can be compensated by (i) advanced programming techniques and (ii) variability-aware training (VAT) algorithms. In both cases, however, detailed knowledge and accurate physics-based statistical models of RRAM are needed to develop programming and VAT methodologies. This work presents an end-to-end approach to the development of highly-accurate IMC circuits with RRAM, encompassing the device modeling, the precise programming algorithm, and the VAT simulations to maximize the DNN classification accuracy in presence of conductance variations. Artem Glukhov, Nicola Lepri, Valerio Milo, Andrea Baroni, Cristian Zambelli, Piero Olivo, Christian Wenger, Daniele Ielmini |
VLSI-SoC | 9 |
| 2021 | A Universal, Analog, In-Memory Computing Primitive for Linear Algebra Using MemristorsabstractThe increasing demand for data-intensive computing applications, such as artificial intelligence (AI) and more specifically machine learning (ML), raises the need for novel computing hardware architectures capable of massive parallelism in performing core algebraic operations. Among the new paradigms, in-memory computing (IMC) with analogue devices is attracting significant interest for its large-scale integration potential, together with unrivaled speed and energy performance. Here, we present a fully-analogue, universal primitive capable of executing linear algebra operations such as regression, generalized least-square minimization and linear system solution with and without preconditioning. We study the impact of the main circuit parameters on accuracy and bandwidth with analytical closed-form expressions and SPICE simulations. Scaling challenges due to parasitic resistance/capacitance and their impact on key parameters such as bandwidth and accuracy are discussed. Finally, a comparison with existing solvers belonging to the same IMC framework is made to assess advantages and disadvantages of the proposed circuit. Piergiulio Mannocci, Giacomo Pedretti, Elisabetta Giannone, Enrico Melacarne, Zhong Sun, Daniele Ielmini |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Optimization Schemes for In-Memory Linear Regression Circuit With Memristor ArraysabstractRecently, an in-memory analog circuit based on crosspoint memristor arrays was reported, which enables solving linear regression problems in one step and can be used to train many other machine learning algorithms. To explore its potential for computing accelerator applications, it is of fundamental importance to improve the computing speed of the circuit,i.e., the circuit response towards correct outputs. In this work, we comprehensively studied the transfer function of this circuit, resulting in a quadratic eigenvalue problem that describes the distribution of poles. The minimal real part of non-zero eigenvalues defines the dominant pole, which in turn dominates the response time. Simulations for multiple linear regression solutions with different datasets evidence that, the computing time does not necessarily increase with problem size. The dominant pole is related to parameters in the circuit, including feedback conductance, and gain bandwidth products of operational amplifiers. By optimizing these parameters synergistically, the dominant pole shifts to higher frequencies and the computing speed is consequently optimized. Our results provide a guideline for design and optimization of in-memory machine learning accelerators with analog memristor arrays. Also, issues including power consumption, impact of noise and variation of sources and memristors are investigated to offer a comprehensive evaluation of the circuit performance. Zhong Sun, Shengyu Bao, Yimao Cai, Daniele Ielmini, Ru Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2020 | A Bio-Inspired Recurrent Neural Network with Self-Adaptive Neurons and PCM Synapses for Solving Reinforcement Learning TasksabstractOne of the main challenges in artificial intelligence is the realization of systems capable of learning from their own experience and adapting themselves to a constantly changing environment. In nature, neurobiological systems modify the morphology of the synaptic connections in response to the past experience in order to optimize the interactions with the surrounding world. The introduction of experience-driven mechanisms in artificial systems would thus enable resilience and reinforcement learning in neural networks. Here, we present a novel brain-inspired recurrent neural network (RNN) with PCM synapses capable of advanced tasks such as maze navigation by reinforcement learning. We experimentally demonstrate that the multilevel synaptic capability provided by PCM devices mimics biology and allows a self-optimization of the navigation task. The autonomous agent can rely on PCM-based plasticity and neuronal spike-frequency adaptation to explore the environment and become its own expert teacher via a penalty/reward scheme. From these results, PCM-based local edge computing appears a key concept to enable learning and autonomous navigation in agents such as robots and cars. Stefano Bianchi, I. Muñoz-Martín, Shahin Hashemkhani, Giacomo Pedretti, Daniele Ielmini |
ISCAS | 5 |
| 2020 | Hardware Implementation of PCM-Based Neurons with Self-Regulating Threshold for Homeostatic Scaling in Unsupervised LearningabstractBrain-inspired neuromorphic engineering aims at designing networks capable of learning from their own experience, in terms of both plasticity and stability. In biology, homeostatic scaling can regulate the frequency of neural processing in the brain and enable efficient synaptic learning activity. Implementing homeostatic regulation into hardware neural networks can thus enable stable, energy-efficient learning. Here, we present a novel artificial neuron based on phase change memory (PCM) devices capable of homeostatic regulation and power saving via self-adaptive threshold control. We experimentally show that this mechanism optimizes multi-pattern learning of the Fashion-MNIST dataset with asynchronous spike-timing-dependent plasticity (STDP). The PCM-based adaptive threshold is shown to act as a spike-frequency modulator of the whole neural network, giving robustness to the system against external perturbations. This work highlights the suitability of PCM devices for the optimization of synaptic dynamics and the implementation of brain-inspired neuromorphic circuits for cognitive agents and edge computing. I. Muñoz-Martín, Stefano Bianchi, Shahin Hashemkhani, Giacomo Pedretti, Daniele Ielmini |
ISCAS | 5 |
| 2020 | A Spiking Recurrent Neural Network with Phase Change Memory Synapses for Decision MakingabstractNeuronal activity of recurrent neural networks (RNNs) experimentally observed in the hippocampus is widely believed to play a key role for mammalian ability to associate concepts and make decisions. For this reason, RNNs have rapidly gained strong interest as computational enabler of brain-inspired cognitive functions in hardware. From the technology viewpoint, nonvolatile memory devices such as phase change memory (PCM) and resistive switching memory (RRAM) have become a key asset to allow for high synaptic density and biorealistic cognitive functionality. In this work, we demonstrate for the first time associative learning and decision making in a hardware Hopfield RNN with 6 spiking neurons and PCM synapses via storage, recall and competition of attractor states. We also experimentally demonstrate the solution of a constraint satisfaction problem (CSP) namely a Sudoku with size 2×2 in hardware and 9×9 in simulation. These results support spiking RNNs with PCM devices for the implementation of decision making capabilities in hardware neuromorphic systems. Giacomo Pedretti, Valerio Milo, Shahin Hashemkhani, Piergiulio Mannocci, Octavian Melnic, Elisabetta Chicca, Daniele Ielmini |
ISCAS | 7 |
| 2018 | Brain-inspired recurrent neural network with plastic RRAM synapsesabstractThe development of neuromorphic systems capable of mimicking the behavior of the human brain has recently received an increasing deal of interest. However, the building of such artificial systems has been hindered by the lack of commercial technologies with nanoscale integration of synaptic devices as well as the complexity of the biological neural architecture in terms of connectivity, parallelism, and plasticity behavior. In particular, there is a wide consensus on the relevance of recurrent connections in the human brain, and their key role for associative learning and pattern classification. Fundamental primitives of cognitive computing can therefore be demonstrated by means of Recurrent Neural Networks (RNNs). In this work, we design and simulate a Hopfield-type RNN with HfO2 RRAM devices capable of learning via spike-timing dependent plasticity (STDP). We first demonstrate learning and recall of a single attractor state in a 4-neuron RNN. Based on this result, we then simulate signal restoration of two orthogonal patterns in a 64-neuron RNN, thus supporting RRAM-based RNN with cognitive computing functionalities. Valerio Milo, Elisabetta Chicca, Daniele Ielmini |
ISCAS | 3 |
| 2018 | Resistive switching synapses for unsupervised learning in feed-forward and recurrent neural networksabstractEmerging memory devices such as resistive switching memory (RRAM) and phase change memory (PCM) are gaining interest as future synapses for smart neuromorphic systems, capable of learning and inference similar to the human brain. Developing neuromorphic systems with emerging memory technologies requires accurate co-design of devices, synapses, and neural networks, aiming at the replication of the fundamental learning processes in the human brain, such as spike-timing dependent plasticity (STDP) and spike-rate dependent plasticity (SRDP). This work addresses the development of RRAM synapses for unsupervised learning via STDP. This learning scheme is implemented in a simple one-transistor/one-resistor (1T1R) structure capable of long term potentiation and depression with standard memory-grade RRAM devices. 1T1R synapses are implemented in a spiking neural network (SNN) with feedforward architecture, allowing for the hardware demonstration of unsupervised learning. Recurrent SNNs employing the same fundamental STDP rule are then addressed by simulation of associative learning, pattern reconstruction, and recall of spatiotemporal sequences. Valerio Milo, Giacomo Pedretti, Mario Laudato, Alessandro Bricalli, Elia Ambrosi, Stefano Bianchi, Elisabetta Chicca, Daniele Ielmini |
ISCAS | 8 |
| 2018 | A 4-Transistors/1-Resistor Hybrid Synapse Based on Resistive Switching Memory (RRAM) Capable of Spike-Rate-Dependent Plasticity (SRDP)abstractMimicking the cognitive functions of the brain in hardware is a primary challenge for several fields, including device physics, neuromorphic engineering, and biological neuroscience. A key element in cognitive hardware systems is the ability to learn via biorealistic plasticity rules, combined with the area scaling capability to enable integration of high-density neuron/synapse networks. To this purpose, resistive switching memory (RRAM) devices have recently attracted a strong interest as potential synaptic elements. Here, we present a novel hybrid 4-transistors/1-resistor synapse capable of spike-rate-dependent plasticity. The frequency-dependent learning behavior of the synapse is shown by experiments on HfO2 RRAM devices. Unsupervised learning, update, and recognition of one or more visual patterns in sequence is demonstrated at the level of neural network, thus, supporting the feasibility of hybrid CMOS/RRAM integrated circuits matching the learning capability in the human brain. Valerio Milo, Giacomo Pedretti, Roberto Carboni, Alessandro Calderoni, Nirmal Ramaswamy, Stefano Ambrogio, Daniele Ielmini |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2016 | Neuromorphic computing with hybrid memristive/CMOS synapses for real-time learningabstractResistive (or memristive) devices, including resistive switching memory (RRAM), phase change memory (PCM) an spin-transfer torque memory (STTRAM), are strong candidates for future high-density memory, embedded memory and storage class memory. The availability of resistive-device technology in the industry would pave the way for several other applications in advanced computing, such as neuromorphic cognitive systems and other non-von Neumann approaches to computing. However, the building-block design, functionality and power consumption need to be carefully evaluated to assess all the advantages of the resistive devices with respect to standard CMOS technology. This work will review the recent progress in developing hybrid memristive/CMOS synapses based on either RRAM or PCM, showing the circuit design, the operation concept and the demonstration of real-time spike-based learning and recognition of visual patterns. The learning accuracy and power consumption of the novel synapse blocks will be finally discussed. Daniele Ielmini, Stefano Ambrogio, Valerio Milo, Simone Balatti, Zhongqiang Wang |
ISCAS | 1 |
| 2014 | Statistical modeling of program and read variability in resistive switching devicesabstractResistive-switching memory (RRAM) based on ion migration in metal oxide layers may provide a scalable, low-power alternative to Flash beyond the 10 nm technology node. However, low current operation in scaled RRAM is prone to switching variability and read noise, e.g., random telegraph noise (RTN). To develop a scalable RRAM technology, the statistical variability of program/read processes must be thoroughly addressed. In this paper we propose a novel Monte-Carlo analytical model for switching variability, accounting for the spread of programmed resistance at variable operation current. A numerical model for RTN is then presented, capable of describing size-dependence of RTN amplitude and its kinetics. Stefano Ambrogio, Simone Balatti, Antonio Cubeta, Daniele Ielmini |
ISCAS | 4 |