VLDB 2026 Research / reviewers in the wild / expert
Cristian Zambelli
dblp:144/4826
· DBLP profile ↗
18ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-8755-0504ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 5 |
| 2026 | Benchmarking a DNN for aortic valve calcium lesions segmentation on FPGA-based DPU using the vitis AI toolchainabstractSemantic segmentation assigns a class to every pixel of an image to automatically locate objects in the context of computer vision applications for autonomous vehicles, robotics, agriculture, gaming, and medical imaging. Deep Neural Network models, such as Convolutional Neural Networks (CNNs), are widely used for this purpose. Among the plethora of models, the U-Net is a standard in biomedical imaging. Nowadays, GPUs efficiently perform segmentation and are the reference architectures for running CNNs, and FPGAs compete for inferences among alternative platforms, promising higher energy efficiency and lower latency solutions. In this contribution, we evaluate the performance of FPGA-based Deep Processing Units (DPUs) implemented on the AMD Alveo U55C for the inference task, using calcium segmentation in cardiac aortic valve computer tomography scans as a benchmark. We design and implement a U-Net-based application, optimize the hyperparameters to maximize the prediction accuracy, perform pruning to simplify the model, and use different numerical quantizations to exploit low-precision operations supported by the DPUs and GPUs to boost the computation time. We describe how to port and deploy the U-Net model on DPUs, and we compare accuracy, throughput, and energy efficiency achieved with four generations of GPUs and a recent dual 32-core high-end CPU platform. Our results show that a complex DNN like the U-Net can run effectively on DPUs using 8-bit integer computation, achieving a prediction accuracy of approximately 95 % in Dice and 91 % in IoU scores. These results are comparable to those measured when running the floating-point models on GPUs and CPUs. On the one hand, in terms of computing performance, the DPUs achieves a inference latency of approximately 3.5 ms and a throughput of approximately 4.2 kPFS, boosting the performance of a 64-core CPU system by approximately 10 % in terms of latency and a factor 2 X in terms of throughput, but still do not overcoming the performance of GPUs when using the same numerical precision. On the other hand, considering the energy efficiency, the improvements are approximately a factor 6.7 X compared to the CPU, and 1.6 X compared to the P100 GPU manufactured with the same technological process (16 nm). Valentina Sisini, Andrea Miola, Giada Minghini, Enrico Calore, Armando Ugo Cavallo, Sebastiano Fabio Schifano, Cristian Zambelli |
Future Gener. Comput. Syst. | 7 |
| 2025 | Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 ProjectabstractRecent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini |
DATE | 9 |
| 2025 | Transforming Digital Electronics Education: Integrating Sustainability Through Eco-Design, Modularity, and Circular PracticesabstractIntegrating eco-design principles and open hardware into digital electronics education is crucial to equip future engineers for addressing the environmental issues of the digital era. This research is part of the DEEPCEL (Digital Electronics with Eco-Designed Paradigm in Collaborative Enhanced Learning) ERASMUS+ project. Initially, it explores the knowledge and perceptions of participants through a targeted survey, emphasizing methodologies, the environmental impact of open hardware, and their role in fostering sustainability in digital electronics. The research also explores effective teaching strategies, tools, and platforms that can improve the adoption of eco-transition principles, with the aim of creating guidelines for curriculum modernization focused on sustainability. Participants affirm that open hardware offers transformative potential in improving innovation, accessibility, and eco-conscious practices and community-driven development. This study underscores the need for a collaborative effort between academia, professors, and industry leaders, along with their demands and needs to prioritize and integrate sustainability principles into digital electronics education. However, significant barriers remain, including insufficient training for professors, limited resources, and resistance to changes in traditional pedagogical practices. The findings should be interpreted with caution due to potential limitations in sample representation. Further research will validate the effectiveness of proposed strategies and assess long-term impacts on educational practices. Ignacio Bravo Muñoz, Ernesto Martín, Etienne Lemaire, Jean-Paul Chemla, Cristian Zambelli, Sebastiano Fabio Schifano, Hélio Sousa Mendonça, José Carlos Alves |
EDUCON | 6 |
| 2025 | ReDiM: An Efficient Strategy for Read Disturb Mitigation in RRAM-Based AcceleratorsabstractResistive RAM (RRAM) has emerged as a promising non-volatile memory technology for implementing energy-efficient hardware accelerators within the in-memory computing (IMC) paradigm. However, due to the immature fabrication process and inherent material instabilities, frequent read operations during computations can induce read disturb effects, leading to unintended resistance drift and potential data corruption. Existing mitigation approaches primarily focus on detecting read disturb effects and triggering memory refresh operations. In this work, we propose an architecture-level solution that mitigates read disturb in RRAM-based accelerators. Our strategy employs crossbar duplication and decomposes the single high input pulse into two lower-amplitude pulses, effectively minimizing the risk of read disturb. To validate our approach, we develop a simulation framework that incorporates measurement data from characterized RRAM devices under read disturb stress conditions. Experimental results on VGG-8 with CIFAR-10 demonstrate that the proposed method significantly mitigates inference accuracy degradation caused by read disturb in RRAM-based accelerators, while incurring modest area and energy overheads of 12.32% and 2.15%, respectively. This work provides a practical and scalable solution for enhancing the robustness of RRAM-based accelerators in edge and high-performance computing applications. Jianan Wen, Andrea Baroni, Alberto Mistroni, Cristian Zambelli, Christian Wenger, Milos Krstic, Letícia Maria Veiras Bolzani |
IOLTS | 5 |
| 2025 | High throughput edit distance computation on FPGA-based accelerators using HLSabstractEdit distance is a computational grand challenge problem to quantify the minimum number of editing operations required to modify one string of characters to the other, finding many applications of natural language processing. In recent years, relevant and increasing interest has also emerged from deoxyribonucleic acid (DNA) applications, like Next Generation Sequencing and DNA storage technologies. Both applications share two crucial features: i) the information is coded into the four bases of DNA and ii) the level of operational noise is still high causing errors in the data, requiring inclusion in the workflow of the computation of algorithms such as the edit distance for finding similarities between sequences. To boost this computation many solutions are available in the literature. Among them, the FPGAs are largely used since the data domain of those applications is strings of 4 characters represented as two-bit values, inconveniently fitting the basic data types of ordinary CPUs and GPUs, with additional benefits of providing a high level of parallelism and low processing latency. This contribution presents a computing- and energy-efficient design implementing the edit distance algorithm combining metaprogramming and High-Level Synthesis. We also assess the performance of our design targeting recent FPGA-based accelerators. Our solution uses nearly 90% of FPGA basic-block hardware resources achieving about 90% of computing efficiency delivering a maximum throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule, enabling the use of FPGAs as a new class of accelerators for High Performance Computing in DNA applications. • Computing- and energy-efficient edit distance algorithms running on FPGA accelerators. • Targeting short reads for DNA data storage and long reads for genomics applications. • Combining metaprogramming and HLS to design and optimize C codes for FPGA devices. • Achieving about 90% of computing efficiency and near 90% FPGA resources occupancy. • Delivering a peak throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule. Sebastiano Fabio Schifano, Marco Reggiani, Enrico Calore, Rino Micheloni, Alessia Marelli, Cristian Zambelli |
Future Gener. Comput. Syst. | 6 |
| 2023 | On the Reliability of RRAM-Based Neural NetworksabstractEmerging device technologies such as Resistive RAMs (RRAMs) are under investigation by many researchers and semiconductor companies; not only to realize e.g., embedded non-volatile memories, but also to enable energy-efficient computing making use of new data processing paradigms such as computation-in-memory. However, such devices suffer from various non-idealities and reliability failure mechanisms (e.g., variability, endurance, and retention); these negatively impact the memory robustness and the computation accuracy. This paper discusses the non-idealities and reliability failure mechanisms for RRAM devices, provides an overview on the most popular ones. In addition, it reports detailed anlysis of some of these based on data measurements. Finally, it presents two different mitigation schemes for RRAM based accelerators; one is based on RRAM non-ideality aware quantization and conductance control for neural network accuracy enhancement while the second is based on reliability-aware biased training technique. Hassen Aziza, Cristian Zambelli, Said Hamdioui, Sumit Diware, Rajendra Bishnoi, Anteneh Gebregiorgis |
VLSI-SoC | 2 |
| 2022 | Experimental verification and benchmark of in-memory principal component analysis by crosspoint arrays of resistive switching memoryabstractIn-memory computing (IMC) is gaining momentum as the most promising candidate for the upcoming non-von-Neumann, machine learning-optimized computing paradigm. Its intrinsic parallelism is well-suited to accelerate matrix-vector multiplications (MVM), which prove challenging for traditional architectures and are a fundamental operation in principal component analysis (PCA), one of the most renowned algorithms for data classification. Here, we show an experimental demonstration of a novel, IMC-based PCA algorithm by in-memory power iteration and deflation executed in a 4-kbit array of resistive random-access memory (RRAM). Our algorithm achieves 95.25% classification accuracy on the Wisconsin Diagnostic Breast Cancer dataset, matching closely results of a floating-point machine while providing a $250\times$ improvement in energy efficiency. Piergiulio Mannocci, Andrea Baroni, Enrico Melacarne, Cristian Zambelli, Piero Olivo, Christian Wenger, Daniele Ielmini |
ISCAS | 4 |
| 2022 | End-to-end modeling of variability-aware neural networks based on resistive-switching memory arraysabstractResistive-switching random access memory (RRAM) is a promising technology that enables advanced applications in the field of in-memory computing (IMC). By operating the memory array in the analogue domain, RRAM-based IMC architectures can dramatically improve the energy efficiency of deep neural networks (DNNs). However, achieving a high inference accuracy is challenged by significant variation of RRAM conductance levels, which can be compensated by (i) advanced programming techniques and (ii) variability-aware training (VAT) algorithms. In both cases, however, detailed knowledge and accurate physics-based statistical models of RRAM are needed to develop programming and VAT methodologies. This work presents an end-to-end approach to the development of highly-accurate IMC circuits with RRAM, encompassing the device modeling, the precise programming algorithm, and the VAT simulations to maximize the DNN classification accuracy in presence of conductance variations. Artem Glukhov, Nicola Lepri, Valerio Milo, Andrea Baroni, Cristian Zambelli, Piero Olivo, Christian Wenger, Daniele Ielmini |
VLSI-SoC | 5 |
| 2018 | Experimental Investigation of 4-kb RRAM Arrays Programming Conditions Suitable for TCAMabstractResistive random access memories (RRAMs) feature high-speed operations, low-power consumption, and nonvolatile retention, thus serving as a promising candidate for future memory applications. To explore the applications of the RRAM, switching variability and cycling endurance need to be addressed. This paper presents extensive characterizations of multi-kb RRAM arrays during forming, set, reset, and cycling operations. The relationships among programming conditions, memory window, and endurance features are presented. The experimental results are then used to perform variability-aware simulations of a 128-bit RRAM-based ternary content-addressable-memory (TCAM) macro. The tradeoff among endurance, search latency, and reliability in terms of match/mismatch detection is explored, identifying the programming conditions that allow to obtain a searching speed comparable to static random access memory-based TCAMs (2 ns on average and 3 ns at 3σ) while guaranteeing good reliability metrics (with a time ratio of 3000 on average and 150 at 3σ). Alessandro Grossi, Elisa Vianello, Cristian Zambelli, Pablo Royer, Jean-Philippe Noël, Bastien Giraud, Luca Perniola, Piero Olivo, Etienne Nowak |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Phase Change and Magnetic Memories for Solid-State Drive ApplicationsabstractThe state-of-the-art solid-state drives (SSDs) now heterogeneously integrate NAND Flash and dynamic random access memories (DRAMs) to partially hide the limitation of the nonvolatile memory technology. However, due to the increased request for storage density coupled with performance that positions the storage tier closer to the latency of the processing elements, NAND Flash are becoming a serious bottleneck. DRAM as well are a limitation in the SSD reliability due to their vulnerability to the power loss events. Several emerging memory technologies are candidate to replace them, namely the storage class memories. Phase change memories and magnetic memories fall into this category. In this work, we review both technologies from the perspective of their possible application in future disk drives, opening up new computation paradigms as well as improving the storage characteristics in terms of latency and reliability. Cristian Zambelli, Gabriele Navarro, Veronique Sousa, Ioan Lucian Prejbeanu, Luca Perniola |
Proc. IEEE | 1 |
| 2017 | Solid-State Drives: Memory Driven Design Methodologies for Optimal PerformanceabstractSolid-state drives (SSDs) faced an astonishing development in the last few years, becoming the cornerstone to new paradigms and markets of the information technology, such as cloud computing and big data centers. So far, the SSD design approach has focused on the optimization of the Flash translation layer, the firmware devoted to fulfill the compatibility with traditional hard-disk drives. For hyperscaled SSDs this strategy is no longer valid since their performance and reliability are strictly linked to that of the NAND Flash memories that constitute the storage medium, in particular when the multilevel cell paradigm is considered. For this reason, the design flow must follow a bottom-up approach that, starting from an accurate knowledge of the time and use dependent reliability of the NAND Flash memories, selects the most appropriate error correction strategy to extend the SSD lifetime while reducing its performance degradation. Then, the design flow moves to that of the SSD controller and of the interface toward the host where the application is running. This paper will thoroughly discuss this bottom-up approach, and finally, it will show how it is possible to leverage new approaches, such as the software-defined storage system that, by exploiting a hardware/software codesign of the SSD controller architecture and of the host application, will be able to revolutionize the traditional computer/memory interaction. Lorenzo Zuolo, Cristian Zambelli, Rino Micheloni, Piero Olivo |
Proc. IEEE | 2 |
| 2016 | ATHENIS_3D: Automotive tested high-voltage and embedded non-volatile integrated SoC platform with 3D technology
Ewald Wachmann, Sergio Saponara, Cristian Zambelli, Pierre Tisserand, J. Charbonnier, Tobias Erlbacher, S. Gruenler, C. Hartler, Jörg Siegert, Pierre Chassard, D. M. Ton, Lorenzo Ferrari, Luca Fanucci |
DATE | 3 |
| 2015 | SSDExplorer: A Virtual Platform for Performance/Reliability-Oriented Fine-Grained Design Space Exploration of Solid State DrivesabstractCurrently available electronic design automation tools for design space exploration of solid state drives (SSDs) are not able to assess: 1) the device architecture inefficiencies; 2) architecture overdesign for a target performance; and 3) performance degradation caused by the disk usage. These tools feature either an overly high abstraction modeling strategy or lack the required flexibility to perform design exploration. To overcome these problems, this paper proposes SSDExplorer, a tool for fine-grained yet reasonably fast design space exploration of different SSD architectures highlighting possible bottlenecks. To prove its accuracy SSDExplorer has been validated with two real SSDs. SSDExplorer efficiency has been assessed by evaluating the impact of the NAND flash read retry algorithm impact on the SSD performance as a function of its internal architecture. Lorenzo Zuolo, Cristian Zambelli, Rino Micheloni, Marco Indaco, Stefano Di Carlo, Paolo Prinetto, Davide Bertozzi, Piero Olivo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Performance and Reliability Analysis of Cross-Layer Optimizations of NAND Flash ControllersabstractNAND flash memories are becoming the predominant technology in the implementation of mass storage systems for both embedded and high-performance applications. However, when considering data and code storage in Non-Volatile Memories (NVMs), such as NAND flash memories, reliability and performance become a serious concern for systems designers. Designing NAND flash-based systems based on worst-case scenarios leads to waste of resources in terms of performance, power consumption, and storage capacity. This is clearly in contrast with the request for runtime reconfigurability, adaptivity, and resource optimization in modern computing systems. There is a clear trend toward supporting differentiated access modes in flash memory controllers, each one setting a differentiated tradeoff point in the performance-reliability optimization space. This is supported by the possibility of tuning the NAND flash memory performance, reliability, and power consumption through several tuning knobs such as the flash programming algorithm and the flash error correcting code. However, to successfully exploit these degrees of freedom, it is mandatory to clearly understand the effect that the combined tuning of these parameters has on the full NVM subsystem. This article performs a comprehensive quantitative analysis of the benefits provided by the runtime reconfigurability of an MLC NAND flash controller through the combined effect of an adaptable memory programming circuitry coupled with runtime adaptation of the ECC correction capability. The full NVM subsystem is taken into account, starting from a characterization of the low-level circuitry to the effect of the adaptation on a wide set of realistic benchmarks in order to provide readers a clear view of the benefit this combined adaptation may provide at the system level. Davide Bertozzi, Stefano Di Carlo, Salvatore Galfano, Marco Indaco, Piero Olivo, Paolo Prinetto, Cristian Zambelli |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2014 | SSDExplorer: A virtual platform for fine-grained design space exploration of Solid State DrivesabstractSolid State Drives (SSDs) are gaining particular momentum in various frameworks such as multimedia, large data centers and cloud environments. Unfortunately, efficient CAD tools for SSD design space exploration able to assess the optimization of the device microarchitecture w.r.t. the target performance are still missing. This paper tries to close this gap by proposing SSDExplorer, a tool for fine-grained and fast design space exploration of SSD devices. SSDExplorer provides unprecedented insights into the architecture behavior and subcomponent interaction efficiency, while avoiding the need for the actual implementation of an FTL or of key hardware components. This is achieved by the introduction of suitable abstractions of the different components. This is confirmed by the thorough validation of SSDExplorer against a commercial SSD device. Lorenzo Zuolo, Cristian Zambelli, Rino Micheloni, Salvatore Galfano, Marco Indaco, Stefano Di Carlo, Paolo Prinetto, Piero Olivo, Davide Bertozzi |
DATE | 2 |
| 2014 | FLARES: An Aging Aware Algorithm to Autonomously Adapt the Error Correction Capability in NAND Flash MemoriesabstractWith the advent of solid-state storage systems, NAND flash memories are becoming a key storage technology. However, they suffer from serious reliability and endurance issues during the operating lifetime that can be handled by the use of appropriate error correction codes (ECCs) in order to reconstruct the information when needed. Adaptable ECCs may provide the flexibility to avoid worst-case reliability design, thus leading to improved performance. However, a way to control such adaptable ECCs' strength is required. This article proposes FLARES, an algorithm able to adapt the ECC correction capability of each page of a flash based on a flash RBER prediction model and on a measurement of the number of errors detected in a given time window. FLARES has been fully implemented within the YAFFS 2 filesystem under the Linux operating system. This allowed us to perform an extensive set of simulations on a set of standard benchmarks that highlighted the benefit of FLARES on the overall storage subsystem performances. Stefano Di Carlo, Salvatore Galfano, Marco Indaco, Paolo Prinetto, Davide Bertozzi, Piero Olivo, Cristian Zambelli |
ACM Trans. Archit. Code Optim. | 7 |
| 2012 | A cross-layer approach for new reliability-performance trade-offs in MLC NAND flash memoriesabstractIn spite of the mature cell structure, the memory controller architecture of Multi-level cell (MLC) NAND Flash memories is evolving fast in an attempt to improve the uncorrected/miscorrected bit error rate (UBER) and to provide a more flexible usage model where the performance-reliability trade-off point can be adjusted at runtime. However, optimization techniques in the memory controller architecture cannot avoid a strict trade-off between UBER and read throughput. In this paper, we show that co-optimizing ECC architecture configuration in the memory controller with program algorithm selection at the technology layer, a more flexible memory sub-system arises, which is capable of unprecedented trade-offs points between performance and reliability. Cristian Zambelli, Marco Indaco, Michele Fabiano, Stefano Di Carlo, Paolo Prinetto, Piero Olivo, Davide Bertozzi |
DATE | 1 |