VLDB 2026 Research / reviewers in the wild / expert
Mariagrazia Graziano
dblp:g/MariagraziaGraziano
· DBLP profile ↗
40ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-8721-9990ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Quantum Circuit Execution Success Estimation via Graph Neural Network-Based Prediction
Antonio Tudisco, Deborah Volpe, Mariagrazia Graziano, Giovanna Turvani |
RC | 3 |
| 2025 | Mage: a Decoupled Access-Execute CGRA tailored for Static Control ApplicationsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) have been thoroughly explored as a promising solution for accelerating compute-intensive applications, offering a balance between flexibility and energy efficiency. Recently, CGRA designs have tried to handle arbitrary complex code constructs, often resulting in increased architectural complexity and inefficient use of Processing Elements (PEs), in particular for Address Generation Instructions (AGIs).This paper introduces Mage, a Decoupled Access-Execute (DAE) CGRA specifically optimised for Static Control Programs (SCPs), which are well-suited for DAE-based acceleration. By leveraging an SCP-tailored Address Generation Unit for affine access patterns computation, Mage maximises PE utilisation for data processing. Compared to other State-of-the-Art DAE CGRAs, Mage reduces area occupation by up to 5.7x while ensuring high area efficiency, reaching 5701.7 MOPs/mm2. Alessio Naclerio, Fabrizio Riente, Giovanna Turvani, Marco Vacca, Maurizio Zamboni, Mariagrazia Graziano |
ISCAS | 6 |
| 2025 | Improving the exploitability of Simulated Adiabatic Bifurcation through a flexible and open-source digital architectureabstractCombinatorial Optimization (CO) problems exhibit exponential complexity, constraining classical computers from providing fast and satisfactory outcomes. Quantum Computers (QCs) can effectively find optimal or near-optimal solutions by exploring the solutions space of a problem encoded in a qubits system, exploiting principles of quantum mechanics. However, non-idealities and high costs limit their availability. These can be overcome by emulating QCs on cheaper and more accessible classical computing platforms, like Field-Programmable Gate Arrays (FPGAs). This article presents a digital architecture, implementing the Ising-compatible Simulated Adiabatic Bifurcation algorithm. It mimics the quantum adiabatic evolution of a network of non-linear Kerr oscillators. The architecture, described in VHDL and targeting FPGAs, consists of processing elements for computing the Kerr oscillators’ evolution, a set of units considering their Ising-related interactions and an evolution variables update unit. The proposed approach includes a speedup-targeting approximation of the algorithm, a method for handling single-variable constraints, and a software model that allows architecture customization for specific problems. Tests were conducted using an Altera Cyclone V SoC with FPGA logic and the Nios II processor for interface purposes. The results demonstrate the functionality of the architecture and its scalability with the problem size, making it suitable for real-world applications. Deborah Volpe, Giovanni Amedeo Cirillo, Maurizio Zamboni, Mariagrazia Graziano, Giovanna Turvani |
ACM Trans. Quantum Comput. | 4 |
| 2024 | WIP: Building an Education Ecosystem for Next Generation Microelectronics Experts in Green and Circular Economy with Digitally-Supported Teaching Methods for Sustainable Chips and Applications (EU Project GreenChips-EDU)abstractThis work in progress innovative practice paper intends to report on the outline and the ongoing progress of the EU-project GreenChips-EDU, which has been started in October 2023, and intends to fundamentally redesign educational microelectronics programs especially but not limited to students and professionals. One of the major goals is the design of a new microelectronics master program to which six European universities are contributing. The contents of this program will be substantially enhanced with green electronics contents innovative teaching methods. Other work will be done in the field of a new MBA program, self-standing modules for professionals, and a new microelectronics bachelor designed by one university of applied sciences. Klaus Hofmann, Ferdinand Keil, David Riehl, Alicja Malgorzata Michalowska-Forsyth, Nikolaus Czepl, Sarah Woywod, Dominik Zupan, Mario R. Casu, Carlo Ricciardi, Massimo Violante, Mariagrazia Graziano, Yuri Ardesi, Fabrizio Mo, Dominik Berger, Sabine Sill, Volker Visotschnig, Panagiota Morfouli, Liliana Prejbeanu, Katell Morin-Allory, Cyrille Chavet, Davide Bucci, Skandar Basrour, Jean-Christophe Crebier, Nhu-Huan Nguyen, Ernesto Quisbert-Trujillo, Christian Defélix, Isabelle Corbett-Etchevers, Johannes Sturm, Jens Peter Konrath, Ulla Birnbacher, Thomas Klinger, Wolfgang Werth, Jorge Fernandes, Marcelino B. Santos, Antonio Rubio 0001, Alba Pagès-Zamora, Jordi Salazar, Beatriz Otero, J. Manuel Moreno, X. Aragones, Israel Martin, Aleix Sole, Dunja Suttnig, Julia Calabro, Floriberto Lima, Eric Jouseau, François Cerisier, Cristian Rivier, Sepp Eisenriegler, Harald Reichl, Miroslav Macan, Dubravko Kruselj, Mladen Puskaric, Mirjana Tatalovic, Vinko Zelenicic, Bernd Deutschmann |
FIE | 11 |
| 2024 | DETECTive: Machine Learning-driven Automatic Test Pattern Prediction for Faults in Digital CircuitsabstractDue to the continuous technology scaling and the ever-increasing complexity and size of the hardware designs, manufacturing defects have become a key obstacle in meeting end-user demand. Despite decades of research, traditional test-generation techniques often struggle to scale to massive and complex designs. Such scalability issues stem from the numerous backtracking the traditional test generation techniques perform before converging to a test pattern. In this work, we present DETECTive that leverages deep learning on graphs to learn fault characteristics and predict test pattern(s) to expose faults without requiring backtracking. DETECTive is trained on small circuits, and its learned knowledge is transferable to predict test patterns for circuits that contain up to 29 × more gates than the training circuits. Since DETECTive avoids backtracking completely, it can predict test patterns up to 15 × faster than academic tools and up to 2 × faster than commercial tools. DETECTive achieves up to 100% pattern accuracy on synthetic designs and up to 95% test pattern accuracy on realistic designs. To our knowledge, DETECTive is the first to leverage deep learning to predict test patterns for digital hardware designs that can complement the traditional test generation techniques for faster design closure. Vincenzo Petrolo, Sourav Medya, Mariagrazia Graziano, Debjit Pal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Taming Molecular Field-Coupling for Nanocomputing DesignabstractMolecular Field-Coupling Nanocomputing (FCN) is one of the most promising technologies for overcoming Complementary Metal Oxide Semiconductor (CMOS) scaling issues. It encodes the information in the charge distribution of nanometric molecules and propagates it through local electrostatic intermolecular interaction. This technology promises very high speed at ambient temperatures with minimal power dissipation. The main research focus on molecular FCN is currently either on single-molecule low-level analysis or circuit design based on naïve assumptions. We aim to fill this gap, assessing the potential and feasibility of FCN. We present a bottom-up analysis and design framework that starts from the physical characterization of molecular and technological parameters and enables physical-aware FCN designs. The framework explicitly considers molecular physics, allowing the designer to tame the molecular interaction to ensure the computational capabilities of the final device. The framework permits studying possible physical effects that create cross-implications and correlations among physical and system-level layers considering possible behavior variability. We characterize and verify molecular propagation in increasingly structured layouts to design complex arithmetic circuits. The results highlight molecular FCN advantages, especially in area occupation, and provide valuable quantitative feedback to designers and technologists to support the assessment of molecular FCN and the realization of an eventual prototype. Yuri Ardesi, Umberto Garlando, Fabrizio Riente, Giuliana Beretta, Gianluca Piccinini, Mariagrazia Graziano |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2022 | Towards Compact Modeling of Noisy Quantum Computers: A Molecular-Spin-Qubit Case of StudyabstractClassical simulation of Noisy Intermediate Scale Quantum computers is a crucial task for testing the expected performance of real hardware. The standard approach, based on solving Schrödinger and Lindblad equations, is demanding when scaling the number of qubits in terms of both execution time and memory. In this article, attempts in defining compact models for the simulation of quantum hardware are proposed, ensuring results close to those obtained with standard formalism. Molecular Nuclear Magnetic Resonance quantum hardware is the target technology, where three non-ideality phenomena—common to other quantum technologies—are taken into account: decoherence, off-resonance qubit evolution, and undesired qubit-qubit residual interaction. A model for each non-ideality phenomenon is embedded into a MATLAB simulation infrastructure of noisy quantum computers. The accuracy of the models is tested on a benchmark of quantum circuits, in the expected operating ranges of quantum hardware. The corresponding outcomes are compared with those obtained via numeric integration of the Schrödinger equation and the Qiskit’s QASMSimulator. The achieved results give evidence that this work is a step forward towards the definition of compact models able to provide fast results close to those obtained with the traditional physical simulation strategies, thus paving the way for their integration into a classical simulator of quantum computers. Mario Simoni, Giovanni Amedeo Cirillo, Giovanna Turvani, Mariagrazia Graziano, Maurizio Zamboni |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | Hybrid-SIMD: A Modular and Reconfigurable Approach to Beyond von Neumann ComputingabstractThe increasing complexity of real-life applications demands constant improvements of microprocessor systems. One of the most frequently adopted microprocessor design scheme is the von Neumann architecture. Central Processing Unit (CPU performs computations and communicates with memory in a constant exchange of information. This unceasing motion of data between these two components became a significant performance bottleneck. A lot of power, energy, and computational time are wasted in this communication. With Beyond von Neumann Computing (BvNC paradigms, calculations are performed inside or very close to a memory array. BvNC approaches are proposed in the literature, mainly based on modifications of existing memories, enabling simple computations. Others exploit emerging technologies to both store and compute data, using analog operations. In this work we follow a different approach, where computational units are placed close to memory cells, improving versatility and performance. We propose a Hybrid-SIMD architecture made of memory and computing elements in an interleaved structure. Hybrid-SIMD can be used both as a low density memory and as SIMD accelerator. We insert our design in a classical von Neumann system based on a RISC-V processor, and we estimate its impact, demonstrating its capability to improve speed reducing at the same time energy consumption. Andrea Coluccio, Umberto Casale, Angela Guastamacchia, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano |
IEEE Trans. Computers | 8 |
| 2021 | FUNCODE: Effective Device-to-System Analysis of Field-Coupled Nanocomputing Circuit DesignsabstractMany beyond-CMOS technologies, based on different switching mechanisms, are arising. Field-coupled technologies are the most promising as they can guarantee an extremely low-power consumption and combine logic and memory into the same device. However, circuit-level explorations, like layout verification and analysis of the circuit performance, considering the constraints of the target technology, cannot be done using existing tools. Here, we propose a methodology to take on this challenge. We present function and connection detection (FUNCODE), an algorithm that can detect element connections, functions, and errors of custom layouts and generate its corresponding very high-speed integrated circuits hardware description language netlist. It is proposed for in-plane and perpendicular nanomagnetic logic as a case study. FUNCODE netlists, which take into account the physical behavior of the technology, were verified using circuits with increasing complexity, from 6 up to 1400 gates with a number of layout elements varying from 200 to 2.3e6. Umberto Garlando, Fabrizio Riente, Mariagrazia Graziano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | SCERPA Simulation of Clocked Molecular Field-Coupling NanocomputingabstractAmong all the possible technologies proposed for post-CMOS computing, molecular field-coupled nanocomputing (FCN) is one of the most promising technologies. The information propagation relies on electrostatic interactions among single molecules, overcoming the need for electron transport, significantly reducing energy dissipation. The expected working frequency is very high, and high throughput may be achieved by introducing an efficient pipeline of information propagation. The pipeline could be realized by adding an external clock signal that controls the propagation of data and makes the transmission adiabatic. In this article, we extend the Self-Consistent Electrostatic Potential Algorithm (SCERPA), previously introduced to analyze molecular circuits with a uniform clock field, to clocked molecular devices. The single-molecule is analyzed by ab initio calculations and modeled as an electronic device. Several clocked devices have been partitioned into clock zones and analyzed: the binary wire, the bus, the inverter, and the majority voter. The proposed modification of SCERPA enables linking the functional behavior of the clocked devices to molecular physics, becoming a possible tool for the eventual physical design verification of emerging FCN devices. The algorithm provides some first quantitative results that highlight the clocked propagation characteristics and provide significant feedback for the future implementation of molecular FCN circuits. Yuri Ardesi, Giovanna Turvani, Mariagrazia Graziano, Gianluca Piccinini |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | SCERPA: A Self-Consistent Algorithm for the Evaluation of the Information Propagation in Molecular Field-Coupled NanocomputingabstractAmong the emerging technologies that are intended to outperform the current CMOS technology, the field-coupled nanocomputing (FCN) paradigm is one of the most promising. The molecular quantum-dot cellular automata (MQCA) has been proposed as possible FCN implementation for the expected very high device density and possible room temperature operations. The digital computation is performed via electrostatic interactions among nearby molecular cells, without the need for charge transport, extremely reducing the power dissipation. Due to the lack of mature analysis and design methods, especially from an electronics standpoint, few attempts have been made to study the behavior of logic circuits based on real molecules, and this reduces the design capability. In this article, we propose a novel algorithm, named self-consistent electrostatic potential algorithm (SCERPA), dedicated to the analysis of molecular FCN circuits. The algorithm evaluates the interaction among all molecules in the system using an iterative procedure. It exploits two optimizations modes named Interaction Radius and Active Region which reduce the computational cost of the evaluation, enabling SCERPA to support the simulation of complex molecular FCN circuits and to characterize consequentially the technology potentials. The proposed algorithm fulfills the need for modeling the molecular structures as electronic devices and provides important quantitative results to analyze the information propagation, motivating and supporting further research regarding molecular FCN circuits and eventual prototype fabrication. Yuri Ardesi, Giovanna Turvani, Gianluca Piccinini, Mariagrazia Graziano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Exploiting the Logic-In-Memory paradigm for speeding-up data-intensive algorithms
Mario Cofano, Marco Vacca, Giulia Santoro, Giovanni Causapruno, Giovanna Turvani, Mariagrazia Graziano |
Integr. | 6 |
| 2018 | Architectural exploration of perpendicular Nano Magnetic Logic based circuits
Umberto Garlando, Fabrizio Riente, Giovanna Turvani, A. Ferrara, Giulia Santoro, Marco Vacca, Mariagrazia Graziano |
Integr. | 7 |
| 2018 | Exploring N3ASIC technology for microwave imaging architectures
Fabrizio Riente, Marco Vacca, Mariagrazia Graziano |
Integr. | 4 |
| 2017 | ToPoliNano: A CAD Tool for Nano Magnetic LogicabstractIn the post-CMOS scenario, field coupled nanotechnologies represent an innovative and interesting new direction for electronic nanocomputing. Among these technologies, nanomagnet logic (NML) makes it possible to finally embed logic and memory in the same device. To fully analyze the potential of NML circuits, design tools that mimic the CMOS design-flow should be used for circuit design. We present, in this paper, the latest and improved version of Torino Politecnico Nanotechnology (ToPoliNano), our design and simulation framework for field coupled nanotechnologies. ToPoliNano emulates the top-down design process of CMOS technology. Circuits are described with a VHSIC hardware description language netlist and layout is then automatically generated considering in-plane NML (iNML) technology. The resulting circuits can be simulated and performance can be analyzed. In this paper, we describe several enhancements to the tool itself, like a circuit editor for custom design of field coupled nanodevices, improved algorithms for netlist optimization and new algorithms for the place and route of iNML circuits. We have validated and analyzed the tool by using extensive metrics, both by using standard circuits and ISCAS'85 benchmarks. This contribution highlights the improvements of ToPoliNano, which is now a innovative and complete tool for the development of iNML technology. Fabrizio Riente, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Domain Wall Interconnections for NMLabstractNanomagnet logic (NML) is one of the most novel solutions studied as complementary technology to CMOS transistors. Information propagation involves only a change in spin orientation, no charge movement is present. Since the basic element is a nanomagnet, NML circuits have no stand-by power consumption and the ability to mix logic and memory in the same device. While CMOS is a multilayer technology, until now NML is confined to one single physical layer. The consequence is that circuit area grows exponentially due to interconnections overhead. In this paper, we present an innovative solution that drastically reduces the area wasted for interconnection wires relying on the properties of domain walls (DWs). We mix DWs and NML technologies in a unique DW logic (DWL) solution that exploits the advantages of both technologies. The proposed solution is technologically compatible with up-to-date fabrication processes. All the results here presented for the NML logic blocks and the DWs interconnections and their combination are obtained through rigorous micromagnetic simulations. Moreover, we implemented as a case study an high performance adder (Pentium 4 adder) and evaluated its features with increasing parallelism and compared with the simple NML implementation in order to explore the potential of DWL technology at circuit and architectural level. The reduction in circuit area corresponds to a notable reduction in both the latency and power consumption. The improvements in NML technology are shown by both the remarkable performance improvement and new possibilities offered by this novel solution. Fabrizio Cairo, Marco Vacca, Giovanna Turvani, Maurizio Zamboni, Mariagrazia Graziano |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Reconfigurable Systolic Array: From Architecture to Physical Design for NMLabstractNanoMagnet logic (NML) is among the emerging technologies that might replace CMOS in the next decades. According to its physical characteristics, to better exploit the potential of this technology-and of other similar ones-the use of parallel architectures with regular layout that avoid long interconnection signals is advised. Systolic arrays (SAs) are among these architectures, being composed of a grid of equal processing elements that are locally interconnected. However, they are usually implemented to execute only a small set of algorithms, and for this reason, throughout the years, they have not been an appealing solution for CMOS. To seriously analyze the potentials of NML, complex architectures must be conceived, and their physical implementation explored considering realistic technological constraints. With the increasing complexity of NML circuits, two issues, then, are noticed: 1) the need for a regular structure arises, that at the same time helps to reduce the intrinsic pipelining nature of NML and can be configured to be used for several applications without developing a dedicated design for each algorithm and 2) the capability to synthesize, place and route NML circuits is fundamental to demonstrate the feasibility of the architecture in two important conditions: efficiently managing the complexity of the design and sticking to the characteristics that are technologically feasible at the time of writing. In this paper, we address these issues presenting a new reconfigurable SA that can be programmed to execute different algorithms, and we provide two examples to show its working principle. Moreover, the array is synthesized and simulated with the aid of the first real tool for nanotechnology circuits that we have conceived, Torino Politecnico Nanotechnology tool. The joint contribution at both the architectural and physical design levels gives a relevant step forward to the state of the art in the demonstration of this emerging technology potential. Giovanni Causapruno, Fabrizio Riente, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2015 | Logic-in-Memory architecture made realabstractThe current trend for intensive computational architectures is to adopt massive parallelism, with several concurrent tasks performed simultaneously, as done for example in GPUs. This approach has many advantages, such as the reduced design time given by circuit replication and an increasing in computational speed without the need of higher frequency. It has however evidenced an important bottleneck in data exchange between memory and processor. We envisage a revolutionary path for the future relation between memory and logic in parallel processors, where a new type of architecture exploits the principle of caching to the limit. Our Logic-in-Memory (LIM) architecture mixes logic and memory in the same device, removing the bottleneck of other existing parallel solutions. The architecture we propose, here in its preliminary version, has an array organization and each element in the array is based on three blocks: a logic unit for processing, a smart memory block and a routing structure for inter block communication. In this article we show the benefits of this approach with an application example in the image processing field. We can achieve a 4X computational time reduction for an image processing algorithm (Summed Area Table) with respect to the best architecture present in the literature, even with a preliminary and not optimized version. Besides the adoption of massive parallelism to increase performance, new technologies to open the post-CMOS era are explored. Among them NanoMagnet Logic (NML) is particularly interesting for its ability to mix logic and memory in the same device. We present here the preliminary results of the NML implementation of the LIM architecture. We thus demonstrate that it is not only a good solution for a standard CMOS technology but can also exploit the potential of an emerging technology as NML. D. Pala, Giovanni Causapruno, Marco Vacca, Fabrizio Riente, Giovanna Turvani, Mariagrazia Graziano, Maurizio Zamboni |
ISCAS | 6 |
| 2015 | Process Variability and Electrostatic Analysis of Molecular QCAabstractMolecular quantum-dot cellular automata (mQCA) is an emerging paradigm for nanoscale computation. Its revolutionary features are the expected operating frequencies (THz), the high device densities, the noncryogenic working temperature, and, above all, the limited power densities. The main drawback of this technology is a consequence of one of its very main advantages, that is, the extremely small size of a single molecule. Device prototyping and the fabrication of a simple circuit are limited by lack of control in the technological process [Pulimeno et al. 2013a]. Moreover, high defectivity might strongly impact the correct behavior of mQCA devices. Another challenging point is the lack of a solid method for analyzing and simulating mQCA behavior and performance, either in ideal or defective conditions. Our contribution in this article is threefold: (i) We identify a methodology based on both ab-initio simulations and post-processing of data for analyzing an mQCA system adopting an electronic point of view (we baptized this method as “MoSQuiTo”); (ii) we assess the performance of an mQCA device (in this case, a bis- ferrocene molecule) working in nonideal conditions, using as a reference the information on fabrication-critical issues and on the possible defects that we are obtaining while conducting our own ongoing experiments on mQCA: (iii) we determine and assess the electrostatic energy stored in a bis-ferrocene molecule both in an oxidized and reduced form. Results presented here consist of quantitative information for an mQCA device working in manifold driving conditions and subjected to defects. This information is given in terms of: (a) output voltage; (b) safe operating area (SOA); (c) electrostatic energy; and (d) relation between SOA and energy, that is, possible energy reduction subject to reliability and functionality constraints. The whole analysis is a first fundamental step toward the study of a complex mQCA circuit. It gives important suggestions on possible improvements of the technological processes. Moreover, it starts an interesting assessment on the energy of an mQCA, one of the most promising features of this technology. Mariagrazia Graziano, Azzurra Pulimeno, Massimo Ruo Roch, Gianluca Piccinini |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2015 | Interleaving in Systolic-Arrays: A Throughput BreakthroughabstractIn past years the most common way to improve computers performance was to increase the clock frequency. In recent years this approach suffered the limits of technology scaling, therefore computers architectures are shifting toward the direction of parallel computing to further improve circuits performance. Not only GPU based architectures are spreading in consideration, but also Systolic Arrays are particularly suited for certain classes of algorithms. An important point in favor of Systolic Arrays is that, due to the regularity of their circuit layout, they are appealing when applied to many emerging and very promising technologies, like Quantum-dot Cellular Automata and nanoarrays based on Silicon NanoWire or on Carbon nanotube Field Effect Transistors. In this work we present a systematic method to improve Systolic Arrays performance exploiting Pipelining and Input Data Interleaving. We tackle the problem from a theoretical point of view first, and then we apply it to both CMOS technology and emerging technologies. On CMOS we demonstrate that it is possible to vastly improve the overall throughput of the circuit. By applying this technique to emerging technologies we show that it is possible to overcome some of their limitations greatly improving the throughput, making a considerable step forward toward the post-CMOS era. Giovanni Causapruno, Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni |
IEEE Trans. Computers | 3 |
| 2015 | Protein Alignment Systolic Array Throughput OptimizationabstractProtein comparison is gaining importance year after year since it has been demonstrated that biologists can find correlation between different species, or genetic mutations that can lead to cancer and genetic diseases. Protein sequence alignment is the most computational intensive task when performing protein comparison. To speed-up alignment, dedicated processors that can perform different computations in parallel have been designed. Among them, the best performance has been achieved using systolic arrays (SAs). However, when the processing elements of the SA have an internal loop, performance could be highly reduced. In this paper, we present an architectural strategy to address this problem applying pipeline interleaving; this strategy is applied to an SA for Smith Waterman algorithm that we designed. Results encourage the adoption of pipeline interleaving for parallel circuits with loop-based processing elements. We demonstrate that important benefits in terms of higher operating frequency can be derived without so relevant costs as increased complexity, area, and power required. Giovanni Causapruno, Gianvito Urgese, Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Feedbacks in QCA: A Quantitative ApproachabstractIn the post-CMOS scenario a primary role is played by the quantum-dot cellular automata (QCA) technology. Irrespective of the specific implementation principle (e.g., either molecular, or magnetic or semiconductive in the current scenario) the intrinsic deep-level pipelined behavior is the dominant issue. It has important consequences on circuit design and performance, especially in the presence of feedbacks in sequential circuits. Though partially already addressed in literature, these consequences still must be fully understood and solutions thoroughly approached to allow this technology any further advancement. This paper conducts an exhaustive analysis of the effects and the consequences derived by the presence of loops in QCA circuits. For each problem arisen, a solution is presented. The analysis is performed using as a test architecture, a complex systolic array circuit for biosequences analysis (Smith-Waterman algorithm), which represents one of the most promising application for QCA technology. The circuit is based on nanomagnetic logic as QCA implementation, is designed down to the layout level considering technological constraints and experimentally validated structures, counts up to approximately 2.3 milion nanomagnets, and is described and simulated with HDL language using as a testbench realistic protein alignment sequences. The results here presented constitute a fundamental advancement in the emerging technologies field since: 1) they are based on a quantitative approach relying on a realistic and complex circuit involving a large variety of QCA blocks; 2) they strictly are reckoned starting from current technological limits without relying on unrealistic assumptions; 3) they provide general rules to design complex sequential circuits with intrinsically pipelined technologies, like QCA; and 4) they prove with a real application benchmark how to maximize the circuits performance. Marco Vacca, Juanchi Wang, Mariagrazia Graziano, Massimo Ruo Roch, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Fault tolerant nanoarray circuits: Automatic design and verificationabstractWe automatically maximize fault-tolerance in nanoarrays based on silicon nanowires and Gate-All-Around transistors optimizing their topology vs. several distributions of faults inherited by technology. We added a Monte Carlo engine in our nanoarchitecture design tool ToPoliNano and verified the effectiveness of the fault-tolerance algorithm over several circuits and faults distributions. Pasquale Ranone, Giovanna Turvani, Fabrizio Riente, Mariagrazia Graziano, Massimo Ruo Roch, Maurizio Zamboni |
VTS | 4 |
| 2014 | Simulation and design of an UWB imaging system for breast cancer detection
Xiaolu Guo, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
Integr. | 3 |
| 2014 | Nanoarray architectures multilevel simulationabstractDensity and regularity are deemed as the major advantages of nanoarray architectures based on nanowires. Literature demonstrated that proper reliability analyzes must be performed and solutions have to be devised to improve nanoarrays yield. Their complexity and high-fault probability claim for specific design automation tools able to explore circuit solutions, performance and fault-tolerant approaches. We envision a simulator conceived to carry on characterizations in terms of logic behavior, defect-induced output error rate assessment, switching activity, power and timing performance. Though already existing for traditional technology, a simulator based on specific technological and topological tiled nanoarray descriptions, and conceived to join both device and architecture levels, has never been attempted at the degree of accuracy we present. Our contribution is twofold. First, marking a difference with respect to the state of the art, we developed an algorithm based on an event-driven engine which works at switch level and is not simply built on top of cost functions evaluations. The straightforward advantage is the possibility to follow the evolution of dynamic control sequences throughout all the inner components of the nanoarray, and, as a consequence, to obtain circuit level characterization as a projection of the real internal parameters. Second, we added to our simulator the capability to inject faults with specific statistical distributions associated to the nanoarray topology. Here we extract output error rates and yield for one of the possible nanoarray structures proposed in literature, the NASIC. Results specificity and accuracy demonstrate the simulator trustworthiness, its effectiveness for extensive nanoarrays characterization and its suitability as a foundation for both higher architectural and lower device simulation levels. The aim of this work, then, is to provide insights into the intertwined relation between actual technology and circuit design for these emerging fabrics, and, as a consequence, to clarify how defects and variability affect circuits and systems performance. Stefano Frache, Mariagrazia Graziano, Maurizio Zamboni |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2014 | Enabling design and simulation of massive parallel nanoarchitectures
Stefano Frache, Diego Chiabrando, Mariagrazia Graziano, Marco Vacca, Luca Boarino, Maurizio Zamboni |
J. Parallel Distributed Comput. | 3 |
| 2014 | UWB microwave imaging for breast cancer detection: Many-core, GPU, or FPGA?abstractAn UWB microwave imaging system for breast cancer detection consists of antennas, transceivers, and a high-performance embedded system for elaborating the received signals and reconstructing breast images. In this article we focus on this embedded system. To accelerate the image reconstruction, the Beamforming phase has to be implemented in a parallel fashion. We assess its implementation in three currently available high-end platforms based on a multicore CPU, a GPU, and an FPGA, respectively. We then project the results applying technology scaling rules to future many-core CPUs, many-thread GPUs, and advanced FPGAs. We consider an optimistic case in which available resources increase according to Moore's law only, and a pessimistic case in which only a fraction of those resources are available due to a limited power budget. In both scenarios, an implementation that includes a high-end FPGA outperforms the other alternatives. Since the number of effectively usable cores in future many-cores will be power-limited, and there is a trend toward the integration of power-efficient accelerators, we conjecture that a chip consisting of a many-core section and a reconfigurable logic section will be the perfect platform for this application. Mario R. Casu, Francesco Colonna, Marco Crepaldi, Danilo Demarchi, Mariagrazia Graziano, Maurizio Zamboni |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | A 130nm PMOS drain-degenerated ratioless level-shifter for near-threshold designsabstractWe present a modified type-I level-up shifter with improved Process-Voltage-Temperature (PVT) robustness, propagation delay and energy consumption. Compared to a standard cross-coupled level-shifter, the circuit comprises a couple of long channel parallel P and N transistors to implement larger PMOS on-resistance maintaining unvaried upstream logic fan-out. Simulation results show significant robustness increase with respect to a standard topology maintaining low NMOS-to-PMOS sizing. Switching energy consumption is reduced from ~ 10pJ to 200fJ and propagation delay from ~ 240ns to 1ns. With Monte Carlo process variation simulations we have verified a reduction in output delay sensitivity from 209ns to 333ps while with transient noise simulation jitter is reduced from 3.5ns to 36ps. Operating ranges are wider in the proposed circuit, while sensitivity to temperature is comparable for high values. A prototype of this drain-degenerated logic-translator has been fabricated in a 130nm CMOS technology and evaluated with measurements. Marco Crepaldi, Paolo Motto Ros, Mariagrazia Graziano, Danilo Demarchi |
ETFA | 3 |
| 2013 | A Hardware Viewpoint on Biosequence Analysis: What's Next?abstractBiosequence alignment recently received an increasing support from both commodity and dedicated hardware platforms. Processing capabilities are constantly rising, but still not satisfying the limitless requirements of this application. We give an insight on the contribution to this need that can possibly be expected from emerging technology devices and architectures, focusing as an example on nanofabrics based on silicon nanowires. By varying a few parameters we explore the solution space, and demonstrate with proper figures of merit how this family of beyond CMOS structures could be considered as the effective disruptive technology for biosequence analysis applications. Mariagrazia Graziano, Stefano Frache, Maurizio Zamboni |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2013 | Nanomagnetic Logic Microprocessor: Hierarchical Power ModelabstractThe interest in emerging nanotechnologies has been recently focused on nanomagnetic logic (NML), which has unique appealing features. NML circuits have very low power consumption and, because of their magnetic nature, maintain the information safely stored even without power supply. The nature of these circuits is much different from that of CMOS circuits. As a consequence, to better understand NML logic, complex circuits and not only simple gates must be designed. This constraint calls for a new design and simulation methodology. It should efficiently encompass manifold properties: 1) being based on commonly used hardware description language (HDL) in order to easily manage complexity and hierarchy; 2) maintaining a clear link with physical characteristics; and 3) modeling performance aspects such as speed and power, together with logic behavior. In this paper, we present a very-high-speed integrated circuits HDL (VHDL) behavioral model for NML circuits, which allows the evaluation of not only the logic behavior but also its power dissipation. It is based on a technological solution called “snake-clock.” We demonstrate this model using a case study which offers the right variety of internal substructures to test the method: a 4-bit microprocessor designed using asynchronous logic. The model enables a hierarchical bottom-up evaluation of the processor logic behavior, area, and power dissipation, which we evaluate using a benchmark division algorithm. The results highlight the flexibility and the efficiency of this model, as well as the remarkable improvements that it brings to the analysis of NML circuits. Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | A Multistandard Digital HD/SD Audio Multiplexer With Modular Ancillary Packet SubstitutionabstractThis article presents a multistandard digital high-definition (HD)/standard-definition (SD)-serial digital interface (SDI) audio multiplexer (embedder) capable of inserting and replacing existing ancillary Audio Engineering Society-European Broadcasting Union packets. The insertion of audio packets in SD-SDI video is achieved with an ad hoc modular unit that shifts the existing audio packets in the video frame. This block enables the use of the entire multiplexer system in cascaded configuration along the same video line, a fundamental requirement in professional television studios. The embedder, implemented on a Virtex-5 field programmable gate array, accepts 48 kHz AES3 synchronous audio, and Society of Motion Picture and Television Engineers (SMPTE) 259 M-C and SMPTE 295 M full HD video. It occupies a small area, and installs on a system board inclusive of dedicated signal conditioning integrated circuit. The full system is tested for robust operation in audio or video production environments and for compatibility with older equipments. Marco Crepaldi, Daniele Franchino, Modesto Scavarda, Mariagrazia Graziano |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | UDSM Trends Comparison: From Technology Roadmap to UltraSparc Niagara2abstractThe increased leakage, yield inefficiency, process, power supply, and temperature variations have significant aftereffects on the performance of complex VLSI architectures especially if mapped on ultra deep sub micrometer (UDSM) technologies. In this paper we assess the technology trend based on three industrial technologies (90, 65, and 45 nm) using a state of the art processor as benchmark: The UltraSparc Niagara 2 from SUN Microsystem. We analyze frequency, dynamic, and static power and area after synthesis varying power supply voltage and temperature. We then compare these exhaustive analyses of system level performance as a function of technology to ITRS device level estimations. The results suggest that this prediction can be of help when addressing both the technological scaling and the variability scenario of the selected technology. We believe that correctly predicting specific values on performance variations when realistic conditions and technologies are changed could provide a valuable information for the architect. Our analysis advises the designer on the effective applicability of the ITRS trends to system performance, but also pinpoints that a reliable system level prediction should better take into account the design complexity. Azzurra Pulimeno, Mariagrazia Graziano, Gianluca Piccinini |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Asynchronous Solutions for Nanomagnetic Logic CircuitsabstractIn the years to come new solutions will be required to overcome the limitations of scaled CMOS technology. One approach is to adopt Nano-Magnetic Logic Circuits, highly appealing for their extremely reduced power consumption. Despite the interesting nature of this approach, many problems arise when this technology is considered for real designs. The wire is the most critical of these problems from the circuit implementation point of view. It works as a pipelined interconnection, and its delay in terms of clock cycles depends on its length. Serious complications arise at the design phase, both in terms of synthesis and of physical design. One possible solution is the use of a delay insensitive asynchronous logic, Null Convention Logic (NCL TM ). Nevertheless its use has many negative consequences in terms of area occupation and speed loss with respect to a Boolean version. In this article we analyze and compare different solutions: nanomagnetic circuits based on full NCL, mixed Boolean-NCL, and fully Boolean logic. We discuss the advantages of these logics, but also the issues they raise. In particular we analyze feedback signals, which, due to their intrinsic pipelined nature, cause errors that still have not found a solution in the literature. The innovative arrangement we propose solves most of the problems and thus soundly increases the knowledge of this technology. The analysis is performed using a VHDL behavioral model we developed and a microprocessor we designed based on this model, as a sound and realistic test bench. Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | A flexible UWB Transmitter for breast cancer detection imaging systemsabstractThis paper presents a flexible architecture for an integrated Ultra-Wideband (UWB) Transmitter capable of generating pulses suited for breast cancer detection imaging systems. A flexible design allows the generation of a large variety of UWB signals fully compatible with the ones used in real experiments in recent state-of-the-art. Flexibility and high degree of programmability of the mixed-signal system allow also to compensate for non-ideal effects of building blocks through a digital calibration. It is also shown how not only internal non-idealities are accounted for but also how channel and antenna responses can be compensated for through a digital pre-emphasis of UWB pulses. The circuit is designed on a 130 nm CMOS technology and simulated at transistor-level. Simulations showed 2% maximum NRMSE pulse error with respect to ideal Gaussian and Modulated and Modified Hermite Polynomial (MMHP) Matlab templates. Massimo Cutrupi, Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano |
DATE | 4 |
| 2010 | A flexible simulation methodology and tool for nanoarray-based architecturesabstractNanoscale arrays based on nanowires are expected to have a promising future thanks to their amazing density and regularity. Experiments demonstrated the feasibility of this technology and pointed out that accurate reliability analyses should be accomplished to assure proper yield requirements. Due to the complexity of these systems and the arising necessity of thorough fault analysis, design automation tools are mandatory in order to explore architectural solutions and fault tolerant approaches deriving information from reliable nanoarray characterisation. We present a simulator, never attempted at this level of detail, based on specific technological and topological tiled nanoarray descriptions, conceived to carry on characterisations in terms of logic behaviour, defect-induced error rate assessment, switching activity and other figures of merit like power and timing performance (not discussed in this paper). It is formulated in a flexible and modular way to assure the simulation of manifold advancing technological solutions, among which the winner has not been determined yet. Marking a difference with respect to the state of the art, the algorithm is based on an event-driven engine and not on cost functions evaluations. Thus even dynamic control sequences can be processed and their evolution followed throughout all the inner components of the array allowing to obtain system level characterization as a projection of the real internal parameters. In this paper we show results attained for one of the possible nanoarray structures proposed in literature, the NASIC: logic behaviour, defect error rates and switching activity for two types of function demonstrate the simulator trustworthiness, its effectiveness for extensive nanoarrays characterisation and its suitability as a foundation for both higher architectural and lower device simulation levels. Stefano Frache, Mariagrazia Graziano, Maurizio Zamboni |
ICCD | 2 |
| 2009 | A mixed-signal demodulator for a low-complexity IR-UWB receiver: Methodology, simulation and design
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
Integr. | 3 |
| 2008 | An Automotive CD-Player Electro-Mechanics Fault Simulation Using VHDL-AMS
Mariagrazia Graziano, Massimo Ruo Roch |
J. Electron. Test. | 1 |
| 2008 | Statistical power supply dynamic noise prediction in hierarchical power grid and package networks
Mariagrazia Graziano, Gianluca Piccinini |
Integr. | 1 |
| 2007 | An effective AMS top-down methodology applied to the design of a mixed-signal UWB system-on-chip
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
DATE | 3 |
| 2004 | An electromigration and thermal model of power wires for a priori high-level reliability predictionabstractIn this paper, a simple power-distribution electrothermal model including the interconnect self-heating is used together with a statistical model of average and rms currents of functional blocks and a high-level model of fanout distribution and interconnect wirelength. Following the 2001 SIA roadmap projections, we are able to predict a priori that the minimum width that satisfies the electromigration constraints does not scale like the minimum metal pitch in future technology nodes. As a consequence, the percentage of chip area covered by power lines is expected to increase at the expense of wiring resources unless proper countermeasures are taken. Some possible solutions are proposed in the paper. Mario R. Casu, Mariagrazia Graziano, Guido Masera, Gianluca Piccinini, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |