Maurizio Zamboni

dblp:74/1735 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 3 since 2021Software engineering, systems software and programming languages · 4Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Mage: a Decoupled Access-Execute CGRA tailored for Static Control Applications
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) have been thoroughly explored as a promising solution for accelerating compute-intensive applications, offering a balance between flexibility and energy efficiency. Recently, CGRA designs have tried to handle arbitrary complex code constructs, often resulting in increased architectural complexity and inefficient use of Processing Elements (PEs), in particular for Address Generation Instructions (AGIs).This paper introduces Mage, a Decoupled Access-Execute (DAE) CGRA specifically optimised for Static Control Programs (SCPs), which are well-suited for DAE-based acceleration. By leveraging an SCP-tailored Address Generation Unit for affine access patterns computation, Mage maximises PE utilisation for data processing. Compared to other State-of-the-Art DAE CGRAs, Mage reduces area occupation by up to 5.7x while ensuring high area efficiency, reaching 5701.7 MOPs/mm2.
Alessio Naclerio, Fabrizio Riente, Giovanna Turvani, Marco Vacca, Maurizio Zamboni, Mariagrazia Graziano
ISCAS5
2025 Improving the exploitability of Simulated Adiabatic Bifurcation through a flexible and open-source digital architecture
abstract
Combinatorial Optimization (CO) problems exhibit exponential complexity, constraining classical computers from providing fast and satisfactory outcomes. Quantum Computers (QCs) can effectively find optimal or near-optimal solutions by exploring the solutions space of a problem encoded in a qubits system, exploiting principles of quantum mechanics. However, non-idealities and high costs limit their availability. These can be overcome by emulating QCs on cheaper and more accessible classical computing platforms, like Field-Programmable Gate Arrays (FPGAs). This article presents a digital architecture, implementing the Ising-compatible Simulated Adiabatic Bifurcation algorithm. It mimics the quantum adiabatic evolution of a network of non-linear Kerr oscillators. The architecture, described in VHDL and targeting FPGAs, consists of processing elements for computing the Kerr oscillators’ evolution, a set of units considering their Ising-related interactions and an evolution variables update unit. The proposed approach includes a speedup-targeting approximation of the algorithm, a method for handling single-variable constraints, and a software model that allows architecture customization for specific problems. Tests were conducted using an Altera Cyclone V SoC with FPGA logic and the Nios II processor for interface purposes. The results demonstrate the functionality of the architecture and its scalability with the problem size, making it suitable for real-world applications.
Deborah Volpe, Giovanni Amedeo Cirillo, Maurizio Zamboni, Mariagrazia Graziano, Giovanna Turvani
ACM Trans. Quantum Comput.3
2022 Towards Compact Modeling of Noisy Quantum Computers: A Molecular-Spin-Qubit Case of Study
abstract
Classical simulation of Noisy Intermediate Scale Quantum computers is a crucial task for testing the expected performance of real hardware. The standard approach, based on solving Schrödinger and Lindblad equations, is demanding when scaling the number of qubits in terms of both execution time and memory. In this article, attempts in defining compact models for the simulation of quantum hardware are proposed, ensuring results close to those obtained with standard formalism. Molecular Nuclear Magnetic Resonance quantum hardware is the target technology, where three non-ideality phenomena—common to other quantum technologies—are taken into account: decoherence, off-resonance qubit evolution, and undesired qubit-qubit residual interaction. A model for each non-ideality phenomenon is embedded into a MATLAB simulation infrastructure of noisy quantum computers. The accuracy of the models is tested on a benchmark of quantum circuits, in the expected operating ranges of quantum hardware. The corresponding outcomes are compared with those obtained via numeric integration of the Schrödinger equation and the Qiskit’s QASMSimulator. The achieved results give evidence that this work is a step forward towards the definition of compact models able to provide fast results close to those obtained with the traditional physical simulation strategies, thus paving the way for their integration into a classical simulator of quantum computers.
Mario Simoni, Giovanni Amedeo Cirillo, Giovanna Turvani, Mariagrazia Graziano, Maurizio Zamboni
ACM J. Emerg. Technol. Comput. Syst.5
2022 Hybrid-SIMD: A Modular and Reconfigurable Approach to Beyond von Neumann Computing
abstract
The increasing complexity of real-life applications demands constant improvements of microprocessor systems. One of the most frequently adopted microprocessor design scheme is the von Neumann architecture. Central Processing Unit (CPU performs computations and communicates with memory in a constant exchange of information. This unceasing motion of data between these two components became a significant performance bottleneck. A lot of power, energy, and computational time are wasted in this communication. With Beyond von Neumann Computing (BvNC paradigms, calculations are performed inside or very close to a memory array. BvNC approaches are proposed in the literature, mainly based on modifications of existing memories, enabling simple computations. Others exploit emerging technologies to both store and compute data, using analog operations. In this work we follow a different approach, where computational units are placed close to memory cells, improving versatility and performance. We propose a Hybrid-SIMD architecture made of memory and computing elements in an interleaved structure. Hybrid-SIMD can be used both as a low density memory and as SIMD accelerator. We insert our design in a classical von Neumann system based on a RISC-V processor, and we estimate its impact, demonstrating its capability to improve speed reducing at the same time energy consumption.
Andrea Coluccio, Umberto Casale, Angela Guastamacchia, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano
IEEE Trans. Computers7
2017 ToPoliNano: A CAD Tool for Nano Magnetic Logic
abstract
In the post-CMOS scenario, field coupled nanotechnologies represent an innovative and interesting new direction for electronic nanocomputing. Among these technologies, nanomagnet logic (NML) makes it possible to finally embed logic and memory in the same device. To fully analyze the potential of NML circuits, design tools that mimic the CMOS design-flow should be used for circuit design. We present, in this paper, the latest and improved version of Torino Politecnico Nanotechnology (ToPoliNano), our design and simulation framework for field coupled nanotechnologies. ToPoliNano emulates the top-down design process of CMOS technology. Circuits are described with a VHSIC hardware description language netlist and layout is then automatically generated considering in-plane NML (iNML) technology. The resulting circuits can be simulated and performance can be analyzed. In this paper, we describe several enhancements to the tool itself, like a circuit editor for custom design of field coupled nanodevices, improved algorithms for netlist optimization and new algorithms for the place and route of iNML circuits. We have validated and analyzed the tool by using extensive metrics, both by using standard circuits and ISCAS'85 benchmarks. This contribution highlights the improvements of ToPoliNano, which is now a innovative and complete tool for the development of iNML technology.
Fabrizio Riente, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2017 Domain Wall Interconnections for NML
abstract
Nanomagnet logic (NML) is one of the most novel solutions studied as complementary technology to CMOS transistors. Information propagation involves only a change in spin orientation, no charge movement is present. Since the basic element is a nanomagnet, NML circuits have no stand-by power consumption and the ability to mix logic and memory in the same device. While CMOS is a multilayer technology, until now NML is confined to one single physical layer. The consequence is that circuit area grows exponentially due to interconnections overhead. In this paper, we present an innovative solution that drastically reduces the area wasted for interconnection wires relying on the properties of domain walls (DWs). We mix DWs and NML technologies in a unique DW logic (DWL) solution that exploits the advantages of both technologies. The proposed solution is technologically compatible with up-to-date fabrication processes. All the results here presented for the NML logic blocks and the DWs interconnections and their combination are obtained through rigorous micromagnetic simulations. Moreover, we implemented as a case study an high performance adder (Pentium 4 adder) and evaluated its features with increasing parallelism and compared with the simple NML implementation in order to explore the potential of DWL technology at circuit and architectural level. The reduction in circuit area corresponds to a notable reduction in both the latency and power consumption. The improvements in NML technology are shown by both the remarkable performance improvement and new possibilities offered by this novel solution.
Fabrizio Cairo, Marco Vacca, Giovanna Turvani, Maurizio Zamboni, Mariagrazia Graziano
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Reconfigurable Systolic Array: From Architecture to Physical Design for NML
abstract
NanoMagnet logic (NML) is among the emerging technologies that might replace CMOS in the next decades. According to its physical characteristics, to better exploit the potential of this technology-and of other similar ones-the use of parallel architectures with regular layout that avoid long interconnection signals is advised. Systolic arrays (SAs) are among these architectures, being composed of a grid of equal processing elements that are locally interconnected. However, they are usually implemented to execute only a small set of algorithms, and for this reason, throughout the years, they have not been an appealing solution for CMOS. To seriously analyze the potentials of NML, complex architectures must be conceived, and their physical implementation explored considering realistic technological constraints. With the increasing complexity of NML circuits, two issues, then, are noticed: 1) the need for a regular structure arises, that at the same time helps to reduce the intrinsic pipelining nature of NML and can be configured to be used for several applications without developing a dedicated design for each algorithm and 2) the capability to synthesize, place and route NML circuits is fundamental to demonstrate the feasibility of the architecture in two important conditions: efficiently managing the complexity of the design and sticking to the characteristics that are technologically feasible at the time of writing. In this paper, we address these issues presenting a new reconfigurable SA that can be programmed to execute different algorithms, and we provide two examples to show its working principle. Moreover, the array is synthesized and simulated with the aid of the first real tool for nanotechnology circuits that we have conceived, Torino Politecnico Nanotechnology tool. The joint contribution at both the architectural and physical design levels gives a relevant step forward to the state of the art in the demonstration of this emerging technology potential.
Giovanni Causapruno, Fabrizio Riente, Giovanna Turvani, Marco Vacca, Massimo Ruo Roch, Maurizio Zamboni, Mariagrazia Graziano
IEEE Trans. Very Large Scale Integr. Syst.6
2015 Logic-in-Memory architecture made real
abstract
The current trend for intensive computational architectures is to adopt massive parallelism, with several concurrent tasks performed simultaneously, as done for example in GPUs. This approach has many advantages, such as the reduced design time given by circuit replication and an increasing in computational speed without the need of higher frequency. It has however evidenced an important bottleneck in data exchange between memory and processor. We envisage a revolutionary path for the future relation between memory and logic in parallel processors, where a new type of architecture exploits the principle of caching to the limit. Our Logic-in-Memory (LIM) architecture mixes logic and memory in the same device, removing the bottleneck of other existing parallel solutions. The architecture we propose, here in its preliminary version, has an array organization and each element in the array is based on three blocks: a logic unit for processing, a smart memory block and a routing structure for inter block communication. In this article we show the benefits of this approach with an application example in the image processing field. We can achieve a 4X computational time reduction for an image processing algorithm (Summed Area Table) with respect to the best architecture present in the literature, even with a preliminary and not optimized version. Besides the adoption of massive parallelism to increase performance, new technologies to open the post-CMOS era are explored. Among them NanoMagnet Logic (NML) is particularly interesting for its ability to mix logic and memory in the same device. We present here the preliminary results of the NML implementation of the LIM architecture. We thus demonstrate that it is not only a good solution for a standard CMOS technology but can also exploit the potential of an emerging technology as NML.
D. Pala, Giovanni Causapruno, Marco Vacca, Fabrizio Riente, Giovanna Turvani, Mariagrazia Graziano, Maurizio Zamboni
ISCAS7
2015 Interleaving in Systolic-Arrays: A Throughput Breakthrough
abstract
In past years the most common way to improve computers performance was to increase the clock frequency. In recent years this approach suffered the limits of technology scaling, therefore computers architectures are shifting toward the direction of parallel computing to further improve circuits performance. Not only GPU based architectures are spreading in consideration, but also Systolic Arrays are particularly suited for certain classes of algorithms. An important point in favor of Systolic Arrays is that, due to the regularity of their circuit layout, they are appealing when applied to many emerging and very promising technologies, like Quantum-dot Cellular Automata and nanoarrays based on Silicon NanoWire or on Carbon nanotube Field Effect Transistors. In this work we present a systematic method to improve Systolic Arrays performance exploiting Pipelining and Input Data Interleaving. We tackle the problem from a theoretical point of view first, and then we apply it to both CMOS technology and emerging technologies. On CMOS we demonstrate that it is possible to vastly improve the overall throughput of the circuit. By applying this technique to emerging technologies we show that it is possible to overcome some of their limitations greatly improving the throughput, making a considerable step forward toward the post-CMOS era.
Giovanni Causapruno, Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni
IEEE Trans. Computers4
2015 Protein Alignment Systolic Array Throughput Optimization
abstract
Protein comparison is gaining importance year after year since it has been demonstrated that biologists can find correlation between different species, or genetic mutations that can lead to cancer and genetic diseases. Protein sequence alignment is the most computational intensive task when performing protein comparison. To speed-up alignment, dedicated processors that can perform different computations in parallel have been designed. Among them, the best performance has been achieved using systolic arrays (SAs). However, when the processing elements of the SA have an internal loop, performance could be highly reduced. In this paper, we present an architectural strategy to address this problem applying pipeline interleaving; this strategy is applied to an SA for Smith Waterman algorithm that we designed. Results encourage the adoption of pipeline interleaving for parallel circuits with loop-based processing elements. We demonstrate that important benefits in terms of higher operating frequency can be derived without so relevant costs as increased complexity, area, and power required.
Giovanni Causapruno, Gianvito Urgese, Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Feedbacks in QCA: A Quantitative Approach
abstract
In the post-CMOS scenario a primary role is played by the quantum-dot cellular automata (QCA) technology. Irrespective of the specific implementation principle (e.g., either molecular, or magnetic or semiconductive in the current scenario) the intrinsic deep-level pipelined behavior is the dominant issue. It has important consequences on circuit design and performance, especially in the presence of feedbacks in sequential circuits. Though partially already addressed in literature, these consequences still must be fully understood and solutions thoroughly approached to allow this technology any further advancement. This paper conducts an exhaustive analysis of the effects and the consequences derived by the presence of loops in QCA circuits. For each problem arisen, a solution is presented. The analysis is performed using as a test architecture, a complex systolic array circuit for biosequences analysis (Smith-Waterman algorithm), which represents one of the most promising application for QCA technology. The circuit is based on nanomagnetic logic as QCA implementation, is designed down to the layout level considering technological constraints and experimentally validated structures, counts up to approximately 2.3 milion nanomagnets, and is described and simulated with HDL language using as a testbench realistic protein alignment sequences. The results here presented constitute a fundamental advancement in the emerging technologies field since: 1) they are based on a quantitative approach relying on a realistic and complex circuit involving a large variety of QCA blocks; 2) they strictly are reckoned starting from current technological limits without relying on unrealistic assumptions; 3) they provide general rules to design complex sequential circuits with intrinsically pipelined technologies, like QCA; and 4) they prove with a real application benchmark how to maximize the circuits performance.
Marco Vacca, Juanchi Wang, Mariagrazia Graziano, Massimo Ruo Roch, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.5
2014 Fault tolerant nanoarray circuits: Automatic design and verification
abstract
We automatically maximize fault-tolerance in nanoarrays based on silicon nanowires and Gate-All-Around transistors optimizing their topology vs. several distributions of faults inherited by technology. We added a Monte Carlo engine in our nanoarchitecture design tool ToPoliNano and verified the effectiveness of the fault-tolerance algorithm over several circuits and faults distributions.
Pasquale Ranone, Giovanna Turvani, Fabrizio Riente, Mariagrazia Graziano, Massimo Ruo Roch, Maurizio Zamboni
VTS6
2014 Simulation and design of an UWB imaging system for breast cancer detection
Xiaolu Guo, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni
Integr.4
2014 Nanoarray architectures multilevel simulation
abstract
Density and regularity are deemed as the major advantages of nanoarray architectures based on nanowires. Literature demonstrated that proper reliability analyzes must be performed and solutions have to be devised to improve nanoarrays yield. Their complexity and high-fault probability claim for specific design automation tools able to explore circuit solutions, performance and fault-tolerant approaches. We envision a simulator conceived to carry on characterizations in terms of logic behavior, defect-induced output error rate assessment, switching activity, power and timing performance. Though already existing for traditional technology, a simulator based on specific technological and topological tiled nanoarray descriptions, and conceived to join both device and architecture levels, has never been attempted at the degree of accuracy we present. Our contribution is twofold. First, marking a difference with respect to the state of the art, we developed an algorithm based on an event-driven engine which works at switch level and is not simply built on top of cost functions evaluations. The straightforward advantage is the possibility to follow the evolution of dynamic control sequences throughout all the inner components of the nanoarray, and, as a consequence, to obtain circuit level characterization as a projection of the real internal parameters. Second, we added to our simulator the capability to inject faults with specific statistical distributions associated to the nanoarray topology. Here we extract output error rates and yield for one of the possible nanoarray structures proposed in literature, the NASIC. Results specificity and accuracy demonstrate the simulator trustworthiness, its effectiveness for extensive nanoarrays characterization and its suitability as a foundation for both higher architectural and lower device simulation levels. The aim of this work, then, is to provide insights into the intertwined relation between actual technology and circuit design for these emerging fabrics, and, as a consequence, to clarify how defects and variability affect circuits and systems performance.
Stefano Frache, Mariagrazia Graziano, Maurizio Zamboni
ACM J. Emerg. Technol. Comput. Syst.3
2014 Enabling design and simulation of massive parallel nanoarchitectures
Stefano Frache, Diego Chiabrando, Mariagrazia Graziano, Marco Vacca, Luca Boarino, Maurizio Zamboni
J. Parallel Distributed Comput.6
2014 UWB microwave imaging for breast cancer detection: Many-core, GPU, or FPGA?
abstract
An UWB microwave imaging system for breast cancer detection consists of antennas, transceivers, and a high-performance embedded system for elaborating the received signals and reconstructing breast images. In this article we focus on this embedded system. To accelerate the image reconstruction, the Beamforming phase has to be implemented in a parallel fashion. We assess its implementation in three currently available high-end platforms based on a multicore CPU, a GPU, and an FPGA, respectively. We then project the results applying technology scaling rules to future many-core CPUs, many-thread GPUs, and advanced FPGAs. We consider an optimistic case in which available resources increase according to Moore's law only, and a pessimistic case in which only a fraction of those resources are available due to a limited power budget. In both scenarios, an implementation that includes a high-end FPGA outperforms the other alternatives. Since the number of effectively usable cores in future many-cores will be power-limited, and there is a trend toward the integration of power-efficient accelerators, we conjecture that a chip consisting of a many-core section and a reconfigurable logic section will be the perfect platform for this application.
Mario R. Casu, Francesco Colonna, Marco Crepaldi, Danilo Demarchi, Mariagrazia Graziano, Maurizio Zamboni
ACM Trans. Embed. Comput. Syst.6
2013 A Hardware Viewpoint on Biosequence Analysis: What's Next?
abstract
Biosequence alignment recently received an increasing support from both commodity and dedicated hardware platforms. Processing capabilities are constantly rising, but still not satisfying the limitless requirements of this application. We give an insight on the contribution to this need that can possibly be expected from emerging technology devices and architectures, focusing as an example on nanofabrics based on silicon nanowires. By varying a few parameters we explore the solution space, and demonstrate with proper figures of merit how this family of beyond CMOS structures could be considered as the effective disruptive technology for biosequence analysis applications.
Mariagrazia Graziano, Stefano Frache, Maurizio Zamboni
ACM J. Emerg. Technol. Comput. Syst.3
2013 Nanomagnetic Logic Microprocessor: Hierarchical Power Model
abstract
The interest in emerging nanotechnologies has been recently focused on nanomagnetic logic (NML), which has unique appealing features. NML circuits have very low power consumption and, because of their magnetic nature, maintain the information safely stored even without power supply. The nature of these circuits is much different from that of CMOS circuits. As a consequence, to better understand NML logic, complex circuits and not only simple gates must be designed. This constraint calls for a new design and simulation methodology. It should efficiently encompass manifold properties: 1) being based on commonly used hardware description language (HDL) in order to easily manage complexity and hierarchy; 2) maintaining a clear link with physical characteristics; and 3) modeling performance aspects such as speed and power, together with logic behavior. In this paper, we present a very-high-speed integrated circuits HDL (VHDL) behavioral model for NML circuits, which allows the evaluation of not only the logic behavior but also its power dissipation. It is based on a technological solution called “snake-clock.” We demonstrate this model using a case study which offers the right variety of internal substructures to test the method: a 4-bit microprocessor designed using asynchronous logic. The model enables a hierarchical bottom-up evaluation of the processor logic behavior, area, and power dissipation, which we evaluate using a benchmark division algorithm. The results highlight the flexibility and the efficiency of this model, as well as the remarkable improvements that it brings to the analysis of NML circuits.
Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Asynchronous Solutions for Nanomagnetic Logic Circuits
abstract
In the years to come new solutions will be required to overcome the limitations of scaled CMOS technology. One approach is to adopt Nano-Magnetic Logic Circuits, highly appealing for their extremely reduced power consumption. Despite the interesting nature of this approach, many problems arise when this technology is considered for real designs. The wire is the most critical of these problems from the circuit implementation point of view. It works as a pipelined interconnection, and its delay in terms of clock cycles depends on its length. Serious complications arise at the design phase, both in terms of synthesis and of physical design. One possible solution is the use of a delay insensitive asynchronous logic, Null Convention Logic (NCL TM ). Nevertheless its use has many negative consequences in terms of area occupation and speed loss with respect to a Boolean version. In this article we analyze and compare different solutions: nanomagnetic circuits based on full NCL, mixed Boolean-NCL, and fully Boolean logic. We discuss the advantages of these logics, but also the issues they raise. In particular we analyze feedback signals, which, due to their intrinsic pipelined nature, cause errors that still have not found a solution in the literature. The innovative arrangement we propose solves most of the problems and thus soundly increases the knowledge of this technology. The analysis is performed using a VHDL behavioral model we developed and a microprocessor we designed based on this model, as a sound and realistic test bench.
Marco Vacca, Mariagrazia Graziano, Maurizio Zamboni
ACM J. Emerg. Technol. Comput. Syst.3
2010 MEDEA: a hybrid shared-memory/message-passing multiprocessor NoC-based architecture
abstract
The shared-memory model has been adopted, both for data exchange as well as synchronization using semaphores in almost every on-chip multiprocessor implementation, ranging from general purpose chip multiprocessors (CMPs) to domain specific multi-core graphics processing units (GPUs). Low-latency synchronization is desirable but is hard to achieve in practice due to the memory hierarchy. On the contrary, an explicit exchange of synchronization tokens among the processing elements through dedicated on-chip links would be beneficial for the overall system performance. In this paper we propose the Medea NoC-based framework, a hybrid shared-memory/message-passing approach. Medea has been modeled with a fast, cycle-accurate SystemC implementation enabling a fast system exploration varying several parameters like number and types of cores, cache size and policy and NoC features. In addition, every SystemC block has its RTL counterpart for physical implementation on FPGAs and ASICs. A parallel version of the Jacobi algorithm has been used as a test application to validate the methodology. Results confirm expectations about performance and effectiveness of system exploration and design.
Sergio Tota, Mario R. Casu, Massimo Ruo Roch, Luca Rostagno, Maurizio Zamboni
DATE5
2010 A flexible simulation methodology and tool for nanoarray-based architectures
abstract
Nanoscale arrays based on nanowires are expected to have a promising future thanks to their amazing density and regularity. Experiments demonstrated the feasibility of this technology and pointed out that accurate reliability analyses should be accomplished to assure proper yield requirements. Due to the complexity of these systems and the arising necessity of thorough fault analysis, design automation tools are mandatory in order to explore architectural solutions and fault tolerant approaches deriving information from reliable nanoarray characterisation. We present a simulator, never attempted at this level of detail, based on specific technological and topological tiled nanoarray descriptions, conceived to carry on characterisations in terms of logic behaviour, defect-induced error rate assessment, switching activity and other figures of merit like power and timing performance (not discussed in this paper). It is formulated in a flexible and modular way to assure the simulation of manifold advancing technological solutions, among which the winner has not been determined yet. Marking a difference with respect to the state of the art, the algorithm is based on an event-driven engine and not on cost functions evaluations. Thus even dynamic control sequences can be processed and their evolution followed throughout all the inner components of the array allowing to obtain system level characterization as a projection of the real internal parameters. In this paper we show results attained for one of the possible nanoarray structures proposed in literature, the NASIC: logic behaviour, defect error rates and switching activity for two types of function demonstrate the simulator trustworthiness, its effectiveness for extensive nanoarrays characterisation and its suitability as a foundation for both higher architectural and lower device simulation levels.
Stefano Frache, Mariagrazia Graziano, Maurizio Zamboni
ICCD3
2009 A mixed-signal demodulator for a low-complexity IR-UWB receiver: Methodology, simulation and design
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni
Integr.4
2009 A Case Study for NoC-Based Homogeneous MPSoC Architectures
abstract
The many-core design paradigm requires flexible and modular hardware and software components to provide the required scalability to next-generation on-chip multiprocessor architectures. A multidisciplinary approach is necessary to consider all the interactions between the different components of the design. In this paper, a complete design methodology that tackles at once the aspects of system level modeling, hardware architecture, and programming model has been successfully used for the implementation of a multiprocessor network-on-chip (NoC)-based system, the NoCRay graphic accelerator. The design, based on 16 processors, after prototyping with field-programmable gate array (FPGA), has been laid out in 90-nm technology. Post-layout results show very low power, area, as well as 500 MHz of clock frequency. Results show that an array of small and simple processors outperform a single high-end general purpose processor.
Sergio Tota, Mario R. Casu, Massimo Ruo Roch, Luca Macchiarulo, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.5
2007 An effective AMS top-down methodology applied to the design of a mixed-signal UWB system-on-chip
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni
DATE4
2004 An electromigration and thermal model of power wires for a priori high-level reliability prediction
abstract
In this paper, a simple power-distribution electrothermal model including the interconnect self-heating is used together with a statistical model of average and rms currents of functional blocks and a high-level model of fanout distribution and interconnect wirelength. Following the 2001 SIA roadmap projections, we are able to predict a priori that the minimum width that satisfies the electromigration constraints does not scale like the minimum metal pitch in future technology nodes. As a consequence, the percentage of chip area covered by power lines is expected to increase at the expense of wiring resources unless proper countermeasures are taken. Some possible solutions are proposed in the paper.
Mario R. Casu, Mariagrazia Graziano, Guido Masera, Gianluca Piccinini, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.5
2003 Wireless sensor networks: a power-scalable motion estimation IP for hybrid video coding
abstract
Wireless Sensor Networks are an emerging phenomenon in the research community. The design and development of network architectures and nodes implementation are fostering many research activities. Due to their wide application fields and pervasive employment possibilities, the investigation of novel classes of wireless sensor nodes is of great concern. In this paper we presented a novel Power-Scalable Motion Estimation IP suitable for video-surveillance over Wireless Sensor Networks. The proposed architecture can achieve low dynamic power, good video quality and reasonable frame-rates. Moreover, it can operate on different video format and can be reconfigured both off-line and on-line. Further researches have to be accomplished in order to reduce the internal critical path, achieving better maximum frequencies performances. Moreover, new FPGA architectures must be exploited searching low-static-power devices suitable for an actual implementation of our IP on a WSN.
Federico Quaglio, Maurizio Martina, Fabrizio Vacca, Guido Masera, Andrea Molino, Gianluca Piccinini, Maurizio Zamboni
FPGA7
2002 Energy Evaluation on a Reconfigurable, Multimedia-Oriented Wireless Sensor
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni
FPL5
2002 Reconfigurable DSP IP for multimedia applications
abstract
In this paper a novel Digital Signal Processor IP for multimedia applications, is presented. Recently, develeper's interest towards SOC architectures has been driven by mobile market explosion. Despite the increasing importance gathered by reconfigurable computing, a lack of easily retargettable cores is felt by developer's community. This IP is intended to be the kernel for many telecommunication and multimedia algorithms computation. During the design flow, much care has been devoted to grant maximum interoperability among this DSP core and other coprocessor units, allowing to easily embed multiple functional blocks on a single FPGA. As far as performance are concerned, the proposed IP shows satisfactory results both in terms of area occupation (11\% on a XILINX XCV1000) and maximum clock frequency (89 MHz after place and route process).
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni
ICASSP5
2002 Architectural strategies for low-power VLSI turbo decoders
abstract
The use of "turbo codes" has been proposed for several applications, including the development of wireless systems, where highly reliable transmission is required at very low signal-to-noise ratios (SNR). The problem of extracting the best coding gains from these kind of codes has been deeply investigated in the last years. Also the hardware implementation of turbo codes is a very challenging topic, mainly due to the iterative nature of the decoding process, which demands an operating frequency much higher than the data rate; in the case of wireless applications, the design constraints became even more strict due to the low-cost and low-power requirements. This paper first presents a new architecture for the decoder core with improved area and power dissipation properties; then partitioning techniques are proposed to reduce the power consumption of the decoder memories. It is proven that most of the power is dissipated by the large RAM units required by the decoder, so the described technique is very efficient: an average power saving of 70% with an area overhead of 23% has been obtained on a set of analyzed architectures.
Guido Masera, Marco Mazza, Gianluca Piccinini, Fabrizio Viglione, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.5
2001 Synthesis of low-leakage PD-SOI circuits with body-biasing
abstract
Article Share on Synthesis of low-leakage PD-SOI circuits with body-biasing Authors: Mario Casu Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Gianluca Piccinini Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Guido Masera Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Maurizio Zamboni Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 287–290https://doi.org/10.1145/383082.383170Online:06 August 2001Publication History 2citation166DownloadsMetricsTotal Citations2Total Downloads166Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Mario R. Casu, Gianluca Piccinini, Guido Masera, Maurizio Zamboni
ISLPED4
2000 A 50 Mbit/s Iterative Turbo-Decoder
abstract
Very low bit error rate has become an important constraint in high performance communication systems that operate at very low signal to noise ratios: due to their impressive coding gains, turbo codes have been proposed for several applications, although they suffer a large decoding delay. This paper presents the design of a turbo decoder with high performances in terms of throughput implemented using TSPC (true single phase clocking) logic family. In order to achieve the best compromise between cost (in terms of area) and throughput, several architectural solutions have been analyzed. The whole system and in particular its core, the SISO module, has been verified through VHDL simulations. HSPICE simulations show that the system can operate with a 1 GHz clock and thus it can reach a throughput of 50 Mbit/s.
Fabrizio Viglione, Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni
DATE5
2000 A high accuracy-low complexity model for CMOS delays
abstract
This paper presents a new model for CMOS structures delays estimation based on a deep analysis of complex gates behavior. This approach can supply a high level of accuracy. A complex structure is reduced first to series-connected MOS, then the delay equations are applied to that reduced rate. The model is based on a time piecewise linearization so that a strongly nonlinear circuit can he solved using well known linear techniques. The delay formulas involve model parameters as MOS width functions, therefore providing routines suitable for optimization algorithms. The high level of accuracy, the low CPU time and the high degree of scaling capability are proved in the paper. These features make the model attractive for deep submicron technologies.
Mario R. Casu, Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni
ISCAS5
1999 New 2 Gbit/s CMOS I/O pads
abstract
A couple of low complexity high performance input and output pads are proposed: they have been designed in 0.7 /spl mu/m CMOS ES2 technology and support bit rates ranging from DC up to 2 Gbit/s. The differential input pad and the differential output pad interface true PECL external logic levels to full swing 5 V CMOS internal levels.
Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni
Great Lakes Symposium on VLSI4
1999 VLSI architectures for turbo codes
abstract
A great interest has been gained in recent years by a new error-correcting code technique, known as "turbo coding", which has been proven to offer performance closer to the Shannon's limit than traditional concatenated codes. In this paper, several very large scale integration (VLSI) architectures suitable for turbo decoder implementation are proposed and compared in terms of complexity and performance; the impact on the VLSI complexity of system parameters like the state number, number of iterations, and code rate are evaluated for the different solutions. The results of this architectural study have then been exploited for the design of a specific decoder, implementing a serial concatenation scheme with 2/3 and 3/4 codes; the designed circuit occupies 35 mm/sup 2/, supports a 2 Mb/s data rate, and for a bit error probability of 10/sup -6/, yields a coding gain larger than 7 dB, with ten iterations.
Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni
IEEE Trans. Very Large Scale Integr. Syst.4
1998 A receiver architecture conforming to the OFDM based digital video broadcasting standard for terrestrial transmission (DVB-T)
abstract
We present the architecture for a four chip integrated receiver decoder compliant with the European Digital Terrestrial Video Transmission standard (DVB-T); both 8 K and 2 K FFT size are supported. The paper focuses on the algorithms used for I/Q signal generation, OFDM demultiplexing, channel estimation, synchronization and channel decoding. The algorithms chosen are robust against the channel and interference conditions to be expected, including Rayleigh fading, co-channel interference from other (e.g. existing analogue) services and tuner imperfections. The channel SNR degradation arising in the digital part of the receiver-with imperfect channel estimation, synchronization and quantization-with respect to the BER at the output of the Viterbi decoder is not more than 2 dB from theory.
Pierre Combelles, Christophe Del Toso, Dietmar Hepper, David Le Goff, J. J. Ma, Patrick Robertson, Fabio Scalise, Laurent Soyer, Maurizio Zamboni
ICC9
1998 Fanout optimization under a submicron transistor-level delay model
abstract
In this paper we present a new fanout optimization algorithm whichisparticularly suitable for digital circuits designed with submicron CMOS technologies. Restricting the class of fanout trees to the so-called bipolar LT-trees, the topology of the optimal fanout tree is found by means of a dynamic programming algorithm. The bu#er selection is in turn performed by using a continuous bu#er sizing technique based on a very accurate delay model especially developed for submicron CMOS processes. The fanout trees can distribute a signal with arbitrary polarity from the root of the tree to a set of sinks with arbitrary required time, required minimum signal slope, polarity and capacitive load. These trees can be constructed to maximize the required time at the root or to minimize the total bu#er area under a required time constraint at the root. The performance of the algorithm shows several improvements with respect to conventional fanout optimization methods. More precisely, the area and del...
Pasquale Cocchini, Massoud Pedram, Gianluca Piccinini, Maurizio Zamboni
ICCAD4
1987 An Experimental VLSI Prolog Interpreter: Preliminary Measurements and Results
abstract
This work presents the preliminary results of a project oriented to the design and VLSI implementation of a Prolog interpreter. Even if the interpretative approach is being considered an inefficient way to execute high level languages when compared to that of compilation, declarative languages with embedded extralogical instructions would require the use of direct execution.
Pierluigi Civera, Franco Maddaleno, Gianluca Piccinini, Maurizio Zamboni
ISCA4
1987 Design considerations on a VLSI Prolog interpreter
Pierluigi Civera, Gianluca Piccinini, Maurizio Zamboni
Microprocess. Microprogramming3
1986 Monitoring tools for multiprocessors
Francesco Gregoretti, Franco Maddaleno, Maurizio Zamboni
Microprocessing and Microprogramming3