Roberto Canegallo

dblp:57/1770 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 A Threshold Voltage Generator Circuit with Automatic Refresh and Dynamic Updating for Ultra-Low-Power Continuous-Time Comparators
abstract
A Threshold Voltage Generator (TVG) circuit for continuous-time comparators is presented. It may be used in ULP IoT systems requiring nanoWatt power consumption, kbps bitrates and the reception of packets with lengths up to hundreds of bits. The circuit is based on a switched capacitor technique to generate the threshold voltage without requiring any large area resistors and with the system clock active only during the reception of data, thus minimizing energy consumption. Latency is also minimized as the threshold is generated within the first received bit of the packet. The reception of packets with no limits on their length is made possible by continuously updating the threshold, which also allows correct operation even in case of amplitude variations in the incoming signal during the data reception. It has been implemented and verified through post-layout simulations in an STMicroelectronics 90-nm CMOS technology with a 0.6-V supply, targeting a 1-kbps bitrate. It occupies an area lower than 0.001 mm2, which is less than 1% of the area of a standard RC-based TVG implemented in the same technology. A prototype of the proposed TVG is currently under fabrication.
Matteo D'Addato, Luca Perilli, Alessia Maria Elgani, Eleonora Franchi, Antonio Gnudi, Roberto Canegallo, Giulio Ricotti
ISCAS6
2022 Phase-Change Memory in Neural Network Layers with Measurements-based Device Models
abstract
The search for energy efficient circuital implementations of neural networks has led to the exploration of phase-change memory (PCM) devices as their synaptic element, with the advantage of compact size and compatibility with CMOS fabrication technologies. In this work, we describe a methodology that, starting from measurements performed on a set of real PCM devices, enables the training of a neural network. The core of the procedure is the creation of a computational model, sufficiently general to include the effect of unwanted non-idealities, such as the voltage dependence of the conductances and the presence of surrounding circuitry. Results show that, depending on the task at hand, a different level of accuracy is required in the PCM model applied at train-time to match the performance of a traditional, reference network. Moreover, the trained networks are robust to the perturbation of the weight values, up to 10% standard deviation, with performance losses within 3.5% for the accuracy in the classification task being considered and an increase of the regression RMS error by 0.014 in a second task. The considered perturbation is compatible with the performance of state-of-the-art PCM programming techniques.
Carmine Paolino, Alessio Antolini, Fabio Pareschi, Mauro Mangia, Riccardo Rovatti, Eleonora Franchi, Gianluca Setti, Roberto Canegallo, Marcella Carissimi, Marco Pasotti
ISCAS8
2021 Compressed Sensing by Phase Change Memories: Coping with Encoder non-Linearities
abstract
Several recent works have shown the advantages of using phase-change memory (PCM) in developing brain-inspired computing approaches. In particular, PCM cells have been applied to the direct computation of matrix-vector multiplications in the analog domain. However, the intrinsic nonlinearity of these cells with respect to the applied voltage is detrimental. In this paper we consider a PCM array as the encoder in a Compressed Sensing (CS) acquisition system, and investigate the effect of the non-linearity of the cells. We introduce a CS decoding strategy that is able to compensate for PCM nonlinearities by means of an iterative approach. At each step, the current signal estimate is used to approximate the average behaviour of the PCM cells used in the encoder. Monte Carlo simulations relying on a PCM model extracted from an STMicrolectronics 90 nm BCD chip validate the performance of the algorithm with various degrees of nonlinearities, showing up to 35 dB increase in median performance as compared to standard decoding procedures.
Carmine Paolino, Alessio Antolini, Fabio Pareschi, Mauro Mangia, Riccardo Rovatti, Eleonora Franchi, Antonio Gnudi, Gianluca Setti, Roberto Canegallo, Marcella Carissimi, Marco Pasotti
ISCAS9
2020 Nanowatt Clock and Data Recovery for Ultra-Low Power Wake-Up Based Receivers
Matteo D'Addato, Alessio Antolini, Francesco Renzini, Alessia Maria Elgani, Luca Perilli, Eleonora Franchi, Antonio Gnudi, Michele Magno, Roberto Canegallo
EWSN9
2018 Dual-Mode Wake-Up Nodes for IoT Monitoring Applications: Measurements and Algorithms
abstract
Internet of Things (IoTs)-based monitoring applications usually involve large-scale deployments of battery-enabled sensor nodes providing measurements at regular intervals. In order to guarantee the service continuity over time, the energy-efficiency of the networked system should be maximized. In this paper, we address such issue via a combination of novel hardware/software solutions including new classes of Wake-up radio IoT Nodes (WuNs) and novel data- and hardware-driven network management algorithms. Three main contributions are provided. First, we present the design and prototype implementation of WuN nodes able to support two different energy-saving modes; such modes can be configured via software, and hence dynamically tuned. Second, we show by experimental measurements that the optimal policy strictly depends on the application requirements. Third, we move from the node design to the network design, and we devise proper orchestration algorithms which select both the optimal set of WuN to wake-up and the proper energy-saving mode for each WuN, so that the application lifetime is maximized, while the redundancy of correlated measurements is minimized. The proposed solutions are extensively evaluated via OMNeT++ simulations under different IoT scenarios and requirements of the monitoring applications.
Luca Bedogni, Luciano Bononi, Roberto Canegallo, Fabio Carbone, Marco Di Felice, Eleonora Franchi, Federico Montori, Luca Perilli, Tullio Salmon Cinotti, Angelo Trotta
ICC3
2014 Multicore Signal Processing Platform With Heterogeneous Configurable Hardware Accelerators
abstract
The computing demand of many signal processing algorithms is dramatically growing because of the increasing complexity of embedded software applications. Concurrently, as process technology scales, the design effort for realizing very large scale integrated circuits and the associated costs are becoming critically high. A possible solution to address this performance/costs challenge is given by customizable multiprocessor system-on-chips. The approach proposed in this paper leads to the customization of multi/many processor system-on-chip at two levels of abstraction: 1) customization through application-specific hardware accelerators implemented on configurable datapath that can target three kinds of structured application-specific integrated circuit technologies: metal, via, and runtime programmable and 2) customization of the architectural parameters of the platform. The proposed platform is equipped with a design framework that assists the user in the high-level design-space exploration of signal processing applications described using the Open Computing Language (OpenCL) language. A peculiar added value of the flow is to support the migration of OpenCL kernels and tasks into pipelined hardware accelerators described using a C-level language. The platform is able to provide an average performance of 90 GOPS on a set of reference signal processing applications, and an average computational energy efficiency of 130 GOPS/W in its metal-programmable configuration. This result shows the benefits in terms of energy efficiency of hardware customization applied to multiprocessor systems with respect to many core devices such as general-purpose graphic processing units, able to provide on average 2.5 GOPS/W for the applications under analysis.
Davide Rossi 0001, Claudio Mucci, Matteo Pizzotti, Luca Perugini, Roberto Canegallo, Roberto Guerrieri
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Input/Output Pad for Direct Contact and Contactless Testing
abstract
Non-contact probing can provide an important contribution for testing complex Systems-on-a-Chip (SoC), Systems-in-a-Package (SiP) and Through-Silicon-Vias (TSV) interconnections. This paper demonstrates the feasibility of wireless testing by capacitive coupling between a cantilever probe card and a pad. In particular a scheme of an I/O pad suitable for both contact and contactless probing is proposed.
Mauro Scandiuzzo, Salvatore Cani, Luca Perugini, Simone Spolzino, Roberto Canegallo, Luca Perilli, Roberto Cardu, Eleonora Franchi, C. Gozzi, F. Maggioni
ETS5
2010 Characterization of chip-to-chip wireless interconnections based on capacitive coupling
abstract
3D chip-to-chip capacitive interconnections are in common practice characterized with FEM solvers as they cannot be modeled as lumped RLC circuits as ohmic 3D interconnects. This paper describes some drawbacks of this procedure and proposes an innovative flow, based on post-layout parasitic extraction tools, to enable the designer to place capacitive interconnects as constrained macros in a digital design flow.
Roberto Cardu, Eleonora Franchi, Roberto Guerrieri, Mauro Scandiuzzo, Salvatore Cani, Luca Perugini, Simone Spolzino, Roberto Canegallo
VLSI-SoC8
2006 Yield prediction for 3D capacitive interconnections
abstract
Capacitive interconnections are very promising structures for high-speed and low-power signaling in 3D packages. Since the performance of AC links, in terms of Band-Width and Bit-Error-Rate (BER), depends on assembly and synchronization accuracy we performed a statistical analysis of assembly procedures and communication circuits. In this paper we present a yield prediction methodology for 3D capacitive links: starting from the analysis of communication circuits and BER measurements, we analyze stacking variability in order to predict reliability and performance. The proposed parametric yield analysis is demonstrated on a test-case, with constrained inter-electrode coupling and operating frequency.
Alberto Fazzi, Luca Magagni, Mario de Dominicis, Paolo Zoffoli, Roberto Canegallo, Pier Luigi Rolandi, Alberto L. Sangiovanni-Vincentelli, Roberto Guerrieri
ICCAD5
2005 A low-power system-on-chip for the documentation of road accidents
abstract
In this letter, the implementation of a system-on-chip for the documentation of road accidents is presented. Key features of the system are the implementation on a programmable architecture of a compression algorithm capable of encoding up to 20 black and white QCIF frames/s, and the computation of a digital signature performed every frame which is applied to the encoded bitstream certifying the source of the video sequence. The acquired images are then stored on a battery of on-chip flash memories that can be used to retrieve accident information to be used in legal confrontations. Having to perform in critical energy conditions (e.g., continued image acquisition for up to 10-20 s after an accident has occurred), the system was designed to minimize energy consumption at all levels. The system-on-chip has been implemented in 6/spl times/6 mm/sup 2/ on a 0.25-/spl mu/m 6-metal standard-cell CMOS technology and works at 40 MHz with a 2.5-V power supply. Performance can be decreased to 12 QCIF frames/s, 24 MHz, and 1.3-V power supply in order to achieve 30-mW power consumption.
Luca Bolcioni, Fabio Campi, Roberto Canegallo, Roberto Guerrieri
IEEE Trans. Circuits Syst. Video Technol.3
1997 Words Recognition using Associative Memory
abstract
Introduces the application of an analog associative memory chip to word recognition, which is a fundamental topic of the text recognition process. The word recognition method takes advantage of a statistical evaluation of the behavior of the optical character recognition system preceding it. That statistical information leads to the creation of a coding that is used to store a lexicon of the most used words in the chip. An input pattern is matched against the full database of the associative memory, and a set of closest patterns is returned. The precision reached by this operation ranges from 93% to 99%. These encouraging results demonstrate the general aptitude of the chip to solve classes of problems that need to use an associative memory.
Loris Navoni, Roberto Canegallo, Mauro Chinosi, Giovanni Gozzini, Alan Kramer, Pier Luigi Rolandi
ICDAR2