Samuel Palermo

dblp:34/10283 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2023
0000-0002-6555-1474ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 since 2021
YearPublicationVenuePosition
2023 Memristor-based Offset Cancellation Technique in Analog Crossbars
abstract
Analog computing platforms have been a popular and promising research area that suggest efficient ways of computation compared to its digital counterparts. Memristor based crossbars drew attention by computing the vector-matrix calculation intensive tasks such as Artificial Intelligence (AI) and Machine Learning (ML) in one time step. Although they provide an energy efficient way of computing these tasks, analog computation in general suffers from non-idealities and systematic errors in the circuitry, which could degrade the performance and accuracy significantly. One of the issues is the random offset associated with the op-amps in the system resulting from the process and mismatch variations. In this paper, a novel technique is offered to reduce the negative effects of the random offset and increase the output accuracy. This newly proposed system uses minimum extra circuitry and additional power consumption and only requires the crossbar to be enlarged by two extra rows. The intrinsic issue of the analog crossbars, interconnect parasitics, must be incorporated into the problem, and a way to separate the offset and wire resistance issues from each other is offered. The functionality of the system has been shown with a case study in the results section where the op-amps have$\sigma_{offset}=3mV$. The effectiveness of the offered technique demonstrates a 6× better accuracy with the mitigation of the offset problem. The proposed method can be used in memristor and other analog crossbars to achieve a greater performance and thus improve their competitiveness.
Anil Korkmaz, Gianluca Zoppo, Francesco Marrone, Fernando Corinto, Su-In Yi, R. Stanley Williams, Samuel Palermo
ISCAS7
2023 Gaussian Process for Nonlinear Regression via Memristive Crossbars
abstract
Over the last decade, Gaussian processes (GPs) have become popular in the area of machine learning and data analysis for their flexibility and robustness. Despite their attractive formulation, practical use in large-scale problems remains out of reach due to computational complexity. Existing direct computational methods for manipulations involving large-scale$n\times n$covariance matrices require$O(n^{3})$calculations. In this work, we present the design and evaluation of a simulated computing platform for exact GP inference, that achieves true model parallelism using memristive crossbars. To achieve a one-shot solution, a linear equation solver and a vector-matrix multiplication solver crossbar configurations are used together, reducing the number of operations from$O(n^{3})$to$O(n)$. The transistor level op-amps, ADC models for quantization, circuit and interconnect parasitics, together with the finite memristor precision are incorporated into the system simulation. The analog system resulted in %1.51 mean error and %2.93 average variance error in solving a nonlinear regression problem. The proposed method achieved 9× to 144× better energy efficiency compared to TPU and 7× compared to a custom analog linear regression solver.
Gianluca Zoppo, Anil Korkmaz, Francesco Marrone, Su-In Yi, Samuel Palermo, Fernando Corinto, R. Stanley Williams
ISCAS5
2022 SEE Sensitivity of a 16GHz LC-Tank VCO in a 22nm FinFET Technology
abstract
Voltage-controlled oscillators (VCOs) are essential components of frequency synthesizers used in space-based applications. Radiation in harsh space environments can cause SEEs (single-event effects) when a charged particle hits the silicon substrate and induces a transient current within the circuit that causes erroneous behavior. Hence, it is essential to understand the radiation sensitivity of a given process in order to design circuits that can tolerate these conditions. An LC-tank VCO is designed in a 22nm FinFET process and achieves 12.5-16.8GHz tuning range and −126dBc/Hz phase noise at a 10MHz offset from a 16.0GHz carrier. SEE sensitivity is quantified with a measured VCO maximum cross-sectional area of 350um2when testing with heavy ions with a linear energy transfer from 10 to 90 MeV.cm2/mg. These results are pertinent for future radiation-hardened VCO designs in this process.
David Dolt, Quintin Livingston, Tong Liu 0041, Samuel Palermo
ISCAS5
2022 Analog Acceleration of the Power Method using Memristor Crossbars
abstract
Determining the dominant eigenvector of matrices and graphs is one of the most fundamental tasks in many machine learning problems, including spectral clustering, Hyperlink Induced Topic Search (HITS), Markov Chains, PageRank and eigenvector centralities. Among the several algorithms used, the Power Method is one of the simplest iterative approaches. It relies on multiple vector-matrix multiplications (VMMs) and a normalization step to prevent divergent behaviours. Recently, efficiency of the memristor crossbars in solving VMMs have been demonstrated using fundamental laws of the circuit theory. In this work, we propose a circuit to accelerate the Power iteration algorithm including current-mode termination for the memristor crossbars and a normalization circuit. The normalization step together with the feedback loop of the complete circuit ensure stability and convergence of the dominant eigenvector. The system allows the observation of the evolution of the outputs. We implement a transistor level peripheral circuitry around the memristor crossbar and take non-idealities such as wire parasitics, source driver resistance and finite memristor precision into account. We compute the eigenvector centrality to demonstrate the performance of the proposed system. We compare our results to the ones coming from the conventional digital computers and observe significant energy savings while maintaining a competitive accuracy.
Anil Korkmaz, Gianluca Zoppo, Francesco Marrone, Fernando Corinto, R. Stanley Williams, Samuel Palermo
ISCAS6
2021 Design of Tunable Analog Filters Using Memristive Crossbars
abstract
Tunable front-end filters are necessary in wireless systems that support multiple frequency bands. N-path filters have the potential for wide tuning ranges, but require multiple high-frequency mixing clocks and suffer from harmonic responses. Another option is digital filtering. However, this requires high-speed analog-to-digital converters and the time complexity to perform Discrete Fourier Transform (DFT) and Inverse Discrete Fourier Transform (IDFT) operations using conventional digital computers is O(N2). Conversely, memristor crossbars have experimentally demonstrated the ability to perform various signal processing tasks, including Discrete Cosine Transform (DCT), in one time-step. In this work, this has been taken a step further by presenting full DFT and IDFT operations using memristive crossbars and proposes the implementation of novel continuous-time tunable analog filters by applying filter coefficients in between these DFT and IDFT crossbars. This highly scalable new method of filtering leverages advantages found both in conventional digital and analog filters and allows any digital filter to be implemented in the analog domain with a filter order as large as half the utilized crossbar size. The proposed architecture allows for the generation of arbitrary filter functions with tunable corner (or center) frequencies and bandwidths from 0.2-20GHz. Stopband attenuation greater than 40dB is achieved with as low as 2-bit memristor precision, while close to 80dB is possible with 8-bit precision. The filter system consumes 106mW to support the 20GHz frequency range, resulting in a 5.3mW/GHz energy metric.
Anil Korkmaz, Chaoyi He, Linda Katehi, R. Stanley Williams, Samuel Palermo
ISCAS5
2021 Analog Solutions of Discrete Markov Chains via Memristor Crossbars
abstract
Problems involving discrete Markov Chains are solved mathematically using matrix methods. Recently, several research groups have demonstrated that matrix-vector multiplication can be performed analytically in a single time step with an electronic circuit that incorporates an open-loop memristor crossbar that is effectively a resistive random-access memory. Ielmini and co-workers have taken this a step further by demonstrating that linear algebraic systems can also be solved in a single time step using similar hardware with feedback. These two approaches can both be applied to Markov chains, in the first case using matrix-vector multiplication to compute successive updates to a discrete Markov process and in the second directly calculating the stationary distribution by solving a constrained eigenvector problem. We present circuit models for open-loop and feedback configurations, and perform detailed analyses that include memristor programming errors, thermal noise sources and element nonidealities in realistic circuit simulations to determine both the precision and accuracy of the analog solutions. We provide mathematical tools to formally describe the trade-offs in the circuit model between power consumption and the magnitude of errors. We compare the two approaches by analyzing Markov chains that lead to two different types of matrices, essentially random and ill-conditioned, and observe that ill-conditioned matrices suffer from significantly larger errors. We compare our analog results to those from digital computations and find a significant power efficiency advantage for the crossbar approach for similar precision results.
Gianluca Zoppo, Anil Korkmaz, Francesco Marrone, Samuel Palermo, Fernando Corinto, R. Stanley Williams
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 A 22 Gb/s Directly Modulated Optical Injection-Locked Quantum-Dot Microring Laser Transmitter with Integrated CMOS Driver
abstract
High-speed optical interconnects are a promising solution for the rapidly increasing bandwidth density demands in data center and high-performance computing applications. This paper presents a transmitter that consists of a co-packaged optically injection-locked quantum-dot (QD) microring laser and a 28nm CMOS driver. Optical injection locking provides an effective solution to overcome the inherent low modulation bandwidth (≤ 5 GHz) of the QD microring laser. Further bandwidth extension is provided with the driver's asymmetric 2-tap feed-forward equalizer (FFE) that compensates the laser's non-linear optical dynamics. Combining these techniques allows for a record 22Gb/s operation of an O-band quantum-dot laser heterogeneously integrated in a silicon photonic platform.
Yang-Hang Fan, Sudharsanan Srinivasan, Yingtao Hu, Di Liang, Ruida Liu, Erwen Li, Raymond G. Beausoleil, Samuel Palermo
ISCAS10
2014 LumiNOC: A Power-Efficient, High-Performance, Photonic Network-on-Chip
abstract
To meet energy-efficient performance demands, the computing industry has moved to parallel computer architectures, such as chip multiprocessors (CMPs), internally interconnected via networks-on-chip (NoC) to meet growing communication needs. Achieving scaling performance as core counts increase to the hundreds in future CMPs, however, will require high performance, yet energy-efficient interconnects. Silicon nanophotonics is a promising replacement for electronic on-chip interconnect due to its high bandwidth and low latency, however, prior techniques have required high static power for the laser and ring thermal tuning. We propose a novel nano-photonic NoC (PNoC) architecture, LumiNOC, optimized for high performance and power-efficiency. This paper makes three primary contributions: a novel, nanophotonic architecture which partitions the network into subnets for better efficiency; a purely photonic, in-band, distributed arbitration scheme; and a channel sharing arrangement utilizing the same waveguides and wavelengths for arbitration as data transmission. In a 64-node NoC under synthetic traffic, LumiNOC enjoys 50% lower latency at low loads and ~40% higher throughput per Watt on synthetic traffic, versus other reported PNoCs. LumiNOC reduces latencies ~40% versus an electrical 2-D mesh NoCs on the PARSEC shared-memory, multithreaded benchmark suite.
Mark Browning, Paul Gratz, Samuel Palermo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2013 A Design Methodology for Power Efficiency Optimization of High-Speed Equalized-Electrical I/O Architectures
abstract
Both power efficiency and per-channel data rates of high-speed input/output (I/O) links must be improved in order to support future inter-chip bandwidth demand. In order to scale data rates over band-limited channels, various types of equalization circuitry are used to compensate for frequency-dependent loss. However, this additional complexity introduces power and area costs, requiring selection of an appropriate I/O equalization architecture in order to comply with system power budgets. This paper presents a design flow for power optimization of high-speed electrical links at a given data rate, channel type, and process technology node, which couples statistical link analysis techniques with circuit power estimates based on normalized transistor parameters extracted with a constant current density methodology. The design framework selects the optimum equalization architecture, circuit logic style (CMOS versus current-mode logic), and transmit output swing for minimum I/O power. Analysis shows that low loss channel characteristics and minimal circuit complexity, together with scaling of transmitter output swing allows excellent power efficiency at high data rates.
Arun Palaniappan, Samuel Palermo
IEEE Trans. Very Large Scale Integr. Syst.2
2012 LumiNOC: a power-efficient, high-performance, photonic network-on-chip for future parallel architectures
abstract
Achieving scaling performance as core counts increase to the hundreds in future chip-multi-processors (CMPs) requires high performing, yet energy-efficient interconnects. Silicon nanophotonics is a promising replacement for electronic on-chip interconnect due to its high bandwidth and low latency, however, prior techniques have required high static power for the laser and ring thermal tuning. We propose a novel nano-photonic NoC architecture, LumiNOC, optimized for high performance and power-efficiency. In a 64-node NoC under synthetic traffic, LumiNOC enjoys 50% lower latency at low loads and 40% higher throughput per Watt on synthetic traffic, versus other reported photonic NoCs. LumiNOC reduces latencies 40% versus an electrical 2D mesh NoCs on the PARSEC shared memory, multithreaded benchmark suite.
Mark Browning, Paul Gratz, Samuel Palermo
PACT4