Necati Uysal

dblp:214/9892 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
4since 2021 · last 2022
0000-0002-9543-3823ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2022 XMAP: Programming Memristor Crossbars for Analog Matrix-Vector Multiplication: Toward High Precision Using Representable Matrices
abstract
Linear transformations are the dominating computation within many important applications. The natural multiply-and-accumulate feature of memristor crossbar arrays promise unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix–vector multiplication (MVM). Unfortunately, the precision of the analog computation may be degraded by parasitics, nonlinear device characteristics, and variations. In this article, we propose a framework, called XMAP, for mapping an arbitrary matrix into appropriate memristor conductance values (or state variables for nonlinear devices). The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. XMAP is based on formulating the mapping problem as a mathematical optimization problem, which can be elegantly minimized using the concept of representable matrices, i.e., the matrices that can be represented on a crossbar. Compared to the state-of-the-art conversion algorithm, the computational accuracy is improved with up to$3.29 \times $at the expense of overhead in runtime. The precision improvements translate into noteworthy application-level benefits within signal compression and neural network inference.
Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Synthesis of Clock Networks with a Mode-Reconfigurable Topology
abstract
Modern digital circuits are often required to operate in multiple modes to cater to variable frequency and power requirements. Consequently, the clock networks for such circuits must be synthesized, meeting different timing constraints in different operational modes. The overall power consumption and robustness to variations of a clock network are determined by the topology. However, state-of-the-art clock networks use the same topology in every mode, despite that timing constraints in low- and high-performance modes can be very different. In this article, we propose a clock network with a mode-reconfigurable topology (MRT) for circuits with positive-edge-triggered sequential elements. In high-performance modes, the MRT structure is reconfigured into a near-tree to provide the required robustness to variations. In low-performance modes, the MRT structure is reconfigured into a tree to save power. Non-tree (or near-tree) structures provide robustness to variations by appropriately constructing multiple alternative paths from the clock source to the clock sinks, which neutralizes the negative impact of variations. In MRT structures, OR-gates are used to join multiple alternative paths into a single path. Hence, the MRT structures consume no short-circuit power because there is only one gate driving each net. Moreover, it is straightforward to reconfigure an MRT structure into a tree topology using a single clock gate. In high-performance modes, the experimental results demonstrate that MRT structures have \( 25\% \) lower power consumption than state-of-the-art near-tree structures. In low-performance modes, the power consumption of the MRT structure is similar to the power consumption of a clock tree.
Necati Uysal, Rickard Ewetz
ACM Trans. Design Autom. Electr. Syst.1
2021 An OCV-Aware Clock Tree Synthesis Methodology
abstract
Closing timing after clock tree synthesis (CTS) is very challenging in the presence of on-chip variations (OCVs). State-of-the-art design flows first synthesize an initial clock tree that contains timing violations introduced by OCVs. Next, aggressive clock tree optimization (CTO) is applied to eliminate the timing violations. Unfortunately, it may be impossible to eliminate all violations given the structure of the initial clock tree. In this paper, we propose an OCV-aware clock tree synthesis methodology that aims to rethink how to account for OCVs. The key idea is to predict the impact of OCVs early in the synthesis process, which allows the variations to be compensated for using non-uniform safety margins. This results in a synthesis flow that is almost correct-by-design. In contrast, state-of-the-art design flows often have an unpredictable success rate because the OCVs are considered too late in the synthesis process. Concretely, this is achieved by top-down constructing a virtual clock tree that is refined bottom-up into a real clock tree implementation. To balance the quality of results (QoR) and runtime, multiple top-level tree topologies are enumerated and pruned in the synthesis process. Compared with the CTO based approach, the experimental results demonstrate that the proposed methodology reduces the total negative slack (TNS) and worst negative slack (WNS) with 90% and 75%, respectively.
Necati Uysal, Rickard Ewetz
ICCAD1
2021 Computational Restructuring: Rethinking Image Compression Using Resistive Crossbar Arrays
abstract
Image compression is performed on billions of edge devices deployed in the Internet of Things (IoT). The bottleneck of the compression is the 2-D discrete cosine transform (2D DCT), which involves performing two matrix-matrix multiplications in series. Earlier studies have explored directly mapping the 2D DCT computation to emerging resistive crossbar arrays (RCAs), which promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product. The main drawback is that the series computation is inherently vulnerable to errors. In this article, we propose to fundamentally rethink how to perform image compression using RCAs. The key idea is to restructure the computation to natively match the properties of the underlying resistive hardware. This allows three of the main design steps within image compression (2D DCT, quantization, and zig-zag reordering) to be integrated into a single analog MVM operation. The integration is facilitated by the development of a 2D DCT reconstruction technique, a frequency spectrum optimization technique, and a quantization optimization technique. The techniques improve the robustness to errors, eliminates the storage of intermediate data, enables processing of small image blocks, facilitates the utilization of large-scale RCAs, and reduces the requirements on the expensive domain interfaces. Compared with the previous work, the experimental results demonstrate significant improvements in image quality while reducing power and latency with up to 62% and 21%, respectively.
Baogang Zhang, Necati Uysal, Rickard Ewetz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Representable Matrices: Enabling High Accuracy Analog Computation for Inference of DNNs using Memristors
abstract
Analog computing based on memristor technology is a promising solution to accelerating the inference phase of deep neural networks (DNNs). A fundamental problem is to map an arbitrary matrix to a memristor crossbar array (MCA) while maximizing the resulting computational accuracy. The state-of-the-art mapping technique is based on a heuristic that only guarantees to produce the correct output for two input vectors. In this paper, a technique that aims to produce the correct output for every input vector is proposed, which involves specifying the memristor conductance values and a scaling factor realized by the peripheral circuitry. The key insight of the paper is that the conductance matrix realized by an MCA is only required to be proportional to the target matrix. The selection of the scaling factor between the two regulates the utilization of the programmable memristor conductance range and the representability of the target matrix. Consequently, the scaling factor is set to balance precision and value range errors. Moreover, a technique of converting conductance values into state variables and vice versa is proposed to handle memristors with non-ideal device characteristics. Compared with the state-of-the-art technique, the proposed mapping results in 4X-9X smaller errors. The improvements translate into that the classification accuracy of a seven-layer convolutional neural network (CNN) on CIFAR-10 is improved from 20.5% to 71.8%.
Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz
ASP-DAC2
2020 Computational Restructuring: Rethinking Image Processing using Memristor Crossbar Arrays
abstract
Image processing is a core operation performed on billions of sensor-devices in the Internet of Things (IoT). Emerging memristor crossbar arrays (MCAs) promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product, which is the dominating computation within the two-dimensional Discrete Cosine Transform (2D DCT). Earlier studies have directly mapped the digital implementation to MCA based hardware. The drawback is that the series computation is vulnerable to errors. Moreover, the implementation requires the use of large image block sizes, which is known to degrade the image quality. In this paper, we propose to restructure the 2D DCT into an equivalent single linear transformation (or MVM operation). The reconstruction eliminates the series computation and reduces the processed block sizes from N×N to √N×√N Consequently, both the robustness to errors and the image quality is improved. Moreover, the latency, power, and area is reduced with 2X while eliminating the storage of intermediate data, and the power and area can be further reduced with up to 62% and 74% using frequency spectrum optimization.
Baogang Zhang, Necati Uysal, Rickard Ewetz
DATE2
2020 Redundant Neurons and Shared Redundant Synapses for Robust Memristor-based DNNs with Reduced Overhead
abstract
The dominating computational workload in the inference phase of deep neural networks (DNNs) is matrix-vector multiplication. An arising solution to accelerate the inference phase is to perform analog matrix-vector multiplication using memristor crossbar arrays (MCAs). A key challenge is that stuck-at-fault defects may degrade the classification accuracy of the memristor-based DNNs. A common technique to reduce the negative impact of stuck-at-faults is to utilize redundant synapses, i.e, each row in a weight matrix is realized using two (or r) parallel rows in an MCA. In this paper, we propose to handle stuck-at-faults by inserting redundant neurons and by sharing redundant synapses. The first technique is based on inserting redundant neurons to surgically repair neurons connected to rows and columns in the MCAs with many stuck-at-faults. The second technique is focused on sharing redundant synapses between different neurons to reduce the hardware overhead, which generalizes (1:r) synapse redundancy in previous studies to (q:r) synapse redundancy. The experimental results demonstrate new trade-offs between robustness and hardware overhead without requiring the neural networks to be retrained. Compared with state-of-the-art, the power and area overhead for a neural network can be reduced with up to 16% and 25%, respectively.
Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz
ACM Great Lakes Symposium on VLSI2
2020 DP-MAP: Towards Resistive Dot-Product Engines with Improved Precision
abstract
The natural multiply and accumulate feature of memristor crossbar arrays promises unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix-vector multiplication. To overcome the challenges of low-precision devices and voltage drop over non-zero array parasitics, each matrix element can be represented using two memristors. In this paper, we propose differential pair map (DP-MAP) - the first matrix to memristor conductance mapping algorithm specifically designed for crossbars with a differential pair configuration. In contrast, previous works consider the differential pair configuration as an afterthought, which limits the achievable precision. The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. Analog computation with high precision is attained by judiciously selecting the conductance range and avoiding to explicitly decompose each matrix into a positive and negative component. Short run-time is achieved using a hierarchical optimization algorithm and two speed-up techniques. Compared with earlier studies, the computational accuracy is improved with 3.36X. This translates into signal and image compression with 61% and 94% higher quality, respectively. The simulation time of complex physical systems modeled using partial differential equations (PDEs) is reduced with 5.87X.
Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz
ICCAD1
2020 Synthesis of Clock Networks with a Mode Reconfigurable Topology and No Short Circuit Current
abstract
Circuits deployed in the Internet of Things operate in low and high performance modes to cater to variable frequency and power requirements. Consequently, the clock networks for such circuits must be synthesized meeting drastically different timing constraints under variations in the different modes. The overall power consumption and robustness to variations of a clock network is determined by the topology. However, state-of-the-art clock networks use the same topology in every mode, despite that the timing constraints in the low and high performance modes are very different. In this paper, we propose a clock network with a mode reconfigurable topology (MRT) for circuits with positive-edge triggered sequential elements. In high performance modes, the required robustness to variations is provided by reconfiguring the MRT structure into a near-tree. In low performance modes, the MRT structure is reconfigured into a tree to save power. Non-tree (or near-tree) structures provide robustness to variations by appropriately constructing multiple alternative paths from the clock source to the clock sinks, which neutralizes the negative impact of variations. In MRT structures, OR-gates are used to join multiple alternative paths into a single path. Consequently, the MRT structures consume no short circuit power because there is only one gate driving each net. Moreover, it is straightforward to reconfigure MRT structures into a tree by gating the clock signal in part of the structure. Compared with state-of-the-art near-tree structures, MRT structures have 8% lower power consumption and similar robustness to variations in high performance modes. In low performance modes, the power consumption is 16% smaller when reconfiguration is used.
Necati Uysal, Juan Ariel Cabrera, Rickard Ewetz
ISPD1
2020 Handling Stuck-at-Fault Defects Using Matrix Transformation for Robust Inference of DNNs
abstract
Matrix-vector multiplication is the dominating computational workload in the inference phase of deep neural networks (DNNs). Memristor crossbar arrays (MCAs) can efficiently perform matrix-vector multiplication in the analog domain. A key challenge is that memristor devices may suffer stuck-at-fault defects, which can severely degrade the classification accuracy. Earlier studies have shown that the accuracy loss can be recovered by utilizing additional hardware or hardware aware training. In this article, we propose a framework that handles stuck-at-faults using matrix transformations, which is called the MT framework. The framework is based on introducing a cost metric that captures the negative impact of the stuck-atfault defects. Next, the cost metric is minimized by applying matrix transformations T. A transformation T changes a weight matrix W into a new weight matrix W̃ = T(W). In particular, a row flipping transformation, a permutation transformation, and a value range transformation are proposed. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the weight matrices, which results in that the stuck-at-faults introduce smaller errors. The experimental results demonstrate that the MT framework is capable of recovering 99% of the accuracy loss on both the MNIST and CIFAR-10 datasets without utilizing hardware aware training. The accuracy improvements come at the expense of an 8.19× and 9.23× overhead in power and area, respectively. Nevertheless, the overhead can be reduced with up to 50% by leveraging hardware aware training.
Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Latency constraint guided buffer sizing and layer assignment for clock trees with useful skew
abstract
Closing timing using clock tree optimization (CTO) is a tremendously challenging problem that may require designer intervention. CTO is performed by specifying and realizing delay adjustments in an initially constructed clock tree. Delay adjustments are typically realized by inserting delay buffers or detour wires. In this paper, we propose a latency constraint guided buffer sizing and layer assignment framework for clock trees with useful skew, called the (BLU) framework. The BLU framework realizes delay adjustments during CTO by performing buffer sizing and layer assignment. Given an initial clock tree, the BLU framework first predicts the final timing quality and specifies a set of delay adjustments, which are translated into latency constraints. Next, buffer sizing and layer assignment is performed with respect to the latency constraints using an extension of van Ginneken's algorithm. Moreover, the framework includes a feature of reducing the power consumption by relaxing the latency constraints and a method of improving the timing performance by tightening the latency constraints. The experimental results demonstrate that the proposed framework is capable of reducing the capacitive cost with 13% on the average. The total negative slack (TNS) and worst negative slack (WNS) are reduced with up to 58% and 20%, respectively.
Necati Uysal, Wen-Hao Liu 0001, Rickard Ewetz
ASP-DAC1
2019 Handling stuck-at-faults in memristor crossbar arrays using matrix transformations
abstract
Matrix-vector multiplication is the dominating computational workload in the inference phase of neural networks. Memristor crossbar arrays (MCAs) can inherently execute matrix-vector multiplication with low latency and small power consumption. A key challenge is that the classification accuracy may be severely degraded by stuck-at-fault defects. Earlier studies have shown that the accuracy loss can be recovered by retraining each neural network or by utilizing additional hardware. In this paper, we propose to handle stuck-at-faults using matrix transformations. A transformation T changes a weight matrix W into a weight matrix, @ = T(W), which is more robust to stuck-at-faults. In particular, we propose a row flipping transformation, a permutation transformation, and a value range transformation. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the matrix, which results in that each stuck-at-fault introduces an error of smaller magnitude. The experimental results demonstrate that the proposed framework is capable of recovering 99% of the accuracy loss introduced by stuck-at-faults without requiring the neural network to be retrained.
Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz
ASP-DAC2
2019 STAT: Mean and Variance Characterization for Robust Inference of DNNs on Memristor-based Platforms
abstract
An emerging solution to accelerate the inference phase of deep neural networks (DNNs) is to utilize memristor crossbar arrays (MCAs) to perform highly efficient matrix-vector multiplication in the analog domain. An adverse challenge is that memristor devices may suffer stuck-at-fault defects, which may compromise the classification accuracy. Stuck-at-fault defects have previously been handled by neuron permutation or by retraining neural networks. In this paper, we propose the STAT framework that utilizes statistics to guide optimization techniques that provide robustness to stuck-at-fault defects. In particular, bias weights are modified to minimize the input error to each neuron with respect to an input vector. The input vector is selected to be equal to the mean from a statistical characterization. Variance statistics are used to define a weight significance metric, which is used to prioritize assigning weights connected to neurons with large (small) variance to non-defective (defective) memristors using neuron permutation, as errors introduced by neurons with small variance can be eliminated by modifying the bias weights. The experimental results demonstrate that the STAT framework improves the normalized classification accuracy from 62.1% to 96.1% without any hardware overhead.
Baogang Zhang, Necati Uysal, Rickard Ewetz
ACM Great Lakes Symposium on VLSI2
2018 OCV guided clock tree topology reconstruction
abstract
The timing performance of clock trees in scaled technology nodes may be severely degraded by on-chip variations (OCV). Clock tree optimization (CTO) is employed to eliminate timing violations by specifying a set of non-negative delay adjustments using a linear programming (LP) formulation. Next, the delay adjustments are realized in the clock tree by inserting delay buffers and detour wires. The drawback is that given the topology of the initial clock tree, it may be impossible to remove all timing violations. In this paper, a framework that performs OCV guided clock tree topology reconstruction is proposed. The framework reconstructs the topology of a clock tree while improving the lower bounds on the worst negative slack (WNS) and the total negative slack (TNS). Next, traditional CTO is employed to reduce WNS and TNS to the improved lower bounds. The reconstruction of the clock tree topology is guided by a predicted leaf buffer slack graph (pLB-SG). The leaf buffers that must be placed closer in the tree topology are identified by detecting cycles (or strongly connected components) in the pLB-SG. The experimental results demonstrate that the proposed framework can on the average reduce WNS and TNS with 84% and 80%, respectively.
Necati Uysal, Rickard Ewetz
ASP-DAC1