EDBT 2026 Demo / reviewers in the wild / expert
Baogang Zhang
dblp:68/8987
· DBLP profile ↗
15ranked-venue papers
10as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | XMAP: Programming Memristor Crossbars for Analog Matrix-Vector Multiplication: Toward High Precision Using Representable MatricesabstractLinear transformations are the dominating computation within many important applications. The natural multiply-and-accumulate feature of memristor crossbar arrays promise unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix–vector multiplication (MVM). Unfortunately, the precision of the analog computation may be degraded by parasitics, nonlinear device characteristics, and variations. In this article, we propose a framework, called XMAP, for mapping an arbitrary matrix into appropriate memristor conductance values (or state variables for nonlinear devices). The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. XMAP is based on formulating the mapping problem as a mathematical optimization problem, which can be elegantly minimized using the concept of representable matrices, i.e., the matrices that can be represented on a crossbar. Compared to the state-of-the-art conversion algorithm, the computational accuracy is improved with up to$3.29 \times $at the expense of overhead in runtime. The precision improvements translate into noteworthy application-level benefits within signal compression and neural network inference. Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Towards Resilient Deployment of In-Memory Neural Networks with High ThroughputabstractResistive computing systems (RCSs) promise exascale computing capabilities to inference engines for deep learning. However, the classification accuracy of the accelerated neural networks may be degraded by defects. While hardware-aware training schemes can restore the accuracy of convolutional neural networks (CNNs) with low throughput, the schemes are rendered futile when weights are replicated to improve throughput. On the other hand, we discover that weight replication provides new opportunities for data layout organization. In this paper, we propose a framework for resilient deployment of high throughput CNNs to RCSs. The framework is based on integrating a data layout organization step and a distribution guided training step into the flow for mapping CNNs to RCSs. The data layout organization step involves modifying the weight matrix to crossbar assignments using channel, pixel, and hybrid channel-pixel data layout transformations. The distribution guided training is focused on training CNNs with weights that are amenable for data layout organization. The experimental results demonstrate that the proposed techniques expand the average solution space for data layout organization with 1.4× 1014X. This translates into that CNNs with high throughput can be deployed onto RCS with up to 10% defects and still attain high classification accuracy. Baogang Zhang, Rickard Ewetz |
DAC | 1 |
| 2021 | Computational Restructuring: Rethinking Image Compression Using Resistive Crossbar ArraysabstractImage compression is performed on billions of edge devices deployed in the Internet of Things (IoT). The bottleneck of the compression is the 2-D discrete cosine transform (2D DCT), which involves performing two matrix-matrix multiplications in series. Earlier studies have explored directly mapping the 2D DCT computation to emerging resistive crossbar arrays (RCAs), which promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product. The main drawback is that the series computation is inherently vulnerable to errors. In this article, we propose to fundamentally rethink how to perform image compression using RCAs. The key idea is to restructure the computation to natively match the properties of the underlying resistive hardware. This allows three of the main design steps within image compression (2D DCT, quantization, and zig-zag reordering) to be integrated into a single analog MVM operation. The integration is facilitated by the development of a 2D DCT reconstruction technique, a frequency spectrum optimization technique, and a quantization optimization technique. The techniques improve the robustness to errors, eliminates the storage of intermediate data, enables processing of small image blocks, facilitates the utilization of large-scale RCAs, and reduces the requirements on the expensive domain interfaces. Compared with the previous work, the experimental results demonstrate significant improvements in image quality while reducing power and latency with up to 62% and 21%, respectively. Baogang Zhang, Necati Uysal, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Representable Matrices: Enabling High Accuracy Analog Computation for Inference of DNNs using MemristorsabstractAnalog computing based on memristor technology is a promising solution to accelerating the inference phase of deep neural networks (DNNs). A fundamental problem is to map an arbitrary matrix to a memristor crossbar array (MCA) while maximizing the resulting computational accuracy. The state-of-the-art mapping technique is based on a heuristic that only guarantees to produce the correct output for two input vectors. In this paper, a technique that aims to produce the correct output for every input vector is proposed, which involves specifying the memristor conductance values and a scaling factor realized by the peripheral circuitry. The key insight of the paper is that the conductance matrix realized by an MCA is only required to be proportional to the target matrix. The selection of the scaling factor between the two regulates the utilization of the programmable memristor conductance range and the representability of the target matrix. Consequently, the scaling factor is set to balance precision and value range errors. Moreover, a technique of converting conductance values into state variables and vice versa is proposed to handle memristors with non-ideal device characteristics. Compared with the state-of-the-art technique, the proposed mapping results in 4X-9X smaller errors. The improvements translate into that the classification accuracy of a seven-layer convolutional neural network (CNN) on CIFAR-10 is improved from 20.5% to 71.8%. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ASP-DAC | 1 |
| 2020 | Computational Restructuring: Rethinking Image Processing using Memristor Crossbar ArraysabstractImage processing is a core operation performed on billions of sensor-devices in the Internet of Things (IoT). Emerging memristor crossbar arrays (MCAs) promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product, which is the dominating computation within the two-dimensional Discrete Cosine Transform (2D DCT). Earlier studies have directly mapped the digital implementation to MCA based hardware. The drawback is that the series computation is vulnerable to errors. Moreover, the implementation requires the use of large image block sizes, which is known to degrade the image quality. In this paper, we propose to restructure the 2D DCT into an equivalent single linear transformation (or MVM operation). The reconstruction eliminates the series computation and reduces the processed block sizes from N×N to √N×√N Consequently, both the robustness to errors and the image quality is improved. Moreover, the latency, power, and area is reduced with 2X while eliminating the storage of intermediate data, and the power and area can be further reduced with up to 62% and 74% using frequency spectrum optimization. Baogang Zhang, Necati Uysal, Rickard Ewetz |
DATE | 1 |
| 2020 | Redundant Neurons and Shared Redundant Synapses for Robust Memristor-based DNNs with Reduced OverheadabstractThe dominating computational workload in the inference phase of deep neural networks (DNNs) is matrix-vector multiplication. An arising solution to accelerate the inference phase is to perform analog matrix-vector multiplication using memristor crossbar arrays (MCAs). A key challenge is that stuck-at-fault defects may degrade the classification accuracy of the memristor-based DNNs. A common technique to reduce the negative impact of stuck-at-faults is to utilize redundant synapses, i.e, each row in a weight matrix is realized using two (or r) parallel rows in an MCA. In this paper, we propose to handle stuck-at-faults by inserting redundant neurons and by sharing redundant synapses. The first technique is based on inserting redundant neurons to surgically repair neurons connected to rows and columns in the MCAs with many stuck-at-faults. The second technique is focused on sharing redundant synapses between different neurons to reduce the hardware overhead, which generalizes (1:r) synapse redundancy in previous studies to (q:r) synapse redundancy. The experimental results demonstrate new trade-offs between robustness and hardware overhead without requiring the neural networks to be retrained. Compared with state-of-the-art, the power and area overhead for a neural network can be reduced with up to 16% and 25%, respectively. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | DP-MAP: Towards Resistive Dot-Product Engines with Improved PrecisionabstractThe natural multiply and accumulate feature of memristor crossbar arrays promises unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix-vector multiplication. To overcome the challenges of low-precision devices and voltage drop over non-zero array parasitics, each matrix element can be represented using two memristors. In this paper, we propose differential pair map (DP-MAP) - the first matrix to memristor conductance mapping algorithm specifically designed for crossbars with a differential pair configuration. In contrast, previous works consider the differential pair configuration as an afterthought, which limits the achievable precision. The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. Analog computation with high precision is attained by judiciously selecting the conductance range and avoiding to explicitly decompose each matrix into a positive and negative component. Short run-time is achieved using a hierarchical optimization algorithm and two speed-up techniques. Compared with earlier studies, the computational accuracy is improved with 3.36X. This translates into signal and image compression with 61% and 94% higher quality, respectively. The simulation time of complex physical systems modeled using partial differential equations (PDEs) is reduced with 5.87X. Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 2 |
| 2020 | Handling Stuck-at-Fault Defects Using Matrix Transformation for Robust Inference of DNNsabstractMatrix-vector multiplication is the dominating computational workload in the inference phase of deep neural networks (DNNs). Memristor crossbar arrays (MCAs) can efficiently perform matrix-vector multiplication in the analog domain. A key challenge is that memristor devices may suffer stuck-at-fault defects, which can severely degrade the classification accuracy. Earlier studies have shown that the accuracy loss can be recovered by utilizing additional hardware or hardware aware training. In this article, we propose a framework that handles stuck-at-faults using matrix transformations, which is called the MT framework. The framework is based on introducing a cost metric that captures the negative impact of the stuck-atfault defects. Next, the cost metric is minimized by applying matrix transformations T. A transformation T changes a weight matrix W into a new weight matrix W̃ = T(W). In particular, a row flipping transformation, a permutation transformation, and a value range transformation are proposed. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the weight matrices, which results in that the stuck-at-faults introduce smaller errors. The experimental results demonstrate that the MT framework is capable of recovering 99% of the accuracy loss on both the MNIST and CIFAR-10 datasets without utilizing hardware aware training. The accuracy improvements come at the expense of an 8.19× and 9.23× overhead in power and area, respectively. Nevertheless, the overhead can be reduced with up to 50% by leveraging hardware aware training. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Handling stuck-at-faults in memristor crossbar arrays using matrix transformationsabstractMatrix-vector multiplication is the dominating computational workload in the inference phase of neural networks. Memristor crossbar arrays (MCAs) can inherently execute matrix-vector multiplication with low latency and small power consumption. A key challenge is that the classification accuracy may be severely degraded by stuck-at-fault defects. Earlier studies have shown that the accuracy loss can be recovered by retraining each neural network or by utilizing additional hardware. In this paper, we propose to handle stuck-at-faults using matrix transformations. A transformation T changes a weight matrix W into a weight matrix, @ = T(W), which is more robust to stuck-at-faults. In particular, we propose a row flipping transformation, a permutation transformation, and a value range transformation. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the matrix, which results in that each stuck-at-fault introduces an error of smaller magnitude. The experimental results demonstrate that the proposed framework is capable of recovering 99% of the accuracy loss introduced by stuck-at-faults without requiring the neural network to be retrained. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ASP-DAC | 1 |
| 2019 | STAT: Mean and Variance Characterization for Robust Inference of DNNs on Memristor-based PlatformsabstractAn emerging solution to accelerate the inference phase of deep neural networks (DNNs) is to utilize memristor crossbar arrays (MCAs) to perform highly efficient matrix-vector multiplication in the analog domain. An adverse challenge is that memristor devices may suffer stuck-at-fault defects, which may compromise the classification accuracy. Stuck-at-fault defects have previously been handled by neuron permutation or by retraining neural networks. In this paper, we propose the STAT framework that utilizes statistics to guide optimization techniques that provide robustness to stuck-at-fault defects. In particular, bias weights are modified to minimize the input error to each neuron with respect to an input vector. The input vector is selected to be equal to the mean from a statistical characterization. Variance statistics are used to define a weight significance metric, which is used to prioritize assigning weights connected to neurons with large (small) variance to non-defective (defective) memristors using neuron permutation, as errors introduced by neurons with small variance can be eliminated by modifying the bias weights. The experimental results demonstrate that the STAT framework improves the normalized classification accuracy from 62.1% to 96.1% without any hardware overhead. Baogang Zhang, Necati Uysal, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Software and Hardware Techniques for Reducing the Impact of Quantization Errors in Memristor Crossbar ArraysabstractMatrix-vector multiplication is the dominating computational workload in the evaluation of neural networks. It has recently been demonstrated that memristor crossbar arrays (MCAs) can perform matrix-vector multiplication with small power consumption and low latency. However, the computational accuracy may be degraded by quantization errors. The quantization errors of mapping a matrix to an MCA are proportional to the the number of distinguishable states of each memristor and the difference between the largest and smallest element in the matrix. In this paper, we propose a framework for mapping an arbitrary matrix A to a grid of MCAs (or a single MCA) while minimizing the negative impact of quantization errors. The framework is guided by a total quantization error bound (TQEB) metric, which is an upper bound on the total quantization errors (TQE). Using the proposed TQEB metric, three techniques of reducing TQE are proposed. The first method is based on scaling and shifting the rows in A with different factors to improve the memristor conductance band utilization. The second technique is based on representing a single column in a matrix A using multiple columns in an MCA, to reduce the magnitude of the smallest and largest elements in A. The third technique is based on permuting the order of the columns in A when the matrix A is required to be mapped to a grid of MCAs. The quantization errors are reduced by assigning matrix values of similar magnitude to the same MCAs in the grid. The experimental results demonstrate that the proposed metric and techniques are capable of greatly reducing the negative impact of quantization errors. Baogang Zhang, Rickard Ewetz |
ICCD | 1 |
| 2013 | An error calibration method based on BPANN algorithm for three-axis magnetometersabstractDue to the effects of nonorthogonality, scale factor and zero-shifts, Three-axis Fluxgate Magnetometers (TFMs) with the attitude changes have certain errors. Therefore, it's quite necessary to establish a calibration method to get high precision magnetic information. Back Propagation Artificial Neural Network (BPANN) algorithm is used to correct the system errors of measured data in this paper. The results show that the BPANN method can well calibrate the data error of TFMs. The measuring error of geomagnetic field decreases obviously after calibration. Lei Jiang 0012, Baogang Zhang |
IGARSS | 3 |
| 2011 | A simplified aeromagnetic compensation model for low magnetism UAV platformabstractMagnetic interference generated by the air-borne platform is one of the most common system noises in aero-magnetic survey. The noise can be divided into 3 parts as the permanent, the induced and the eddy-current fields. A new air-borne platform called the Unmanned Aerial Vehicle (UAV) can be designed and manufactured with low magnetic interference. Analyzing the test measured data has shown that the UAV, compare with the traditional platform, has lower magnetic noises especially in eddy-current fields. Thus, the eddy-current terms may be ignored in magnetic compensation model. The simulation result shows that the new model reduced the multi-colinearity effect and can be useful in magnetic compensation for low magnetism UAV platform. Baogang Zhang, Yanchao Qiao |
IGARSS | 1 |
| 2010 | Using local transition probability models in Markov Random Field for multi-temporal image classificationabstractMaking use full of multi-source and multi-temporal information to extract richer and interesting information is a tendency in analysis of remote sensing images. In this paper, spatial and temporal contextual classification based on Markov Random Field (MRF) is used to classify ecological function vegetation in Poyang Lake. The results show that spatial and temporal neighborhood complementary information from different images can be used to remove the spectral confusion of different kinds of vegetation on single image and improve classification accuracy compared to MLC method. The local transition model is more accurate than global transition model and also effective in computation. Building effective spatial and temporal neighborhood model for information extraction in special application is the key of multi-source and multi-temporal image analysis. Although spatial and temporal contextual classification method is computation demanding, it's promising in the application emphasizing classification accuracy. Caixia Liu 0002, Baogang Zhang |
IGARSS | 5 |
| 2010 | A simplified image fusion technique with sensor spectral responseabstractIntroducing the sensor spectral response into the wavelet-based image fusion methods may produce images closer to which obtained by the ideal sensor. But the wavelet-based image fusion methods are complex because of the wavelet decomposition. This paper presents a simplified image fusion method using sensor spectral response which derives from a wavelet-based one, but need not to do the wavelet decomposition at the last calculation. The experimental results demonstrate that it provides good performance both in processing speed and image fusion quality. Peng Gong 0002, Caixia Liu 0002, Baogang Zhang |
IGARSS | 6 |