EDBT 2026 Demo / reviewers in the wild / expert
Jie Han 0001
dblp:09/2621-1
· DBLP profile ↗
106ranked-venue papers
6as first author
34since 2021 · last 2026
0000-0002-8849-4994ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 96 · 6 first-author · 29 since 2021Software engineering, systems software and programming languages · 18 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-ANT: Efficient Low-Bit Group-Wise Quantization for Large Language Models via Sign-Asymmetric Adaptive Numeric TypeabstractLarge language models (LLMs) have demonstrated remarkable potential across diverse domains; meanwhile, their large parameter sizes pose substantial inference costs, motivating the need for efficient low-bit quantization. Group-wise quantization, which adopts finer granularity, has been widely used to improve low-bit quantization performance. Several adaptive numeric types have been proposed to further enhance low-bit group-wise quantization; however, they construct quantization grids based on symmetric numeric types, which limits their ability to model asymmetric distributions. To address this limitation, we propose SA-ANT, a sign-asymmetric adaptive numeric type for efficient low-bit group-wise quantization. SA-ANT constructs quantization grids separately on the positive and negative sides, enabling adaptive support for asymmetric and non-uniform distributions. Furthermore, the carefully designed SA-ANT not only reduces quantization errors but also ensures a unified computing across different sub numeric types, thereby facilitating hardware efficiency. To accelerate LLM inference, we develop (1) a quantization framework that transforms LLM weights into the SA-ANT and adaptively selects the sub numeric type for each group, and (2) an accelerator that maps SA-ANT inference to low-bit INT operations. Experimental results show that SA-ANT delivers 3.92%– 5.57% higher accuracy than state-of-the-art adaptive numeric types under 3-bit weight quantization, while also enabling 7.84%– 44.65% area savings and 7.80%–43.88% power reductions. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 4 |
| 2026 | Bridging the Power Estimation Gap: A GNN-Based Prediction Model for Approximate Logic SynthesisabstractApproximate computing is an application-related paradigm that trades limited accuracy for improvements in hardware cost. As a key technique of approximate computing, approximate logic synthesis (ALS) automatically generates approximate circuits with reduced area, power, and delay while satisfying predefined quality-of-result (QoR) constraints. However, in typical gate-level ALS workflows, synthesis tools are invoked at the final stage for optimization, leading to a discrepancy between the circuit in design space exploration (DSE) and the final obtained circuit. Thus, the power estimation for a candidate circuit during DSE may exhibit a significant gap from the actual power consumed by its post-synthesis circuit. This gap may mislead the DSE to a sub-optimal design. To address this issue, we propose a graph neural network (GNN)-based power prediction model that operates on gate-level circuits. The model incorporates multi-head channel attention, which extracts high-level topological and functional features that correlate with power dissipation and implicitly captures the optimization behavior of synthesis tools. Thus, it enables a direct prediction of post-synthesis power from pre-synthesis gate-level circuits. Experimental results show that the proposed model improves the concordance index (C-index) for power ranking by up to 14.0% over traditional methods. Furthermore, we construct an ALS framework by integrating the proposed model with Cartesian genetic programming (CGP). Compared to state-of-the-art ALS approaches, our GNN-CGP framework generates circuits with up to 26.8% power savings under the same error constraints. Fuxuan Li, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 5 |
| 2026 | HAP: Accelerating DNNs with Resolution-Preserved Quantization by Harnessing Adaptive-PrecisionabstractReducing the precision in post-training quantization can cause catastrophic accuracy loss in Deep Neural Networks, especially when compressing the activations. To address this problem, we present a novel adaptive-precision quantization (APQ) and accelerator design that achieves lossless activation compression by exploiting the inherent coding redundancy. Compared to existing APQ methods, this design can be generalized to implement asymmetric quantization, making it particularly suitable for activations. The accelerator offers a practical solution to mitigate the computational workload imbalance problem incurred by variable precision. A dual-precision quantization scheme further provides the flexibility to trade off accuracy and performance. Erjing Luo, Xinkuang Geng, Honglan Jiang, Leibo Liu, Jie Han 0001 |
DATE | 5 |
| 2026 | Model-based speech enhancement with spectral envelope correction using stacked autoencoders
Wenhao Lu, Zhenya Zang, Xia Dong, Jie Han 0001, Zuozhou Pan, Yiping Ke |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | LUT-ALMs: Trading Off Accuracy and Power for Approximate Logarithmic Multipliers via LUT OptimizationabstractLogarithmic multiplier (LM) converts fixed-point (FxP) input operands to logarithmic numbers and performs multiplication with simple shift and addition operations, which achieves distinct power reduction, yet with significant single-sided errors. This paper proposes to fuse error compensation with logarithmic conversion by using customized look-up tables (LUTs). To avoid the use of large LUTs, partition strategies are designed for the optimization of LUTs. In addition, to effectively balance the accuracy and hardware costs, two iterative algorithms are proposed for generating precision-configurable LUTs. Based on the optimized LUTs, high-accuracy and low-power approximate LMs (LUT-ALMs) are constructed for 8-bit and 16-bit multiplications. Furthermore, to enhance the flexibility of data type and bit width, a mixed-mode LM (MM-ALM) supporting eight multiplication modes is devised. Compared with the exact 8-bit signed multiplier from Synopsys DesignWare (DW) library (DW-Exact), LUT-ALMs show up to 30.96% and 23.99% reductions in the power-delay product (PDP) and area-delay product (ADP), respectively, with a mean relative error distance (MRED) of 3.45%. Compared with state-of-the-art approximate 8- bit signed multipliers, LUT-ALMs form the Pareto front in terms of PDP and MRED. For 16-bit multiplication, LUT-ALMs obtain up to 70.90% and 62.13% savings in PDP and ADP with a MRED of 2.75%, compared with the corresponding DW-Exact. Compared with the corresponding mixed-mode exact multiplier constructed of DW multipliers, MM-ALM performing 8-bit multiplication achieves up to 57.44% and 47.99% reductions in PDP and ADP, respectively. When performing 4-bit multiplications, MM-ALM can save up to 37.16% savings in PDP. With lower hardware over-heads, LUT-ALMs and MM-ALM present comparable accuracy to the corresponding exact designs in the considered convolutional neural networks (CNNs) and image processing applications. The hardware description of the devised LMs is open-sourced athttps://anonymous.4open.science/r/LUT-ALM-5FD1. Xinkuang Geng, Xiaolu Hu, Hui Wang 0023, Jianfei Jiang 0001, Qin Wang 0009, Siting Liu 0001, Jie Han 0001, Honglan Jiang |
IEEE Trans. Computers | 8 |
| 2025 | Lookup Table Refactoring: Towards Efficient Logarithmic Number System Addition for Large Language ModelsabstractCompared to integer quantization, logarithmic quantization aligns more effectively with the long-tailed distribution of data in large language models (LLMs), resulting in lower quantization errors. Moreover, the logarithmic number system (LNS) employs a fixed-point adder to perform multiplication, indicating a potential reduction in computational complexity for LLM accelerators that require extensive multiply-accumulate (MAC) operations. However, a key bottleneck is that LNS addition requires complex nonlinear functions, which are typically approximated using lookup tables (LUTs). This study aims to reduce the hardware resources needed for LUTs in LNS addition while maintaining high precision. Specifically, we investigate the specific nature of addition operations within LLMs; the relationship between the hardware parameters of the LUT and the computing errors is then mathematically derived. Based on these insights, we propose LUT refactoring to optimize the LUT for enhanced efficiency in LNS addition. With 10.93% and 19.78% reductions in area-delay product (ADP) and power-delay product (PDP), respectively, LUT refactoring results in an accuracy improvement of up to 33.5% in LLM benchmarks compared to the naive design. When compared to integer quantization, our method achieves higher accuracy while reducing area by 18.27% and power by 42.61%. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 4 |
| 2025 | A Low-Power Mixed-Precision Integrated Multiply-Accumulate Architecture for Quantized Deep Neural NetworksabstractAs mixed-precision quantization techniques have been widely considered for balancing computational efficiency and flexibility in quantized deep neural networks (DNNs), mixed-precision multiply-accumulate (MAC) units are increasingly important in DNN accelerators. However, conventional mixed-precision MAC architectures support either signed × signed or unsigned ×unsigned multiplications. The signed ×unsigned multiplication enhancing the computing efficiency of DNNs with ReLU activations has never been considered in the design of mixed-precision MAC. Thus, this work proposes a mixed-precision MAC architecture supporting six operation modes, int8 × int8, int8 × uint8, two int4 × int4, two int4 × uint4, four int2 × int2, and four int2 × uint2. In this design, to balance the power and delay of different modes, the multiplication is implemented based on four precision-split 4×4 multipliers (PS4Ms). The accumulation is integrated into the partial product accumulation of the multiplication to eliminate redundant switching activities in separate compression. With 10% area reduction, the proposed MAC denoted as PS4MAC, reduces the power by over 35%, 42%, and 56% for 8-bit, 4-bit, and 2-bit operations, respectively, compared with the design based on the Synopsys DesignWare (DW) multipliers. Additionally, it achieves over 23% power savings for 8-bit operations compared to state-of-the-art (SotA) mixed-precision MAC designs. To save more power, an approximate computing mode for 8-bit multiplication is further designed, resulting in a MAC unit enabling eight operation modes, referred to as PS4MAC_AP. Finally, output-stationary systolic arrays (SAs) are explored using the above-mentioned MAC designs to implement DNNs operating under a 1 GHz clock. Our designs show the highest energy efficiency and outstanding area efficiency in all 8-bit, 4-bit, and 2-bit operation modes. Compared with the traditional SA with high-precision-split multipliers, PS4MAC_AP improves the energy efficiency for 8-bit operations by 0.6 TOPS/W, and PS4MAC achieves 0.4 TOPS/W - 0.7 TOPS/W improvement for all operation modes. Xiaolu Hu, Xinkuang Geng, Zhigang Mao, Jie Han 0001, Honglan Jiang |
DATE | 4 |
| 2025 | An SRAM-based Stochastic Number Generator for Stochastic ComputingabstractStochastic computing (SC) features a unique number representation, where real values are encoded by the probability of "1"s in a random binary bit stream or a stochastic sequence. It enables hardware-efficient arithmetic circuit designs with simple logic gates. However, stochastic number generators (SNGs) are required to produce stochastic sequences. The high hardware cost of an SNG offsets the advantage of SC. To reduce the hardware cost of an SNG, we propose an SRAM-based SNG using voltage under-scaling. It generates random bits by leveraging the access instability of selected SRAM cells, induced by a reduced supply voltage. It is suitable for energy-efficient SC. We implemented the SRAM-based SNG on a Xilinx ZC702 FPGA using block RAMs and evaluated its performance across multiple SC applications, including finite-state machine-based tanh function generation and an Ising machine that solves max-cut problems (MCPs). For the tanh function, our design achieves a comparable mean-squared error (MSE) (8.7 × 10−3) compared to the use of traditional SNGs, such as Sobol- (4.67 × 10−2) and linear feedback shift register (LFSR)-based (1.53 × 10−2) ones. For MCPs, a maximum cut value comparable to that of a cutting-edge design is achieved. Compared with LFSR- and Sobol-based designs, the proposed design consumes 84.3% and 92.5% less energy, respectively. Heng Shi 0003, Zhengkun Yu, Jie Han 0001, Siting Liu 0001 |
ISCAS | 4 |
| 2024 | QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceabstractWhile exhibiting superior performance in many tasks, vision transformers (ViTs) face challenges in quantization. Some existing low-bit-width quantization techniques cannot effectively cover the whole inference process of ViTs, leading to an additional memory overhead (22.3%-172.6%) compared with corresponding fully quantized models. To address this issue, we propose quadruplet uniform quantization (QUQ) to deal with data of various distributions in ViT. QUQ divides the entire data range into at most four subranges that are uniformly quantized with different scale factors. To determine the partition scheme and quantization parameters, an efficient relaxation algorithm is proposed accordingly. Moreover, dedicated encoding and decoding strategies are devised to facilitate the design of an efficient accelerator. Experimental results show that QUQ surpasses state-of-the-art quantization techniques; it is the first viable scheme that can fully quantize ViTs to 6-bit with acceptable accuracy. Compared with conventional uniform quantization, QUQ leads to not only a higher accuracy but also an accelerator with lower area and power. Xinkuang Geng, Siting Liu 0001, Leibo Liu, Jie Han 0001, Honglan Jiang |
DAC | 4 |
| 2024 | Efficient Approximate Decomposition Solver using Ising ModelabstractComputing with memory is an energy-efficient computing approach. It pre-computes a function and stores its values in a lookup table (LUT), which can be retrieved at runtime. Approximate Boolean decomposition reduces the LUT size for implementing complex functions, but it takes a long time to find a decomposition with a minimal error. In this work, to address this issue, we propose an efficient Ising model-based approximate Boolean decomposition solver. First, a new column-based approximate disjoint decomposition method is proposed to fit the Ising model. Then, it is adapted to the Ising model-based optimization solver. Moreover, two improvement techniques are developed for an efficient search of the approximate disjoint decomposition when using simulated bifurcation to solve the Ising model. Experimental results show that compared to the state-of-the-art work, our approach achieves a 11% smaller mean error distance with an average 1.16× speedup when approximately decomposing 16-input Boolean functions. Weihua Xiao, Xingyue Qian, Jie Han 0001, Weikang Qian |
DAC | 4 |
| 2024 | A High-Performance Stochastic Simulated Bifurcation Ising MachineabstractIsing model-based computers, or Ising machines, have recently emerged as high-performance solvers for combinatorial optimization problems (COPs). A simulated bifurcation (SB) Ising machine searches for the solution by solving pairs of differential equations related to the oscillator positions and momenta. It benefits from massive parallelism but suffers from high energy. As an unconventional computing paradigm, dynamic stochastic computing implements accumulation-based operations efficiently. By exploiting the advantages in algorithm and hardware codesign, this article proposes a high-performance stochastic SB machine (SSBM) with efficient hardware. To this end, we develop a stochastic SB (sSB) algorithm such that the multiply-and-accumulate (MAC) operation is converted to multiplexing and addition while the numerical integration is implemented by using signed stochastic integrators (SSIs). Specifically, the sSB stochastically ternarizes position values used for the MAC operation. Two types of SB cells are constructed. A stochastic computing SB cell contains two SSIs with a high area efficiency, while a binary-stochastic computing SB cell contains one binary integrator and one SSI with a reduced delay. Based on sSB, an SSBM is then built by using the proposed SB cells as the basic building block. The designs and syntheses of two SSBMs with 2000 fully connected spins require at least 10.62% smaller area than the state-of-the-art designs. It shows the potential of stochastic computing for SB to efficiently solve COPs. Hongqiao Zhang, Zhengkun Yu, Siting Liu 0001, Jie Han 0001 |
DAC | 5 |
| 2024 | A Low-Power and High-Accuracy Approximate Adder for Logarithmic Number SystemabstractThe Logarithmic Number System (LNS) exploits the non-uniform distribution of data in convolutional neural networks (CNNs), so it leads to a high accuracy for image classification. An LNS provides an easier way to implement complex operations such as multiplication and division. However, addition and subtraction in the LNS require huge hardware resources due to the involved nonlinear operations. To mitigate this problem, we design a low-power approximate logarithmic adder with high-accuracy. Initially, a compact piecewise linear approximation (CPLA) algorithm is proposed to approximately compute the binary exponentiation and logarithm. Implemented by using simple circuits, the CPLA algorithm results in higher accuracy than the classical Mitchell’s algorithm. Consequently, three approximate logarithmic adders are devised, denoted as LA_CPLA1, LA_CPLA2, and LA_CPLA3. Compared with the logarithmic adder design based on lookup tables, the proposed LA_CPLA3 with a configuration of (e, f, n) = (7, 6, 3) achieves 35.05% and 39.80% reductions in area and power dissipation respectively, with a 0.01% mean relative error distance (MRED). We define (e, f) as the bit width of the logarithmic adder, where e and f are the bit widths of the integer and fractional parts, respectively. n is the approximate LSBs in the proposed LA_CPLAs processed by using OR gates. Compared with the multiply and accumulate (MAC) unit in a conventional system using fixed-point numbers, the MAC in the LNS using the proposed LA_CPLAs achieve a lower power by 5.96% to 32.02%, and a smaller area by 6.48% to 32.40%. To assess the efficiency of the proposed approximate adders, they are applied to the implementations of two image processing and CNN applications. The simulation results show that LA_CPLAs result in marginal accuracy loss compared to the corresponding accurate implementations. Xinkuang Geng, Qin Wang 0009, Jie Han 0001, Honglan Jiang |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Learning the Error Features of Approximate Multipliers for Neural Network ApplicationsabstractApproximate multipliers (AMs) have widely been investigated to pursue high-performance and energy-efficient hardware designs for error-tolerant applications, such as neural networks (NNs). The computing accuracy of an AM has been evaluated by using statistical error features; however, it is difficult to estimate the quality of a specific application using AMs. Thus, it is a great challenge to select or design appropriate AMs for an accuracy-constrained application. This paper proposes an application-oriented error evaluation framework for AMs with the aim of exploring the correlation between statistical error features of AMs and the accuracy degradation in AM-based NN applications. Specifically, based on the Dropout Feature Ranking technique, statistical error features of AMs are extensively studied and ranked by their importance to the accuracy of AM-based NN applications. The three most informative features are obtained to construct error models to predict the accuracy loss of AM-based NN applications. The constructed classification models show a probability higher than 96% for correctly classifying the AMs into three categories in accordance with the induced accuracy loss in AM-based NN applications. Furthermore, regression models can predict the accuracy of NN applications using an AM with a deviation as low as 6%. These results show that the proposed error evaluation framework can guide an efficient selection of AMs for NN applications by using just several AM error features, instead of running time-consuming and complicated hardware simulation. The obtained statistical error features can also provide a guidance for the design or generation of application-oriented AMs. Moreover, the proposed framework is applicable for quickly analyzing and selecting other approximate circuits for error-tolerant applications. Hai Mo, Yong Wu 0009, Honglan Jiang, Zining Ma, Fabrizio Lombardi, Jie Han 0001, Leibo Liu |
IEEE Trans. Computers | 6 |
| 2024 | Hardware-Efficient Logarithmic Floating-Point Multipliers for Error-Tolerant ApplicationsabstractThe increasing computational intensity of important new applications poses a challenge for their use in resource-restricted devices. Approximate computing using power-efficient arithmetic circuits is one of the emerging strategies to reach this objective. In this article, five hardware-efficient logarithmic floating-point (FP) multipliers are proposed, which all use simple operators, such as adders and multiplexers, to replace complex and more costly conventional FP multipliers. Radix-4 logarithms are used to further reduce the hardware complexity. These designs produce double-sided error distributions to mitigate error accumulation in complex computations. The proposed multipliers provide superior trade-offs between accuracy and hardware, with up to 30.8% higher accuracy than a recent logarithmic FP design or up to$68\times $less energy than the conventional FP multiplier. Using the proposed FP logarithmic multipliers in JPEG image compression achieves higher image quality than a recent logarithmic multiplier design with up to 4.7 dB larger peak signal-to-noise ratio. For training in benchmark NN applications, the proposed FP multipliers can slightly improve the classification accuracy while achieving$4.2\times $less energy and$2.2\times $smaller area than the state-of-the-art design. Zijing Niu, Honglan Jiang, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | A Survey on Approximate Multiplier Designs for Energy Efficiency: From Algorithms to CircuitsabstractGiven the stringent requirements of energy efficiency for Internet-of-Things edge devices, approximate multipliers, as a basic component of many processors and accelerators, have been constantly proposed and studied for decades, especially in error-resilient applications. The computation error and energy efficiency largely depend on how and where the approximation is introduced into a design. Thus, this article aims to provide a comprehensive review of the approximation techniques in multiplier designs ranging from algorithms and architectures to circuits. We have implemented representative approximate multiplier designs in each category to understand the impact of the design techniques on accuracy and efficiency. The designs can then be effectively deployed in high-level applications, such as machine learning, to gain energy efficiency at the cost of slight accuracy loss. Chuangtao Chen 0001, Weihua Xiao, Xuan Wang 0027, Chenyi Wen, Jie Han 0001, Xunzhao Yin, Weikang Qian, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2024 | Design of a Stochastic Computing Architecture for the Phansalkar AlgorithmabstractBinarization plays a key role in image processing. Its performance directly affects the success of subsequent character segmentation and recognition. The Phansalkar algorithm performs excellent in processing heavily degraded or poor-quality images. However, this algorithm incurs significant hardware costs. In this article, efficient stochastic computing (SC) functions and an architecture are proposed for the Phansalkar algorithm. Highly accurate stochastic elements are designed for this architecture, including a stochastic mean circuit (SMC), a stochastic unipolar subtractor (USUB), a stochastic square root circuit (SQRT), and a stochastic exponential circuit (SEXP). Simulation results show that the SC architecture using 64-bit streams for the Phansalkar algorithm provides sufficient accuracy. Physical implementation indicates the effectiveness of the proposed architecture in lowering hardware costs for this algorithm compared with the binary counterpart. Yongqiang Zhang 0006, Jiao Qin, Jie Han 0001, Guangjun Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Improving Adversarial Robustness with Self-Paced Hard-Class Pair ReweightingabstractDeep Neural Networks are vulnerable to adversarial attacks. Among many defense strategies, adversarial training with untargeted attacks is one of the most effective methods. Theoretically, adversarial perturbation in untargeted attacks can be added along arbitrary directions and the predicted labels of untargeted attacks should be unpredictable. However, we find that the naturally imbalanced inter-class semantic similarity makes those hard-class pairs become virtual targets of each other. This study investigates the impact of such closely-coupled classes on adversarial attacks and develops a self-paced reweighting strategy in adversarial training accordingly. Specifically, we propose to upweight hard-class pair losses in model optimization, which prompts learning discriminative features from hard classes. We further incorporate a term to quantify hard-class pair consistency in adversarial training, which greatly boosts model robustness. Extensive experiments show that the proposed adversarial training method achieves superior robustness performance over state-of-the-art defenses against a wide range of adversarial attacks. The code of the proposed SPAT is published at https://github.com/puerrrr/Self-Paced-Adversarial-Training. Pengyue Hou, Jie Han 0001 |
AAAI | 2 |
| 2023 | Hardware Efficient Weight-Binarized Spiking Neural NetworksabstractThe advancement in spiking neural networks (SNNs) provides a promising and alternative approach to conventional artificial neural networks (ANNs) with higher energy efficiency. However, the significant requirements on memory usage presents a performance bottleneck on resource constrained devices. Inspired by the notion of binarized neural networks (BNNs), we incorporate the design principles in BNNs into that of SNNs to reduce the stringent resource requirements. Specifically, the weights are binarized to 1 and -1 for implementing the functions of excitatory and inhibitory synapses. Hence, the proposed design is referred to as a weight-binarized spiking neural network (WB-SNN). In the WB-SNN, only one bit is used for the weight or a spike; for the latter, 1 and 0 indicate a spike and no spike, respectively. A priority encoder is used to identify the index of an active neuron as a basic unit to construct the WB-SNN. We further design a fully connected neural network that consists of an input layer, an output layer, and fully connected layers of different sizes. A counter is utilized in each neuron to complete the accumulation of weights. The WB-SNN design is validated by using a multi-layer perceptron on the MNIST dataset. Hardware implementations on FPGAs show that the WB-SNN attains a significant saving of memory with only a limited accuracy loss compared with its SNN and BNN counterparts. Chengcheng Tang, Jie Han 0001 |
DATE | 2 |
| 2023 | Approximate Processing Element Design and Analysis for the Implementation of CNN Accelerators
Honglan Jiang, Hai Mo, Jie Han 0001, Leibo Liu, Zhigang Mao |
J. Comput. Sci. Technol. | 4 |
| 2023 | An Energy-Efficient Binary-Interfaced Stochastic Multiplier Using Parallel DatapathsabstractStochastic computing (SC) typically requires a low design complexity compared with weighted binary computing, so it has been successfully applied in neural networks (NNs). Usually, SC utilizes random bitstreams as its medium, which makes it suffer from a long delay that offsets its advantages. This drawback can be alleviated by utilizing parallel datapaths, which, however, will significantly increase the hardware cost due to the requirement of multiple parallel computing units. In this article, a hybrid bit-splitting generator (HBSG) is proposed to efficiently produce parallel bitstreams in a single clock cycle to reduce delay. The HBSG uniformly splits binary numbers into R segments, each of which is encoded in parallel by using hardwired connections according to the weight of each bit. A binary-interfaced parallel stochastic multiplier (BipSMul) using the HBSG is then proposed to accelerate the multiplication in SC. Experimental results show that the BipSMul is more energy efficient than the state-of-the-art parallel and serial stochastic designs, as well as their binary and Booth counterparts, in delay, power-delay product (PDP), and area-delay product (ADP). Yongqiang Zhang 0006, Siting Liu 0001, Jie Han 0001, Zhendong Lin, Xin Cheng 0001, Guangjun Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Solving traveling salesman problems via a parallel fully connected ising machineabstractAnnealing-based Ising machines have shown promising results in solving combinatorial optimization problems. As a typical class of these problems, however, traveling salesman problems (TSPs) are very challenging to solve due to the constraints imposed on the solution. This article proposes a parallel annealing algorithm for a fully connected Ising machine that significantly improves the accuracy and performance in solving constrained combinatorial optimization problems such as the TSP. Unlike previous parallel annealing algorithms, this improved parallel annealing (IPA) algorithm efficiently solves TSPs using an exponential temperature function with a dynamic offset. Compared with digital annealing (DA) and momentum annealing (MA), the IPA reduces the run time by 44.4 times and 19.9 times for a 14-city TSP, respectively. Large scale TSPs can be more efficiently solved by taking a k-medoids clustering approach that decreases the average travel distance of a 22-city TSP by 51.8% compared with DA and by 42.0% compared with MA. This approach groups neighboring cities into clusters to form a reduced TSP, which is then solved in a hierarchical manner by using the IPA algorithm. Qichao Tao, Jie Han 0001 |
DAC | 2 |
| 2022 | Efficient Traveling Salesman Problem Solvers using the Ising Model with Simulated BifurcationabstractAn Ising model-based solver has shown efficiency in obtaining suboptimal solutions for combinatorial optimization problems. As an NP-hard problem, the traveling salesman problem (TSP) plays an important role in various routing and scheduling applications. However, the execution speed and solution quality significantly deteriorate using a solver with simulated annealing (SA) due to the quadratically increasing number of spins and strong constraints placed on the spins. The ballistic simulated bifurcation (bSB) algorithm utilizes the signs of Kerr-nonlinear parametric oscillators' positions as the spins' states. It can update the states in parallel to alleviate the time explosion problem. In this paper, we propose an efficient method for solving TSPs by using the Ising model with bSB. Firstly, the TSP is mapped to an Ising model without external magnetic fields by introducing a redundant spin. Secondly, various evolution strategies for the introduced position and different dynamic configurations of the time step are considered to improve the efficiency in solving TSPs. The effectiveness is specifically discussed and evaluated by comparing the solution quality to SA. Experiments on benchmark datasets show that the proposed bSB-based TSP solvers offer superior performance in solution quality and achieve a significant speed up in runtime than recent SA-based ones. Jie Han 0001 |
DATE | 2 |
| 2022 | Upward Packet Popup for Deadlock Freedom in Modular Chiplet-Based SystemsabstractMonolithic SoCs can be decomposed into disparate chiplets that support integration with advanced pack-aging technologies. This concept is promising in reducing the manufacturing cost of large scale SoCs due to the higher yield rate and reusability of chiplets. The chiplets should be designed in a modular manner without holistic system knowledge so that they can be reused in different SoCs. However, the design modularity is a major challenge to the networks-on-chip (NoCs) of chiplets.New deadlocks may occur across both the chiplets and the interposer due to the integration, even if the NoC of each individually designed chiplet is deadlock free. However, conventional deadlock freedom approaches are unsuitable to handle such deadlocks because they require holistic knowledge and violate the modularity. Although there are several modular approaches that specifically target at integration-induced deadlocks, their routing is overly restricted and the injection control incurs additional latency. They also lack flexibility in dynamically changing topologies due to their complex software algorithm and the hard-wired components.In this paper, a key insight on the chiplet integration-induced deadlocks is gained, inspired by which a deadlock recovery framework (named UPP) is proposed. Specifically, it is verified that an integration-induced deadlock always involves a stalled upward packet moving from the interposer to the connected chiplet via the vertical link. Thus, UPP detects a deadlock by discovering the upward packet and recovers the system from deadlock by transmitting the upward packet to its destination. Hybrid flow control mechanisms are proposed to enable the upward packet to bypass the buffers and be transmitted via the normal router datapath. To guarantee the ejection of the upward packet after transmission, a lightweight protocol is proposed to reserve ejection queue entries of the network interface. Experimental results show that while adhering to design modularity, UPP provides an average runtime speedup of 3.1%∼10.3% with an area overhead of less than 4%. Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Jianfeng Zhu 0001, Honglan Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu |
HPCA | 4 |
| 2022 | Adversarial Fine-tune with Dynamically Regulated AdversaryabstractAdversarial training is an effective method to boost model robustness to malicious, adversarial attacks. However, such improvement in model robustness often leads to a significant sacrifice of standard performance on clean images. In many real-world applications such as health diagnosis and autonomous surgical robotics, the standard performance is more valued over model robustness against such extremely malicious attacks. This leads us to the question: to what extent can we improve the robustness of the model without sacrificing standard performance? This work tackles this problem and proposes a simple yet effective transfer learning based adversarial training strategy that disentangles the negative effects of adversarial samples on model's standard performance. In addition, we introduce a training-friendly adversarial attack algorithm, which facilitates the boost of adversarial robustness without introducing significant training complexity. Extensive experiments show that the proposed approach outperforms previous adversarial training algorithms with the following objective: to improve the robustness of the model while preserving model's standard accuracy on clean data. Pengyue Hou, Jie Han 0001, Petr Musilek |
IJCNN | 3 |
| 2022 | A Review of Simulation Algorithms of Classical Ising Machines for Combinatorial optimizationabstractCombinatorial optimization problems are difficult to solve due to the space explosion in an exhaustive search. Using Ising model-based solvers can efficiently find near-optimal solutions by minimizing the energy of a nonlinear Hamiltonian system. In contrast to Ising machines based on quantum mechanics, classical Ising machines using conventional technologies, such as the complementary metal-oxide-semiconductor, offer efficient implementations with competitive performance. In this paper, we briefly review recently developed simulation algorithms of classical Ising machines. These algorithms are classified by considering various inherent mechanisms in the simulation of physical phenomena. Then, strategies to improve the simulation efficiency are discussed by generalizing their characteristics and behaviours. These simulation algorithms are key for improving the efficiency of classical Ising machines in solving combinatorial optimization problems. Qichao Tao, Bailiang Liu, Jie Han 0001 |
ISCAS | 4 |
| 2022 | A Genetic-algorithm-based Approach to the Design of DCT Hardware AcceleratorsabstractAs modern applications demand an unprecedented level of computational resources, traditional computing system design paradigms are no longer adequate to guarantee significant performance enhancement at an affordable cost. Approximate Computing (AxC) has been introduced as a potential candidate to achieve better computational performances by relaxing non-critical functional system specifications. In this article, we propose a systematic and high-abstraction-level approach allowing the automatic generation of near Pareto-optimal approximate configurations for a Discrete Cosine Transform (DCT) hardware accelerator. We obtain the approximate variants by using approximate operations, having configurable approximation degree, rather than full-precise ones. We use a genetic searching algorithm to find the appropriate tuning of the approximation degree, leading to optimal tradeoffs between accuracy and gains. Finally, to evaluate the actual HW gains, we synthesize non-dominated approximate DCT variants for two different target technologies, namely, Field Programmable Gate Arrays (FPGAs) and Application Specific Integrated Circuits (ASICs). Experimental results show that the proposed approach allows performing a meaningful exploration of the design space to find the best tradeoffs in a reasonable time. Indeed, compared to the state-of-the-art work on approximate DCT, the proposed approach allows an 18% average energy improvement while providing at the same time image quality improvement. Mario Barbareschi, Salvatore Barone, Alberto Bosio, Jie Han 0001, Marcello Traiola |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | Characterizing Approximate Adders and Multipliers for Mitigating Aging and Temperature DegradationsabstractThe performance of nanoscale semiconductor technologies has become susceptible to high temperatures and aging phenomena. While guard-bands have conventionally been used to combat degradation-induced timing violations, approximations have recently been leveraged to compensate for degradations in lieu of adding timing guard-bands, without a loss in performance. However, only simple approximation techniques such as truncation have been considered in prior work. In this paper, a wide range of approximate arithmetic circuits including adders and multipliers using various sophisticated approximation techniques are investigated to cope with aging- and temperature-induced degradations. To this end, approximate circuits are first characterized for their delay increase under degradations. With this, we then determine the approximation level required to compensate for guard-bands under different degradations. Degradation-aware logic synthesis results show that the simple use of truncated arithmetic circuits leads to a higher quality loss compared to using other approximate circuits. However, a truncated multiplier has the lowest error distance towards a reliable operation in 10 years. The approximate multipliers with configurable error recovery are most suitable when the level of degradation is higher, e.g., at a temperature of 70 °C. The characterization of degradation at the circuit level is then used for design exploration at the architecture level without the need for further gate-level simulations. For three different image processing applications, experimental results show that guard-bands can be mitigated while maintaining an output result with a high visual quality. Francisco J. H. Santiago, Honglan Jiang, Hussam Amrouch, Andreas Gerstlauer, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | An Energy-Efficient Approximate Divider Based on Logarithmic Conversion and Piecewise Constant ApproximationabstractApproximate computing (AC) has been considered as a promising paradigm to improve the energy-efficiency of computing hardware for error-tolerant applications, with negligible quality degradation to the output. Dividers frequently limit the performance of a computing system; however, they have not received as much attention as multipliers and adders in AC. In this paper, an energy-efficient and high-performance approximate divider is proposed based on logarithmic conversion and piecewise constant approximation. In this design, the range for the conversion between binary and logarithmic numbers is first expanded from$\mathbf {[{0,1}]}$to$\mathbf {[-0.5,1]}$. A heuristic search algorithm is then devised to find the most accurate constant set to approximate the reciprocal of the divisor, by minimizing a statistical error. The hardware implementation is presented for both floating-point (FP) and integer dividers. With a high configurability, the proposed divider results in a mean relative error distance (MRED) from 2.78% to 0.046%, indicating a high accuracy among state-of-the-art approximate dividers. Compared to the half-precision FP divider, the proposed divider with a MRED of 0.74% can achieve nearly$\mathbf {90\times }$improvement in PDP. Moreover, compared to state-of-the-art approximate dividers, the proposed design is in the Pareto Frontier in terms of power delay product (PDP) and MRED. The three image processing application results demonstrate that the proposed divider can result in the highest peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) even with truncation. Yong Wu 0009, Honglan Jiang, Zining Ma, Pengfei Gou, Jie Han 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | A Logarithmic Floating-Point Multiplier for the Efficient Training of Neural NetworksabstractThe development of important applications of increasingly large neural networks (NNs) is spurring research that aims to increase the power efficiency of the arithmetic circuits that perform the huge amount of computation in NNs. The floating-point (FP) representation with a large dynamic range is usually used for training. In this paper, it is shown that the FP representation is naturally suited for the binary logarithm of numbers. Thus, it favors a design based on logarithmic arithmetic. Specifically, we propose an efficient hardware implementation of logarithmic FP multiplication that uses simpler operations to replace complex multipliers for the training of NNs. This design produces a double-sided error distribution that mitigates the accumulative effect of errors in iterative operations, so it is up to 45% more accurate than a recent logarithmic FP design. The proposed multiplier also consumes up to 23.5x less energy and 10.7x smaller area compared to exact FP multipliers. Benchmark NN applications, including a 922-neuron model for the MNIST dataset, show that the classification accuracy can be slightly improved using the proposed multiplier, while achieving up to 2.4x less energy and 2.8x smaller area with a better performance. Zijing Niu, Honglan Jiang, Mohammad Saeed Ansari, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2021 | An Improved Logarithmic Multiplier for Energy-Efficient Neural ComputingabstractMultiplication is the most resource-hungry operation in neural networks (NNs). Logarithmic multipliers (LMs) simplify multiplication to shift and addition operations and thus reduce the energy consumption. Since implementing the logarithm in a compact circuit often introduces approximation, some accuracy loss is inevitable in LMs. However, this inaccuracy accords with the inherent error tolerance of NNs and their associated applications. This article proposes an improved logarithmic multiplier (ILM) that, unlike existing designs, rounds both inputs to their nearest powers of two by using a proposed nearest-one detector (NOD) circuit. Considering that the output of the NOD uses a one-hot representation, some entries in the truth table of a conventional adder cannot occur. Hence, a compact adder is designed for the reduced truth table. The 8x8 ILM achieves up to 17.48 percent saving in power consumption compared to a recent LM in the literature while being almost 8 percent more accurate. Moreover, the evaluation of the ILM for two benchmark NN workloads shows up to 21.85 percent reduction in energy consumption compared to the NNs implemented with other LMs. Interestingly, using the ILM increases the classification accuracy of the considered NNs by up to 1.4 percent compared to a NN implementation that uses exact multipliers. Mohammad Saeed Ansari, Bruce F. Cockburn, Jie Han 0001 |
IEEE Trans. Computers | 3 |
| 2021 | A Deflection-Based Deadlock Recovery Framework to Achieve High Throughput for Faulty NoCsabstractDeadlock is a critical issue in faulty Networks-on-Chips (NoCs). Existing deadlock-free approaches on faulty NoCs suffer from low throughput and poor fairness when the network becomes oversaturated. This problem hinders their practical use as oversaturation scenarios are more frequent on faulty NoCs. To address this issue, a deflection-based deadlock recovery framework is proposed for higher oversaturation performance on faulty NoCs. First, we observe the low oversaturation performance of existing deadlock recovery approaches, and analyze the positive feedback loop that can amplify the negative impact of deadlocks and congestions, which necessitate handling both deadlocks and congestions in a deadlock recovery framework. Second, we propose a novel deadlock recovery framework, which includes an accurate, timely deadlock detection and a highly efficient deadlock recovery. Both the deadlock detection and recovery reduce the average packet traversal latency, thereby improving the average oversaturation throughput. Third, we propose a distributed implementation to make the entire network enter and exit the deflection mode, which is conducted by broadcasting special messages via a bufferless subnetwork. An average oversaturation throughput improvement of 1.1 ~ 8.1× over state-of-the-art approaches is achieved. In terms of fairness, the minimal oversaturation throughput is improved from near zero to half of the peak throughput. Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Non-Volatile Approximate Arithmetic Circuits Using Scalable Hybrid Spin-CMOS Majority GatesabstractIn the nanoscale era, leakage/static power dissipation has become an inevitable and important issue for CMOS devices. To alleviate this issue, we propose to use spintronic devices with near-zero leakage power and non-volatility as key components in arithmetic circuits for error-resilient applications. To this end, spintronic threshold devices are first utilized to construct highly-scalable majority gates (MGs) based on spin-CMOS technology. These MGs are then used in the design of compressors for constructing multipliers and accumulators. For an MG-based compressor, the truth table of a conventional compressor is transformed to ensure that the outputs depend only on the number of input “1”s. To synthesize and optimize the MG-based circuits, a heuristic majority-inverter graph (HMIG) is further proposed for the design of an accurate and two approximate non-volatile 4-2 compressors (denoted as MG-EC, MG-AC1 and MG-AC2). Due to the high scalability of the MGs, approximate compressors with a larger number of inputs can be devised using the same method. Compared to previous designs, the proposed 4-2 compressors show shorter critical path delays and lower energy consumption; MG-AC1 and MG-AC2 also achieve a higher accuracy than state-of-the-art approximate designs. For achieving a similar image quality in image compression, the multiplier implementations using MG-AC1 and MG-AC2 result in more significant reductions in delay and energy than those using other approximate designs. Honglan Jiang, Shaahin Angizi, Deliang Fan, Jie Han 0001, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | High Performance CNN Accelerators Based on Hardware and Algorithm Co-OptimizationabstractConvolutional neural networks (CNNs) have been widely used in image classification and recognition due to their effectiveness; however, CNNs use a large volume of weight data that is difficult to store in on-chip memory of embedded designs. Pruning can compress the CNN model at a small accuracy loss; however, a pruned CNN model operates slower when implemented on a parallel architecture. In this paper, a hardware-oriented CNN compression strategy is proposed; a deep neural network (DNN) model is divided into “no-pruning layers ($NP$-layers)” and “pruning layers ($P$-layers)”. A$NP$-layer has a regular weights distribution for parallel computing and high performance. A$P$-layer is irregular due to pruning, but it generates a high compression ratio. Uniform and incremental quantization schemes are used to achieve a tradeoff between compression ratio and processing efficiency at a small loss in accuracy. A distributed convolutional architecture with several parallel finite impulse response (FIR) filters is further proposed for the regular model in the$NP$-layers. A shift-accumulator based processing element with an activation-driven data flow (ADF) is proposed for the irregular sparse model in the$P$-layers. Based on the proposed compression strategy and hardware architecture, a hardware/algorithm co-optimization (HACO) approach is proposed for implementing a$NP-P$hybrid compressed CNN model on FPGAs. For a hardware accelerator on a single FPGA chip without the use of off-chip memory, a$27.5\times $compression ratio is achieved with 0.44% top-5 accuracy loss for VGG-16. The implementation of the compressed VGG-16 model on a Xilinx VCU118 evaluation board processes 83.0 frames per second (FPS) for image applications, this is$1.8\times $superior than the state-of-the-art design found in the technical literature. Weiqiang Liu 0001, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A Survey of Stochastic Computing Neural Networks for Machine Learning ApplicationsabstractNeural networks (NNs) are effective machine learning models that require significant hardware and energy consumption in their computing process. To implement NNs, stochastic computing (SC) has been proposed to achieve a tradeoff between hardware efficiency and computing performance. In an SC NN, hardware requirements and power consumption are significantly reduced by moderately sacrificing the inference accuracy and computation speed. With recent developments in SC techniques, however, the performance of SC NNs has substantially been improved, making it comparable with conventional binary designs yet by utilizing less hardware. In this article, we begin with the design of a basic SC neuron and then survey different types of SC NNs, including multilayer perceptrons, deep belief networks, convolutional NNs, and recurrent NNs. Recent progress in SC designs that further improve the hardware efficiency and performance of NNs is subsequently discussed. The generality and versatility of SC NNs are illustrated for both the training and inference processes. Finally, the advantages and challenges of SC NNs are discussed with respect to binary counterparts. Yidong Liu, Siting Liu 0001, Yanzhi Wang 0001, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus TopologyabstractFuture NoCs should be highly flexible to adapt to communication demands to achieve high scalability and low power consumption. However, the flexibility is still quite limited by the high complexity of reconfiguration for globally reconfigured channels. In this paper, we propose to augment a router-based buffered NoC with a reconfigurable ring architecture by exploiting cycle decomposition of a torus bufferless network. At runtime, the topologies of the rings can be reconfigured according to the workloads by choosing different cycle decompositions of the torus network. Because the shapes of the rings are restricted to a specified regular shape, the reconfiguration time can be reduced to a linear complexity with respect to network size, and the reconfiguration algorithm can be implemented in a distributed hardware. The experimental results show that the reconfigurable rings provide 54% and 26% improvements on packet latency and static power saving, respectively, for realistic workloads. Liang Wang 0020, Leibo Liu, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei |
DAC | 4 |
| 2020 | Dynamic Stochastic Computing for Digital Signal Processing ApplicationsabstractStochastic computing (SC) utilizes a random binary bit stream to encode a number by counting the frequency of 1's in the stream (or sequence). Typically, a small circuit is used to perform a bit-wise logic operation on the stochastic sequences, which leads to significant hardware and power savings. Energy efficiency, however, is a challenge for SC due to the long sequences required for accurately encoding numbers. To overcome this challenge, we consider to use a stochastic sequence to encode a continuously variable signal instead of a number to achieve higher accuracy, higher energy efficiency and greater flexibility. Specifically, one single bit is used to encode a sample from a signal for efficient processing. This type of sequences encodes constantly variable values, so it is referred to as dynamic stochastic sequences (DSS's). The DSS enables the use of SC circuits to efficiently perform tasks such as frequency mixing and function estimation. It is shown that such a dynamic SC (DSC) system achieves savings up to 98.4% in energy and up to 96.8% in time with a slightly higher accuracy compared to conventional SC. It also achieves energy and time savings of up to 60% compared to a fixed-width binary implementation. Siting Liu 0001, Jie Han 0001 |
DATE | 2 |
| 2020 | Approximate Arithmetic Circuits: A Survey, Characterization, and Recent ApplicationsabstractApproximate computing has emerged as a new paradigm for high-performance and energy-efficient design of circuits and systems. For the many approximate arithmetic circuits proposed, it has become critical to understand a design or approximation technique for a specific application to improve performance and energy efficiency with a minimal loss in accuracy. This article aims to provide a comprehensive survey and a comparative evaluation of recently developed approximate arithmetic circuits under different design constraints. Specifically, approximate adders, multipliers, and dividers are synthesized and characterized under optimizations for performance and area. The error and circuit characteristics are then generalized for different classes of designs. The applications of these circuits in image processing and deep neural networks indicate that the circuits with lower error rates or error biases perform better in simple computations, such as the sum of products, whereas more complex accumulative computations that involve multiple matrix multiplications and convolutions are vulnerable to single-sided errors that lead to a large error bias in the computed result. Such complex computations are more sensitive to errors in addition than those in multiplication, so a larger approximation can be tolerated in multipliers than in adders. The use of approximate arithmetic circuits can improve the quality of image processing and deep learning in addition to the benefits in performance and power consumption for these applications. Honglan Jiang, Francisco J. H. Santiago, Hai Mo, Leibo Liu, Jie Han 0001 |
Proc. IEEE | 5 |
| 2020 | A Novel Heuristic Search Method for Two-Level Approximate Logic SynthesisabstractRecently, much attention has been paid to approximate computing, a novel design paradigm for error-tolerant applications. It can significantly reduce area, power, and delay of circuits by introducing an acceptable amount of error. In this paper, we propose a new heuristic method for two-level approximate logic synthesis. The problem is to identify an approximate sum-of-product (SOP) expression under a given error rate (ER) constraint so that it has the fewest literals. The basic idea of our method is to find an optimal set of input combinations for 0-to-1 output complement (SICC). For this purpose, we first identify all prime SICCs, which are fundamental SICCs in the sense that the optimal SICC is very likely to be a union of a subset of the prime SICCs. Then, we search among all subsets of the prime SICCs the optimal subset, which leads to a final good approximate SOP. We further propose four speed-up techniques. The experiments on benchmarks showed that our method is better than the previous state-of-the-art method and our speed-up techniques are effective. For an ER threshold of 0.8%, our method can reduce 15.8% literals on average. Sanbao Su, Chen Zou 0001, Weijiang Kong, Jie Han 0001, Weikang Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Aggressive Fine-Grained Power Gating of NoC BuffersabstractPower gating is effective for networks-on-chip (NoCs) to reduce the excessive leakage power dissipated by idle network components. Most existing NoC power-gating approaches rely on the routing algorithms to mitigate the power-gating blocking latency problem. When the network becomes faulty and fault-tolerant routing algorithms are applied, these approaches are no longer applicable or can seriously degrade the performance. Other approaches propose fine-grained buffer power gating, but they are too conservative in power saving due to the buffer backpressure flow control. To address these problems, we propose an aggressive fine-grained power gating of flit-sized buffer entries by adopting backpressureless flow control in an input-buffered network. The power-gating decisions are made based on the flit deflection rate. However, directly applying the backpressureless flow control leads to the difficulties of multiflit packet truncation and protocol deadlocks. Therefore, we modify the packet injection architecture to avoid packet truncation. This is done by chaining the local input port with a randomly chosen input port. Finally, we design a progressive recovery framework to handle both livelocks and protocol deadlocks. It does not need to truncate packets or strictly separate different message classes when the network is free of livelocks or protocol deadlocks. The experimental results show that with a hardware overhead of 9.6%, our design can save up to 59% network power consumption in both a fault-free and a faulty NoC with little zero-load latency penalty. Our design also approaches an ideal energy-proportional NoC because it can constantly reduce power consumption over a wide range of injection rates. Leibo Liu, Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Achieving Flexible Global Reconfiguration in NoCs Using Reconfigurable RingsabstractThe communication behaviors in NoCs of chip-multiprocessors exhibit great spatial and temporal variations, which introduce significant challenges for the reconfiguration in NoCs. Existing reconfigurable NoCs are still far from ideal reconfiguration scenarios, in which globally reconfigurable interconnects can be immediately reconfigured to provide bandwidths on demand for varying traffic flows. In this paper, we propose a hybrid NoC architecture that globally reconfigures the ring-based interconnect to adapt to the varying traffic flows with a high flexibility. The ring-based interconnect has the following advantages. First, it includes horizontal rings and vertical rings, which can be dynamically combined or split to provide low-latency channels for heavy traffic flows. Second, each combined ring connects a number of nodes, thereby improving both the utilization of each ring and the probability to reuse previous reconfigurable interconnects. Finally, the reconfiguration algorithm has a linear-time complexity and can be implemented using a low-overhead hardware design, making it possible to achieve a fast reconfiguration in NoCs. The experimental results show that compared to recent reconfigurable NoCs, the proposed NoC architecture can greatly improve the saturation throughput for synthetic traffic patterns, and reduce the packet latency over 40 percent for realistic benchmarks without incurring significant area and power overhead. Liang Wang 0020, Leibo Liu, Jie Han 0001, Xiaohang Wang 0001, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Improving the Accuracy and Hardware Efficiency of Neural Networks Using Approximate MultipliersabstractImproving the accuracy of a neural network (NN) usually requires using larger hardware that consumes more energy. However, the error tolerance of NNs and their applications allow approximate computing techniques to be applied to reduce implementation costs. Given that multiplication is the most resource-intensive and power-hungry operation in NNs, more economical approximate multipliers (AMs) can significantly reduce hardware costs. In this article, we show that using AMs can also improve the NN accuracy by introducing noise. We consider two categories of AMs: 1) deliberately designed and 2) Cartesian genetic programing (CGP)-based AMs. The exact multipliers in two representative NNs, a multilayer perceptron (MLP) and a convolutional NN (CNN), are replaced with approximate designs to evaluate their effect on the classification accuracy of the Mixed National Institute of Standards and Technology (MNIST) and Street View House Numbers (SVHN) data sets, respectively. Interestingly, up to 0.63% improvement in the classification accuracy is achieved with reductions of 71.45% and 61.55% in the energy consumption and area, respectively. Finally, the features in an AM are identified that tend to make one design outperform others with respect to NN accuracy. Those features are then used to train a predictor that indicates how well an AM is likely to work in an NN. Mohammad Saeed Ansari, Vojtech Mrazek, Bruce F. Cockburn, Lukás Sekanina, Zdenek Vasícek, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | A Hardware-Efficient Logarithmic Multiplier with Improved AccuracyabstractLogarithmic multipliers take the base-2 logarithm of the operands and perform multiplication by only using shift and addition operations. Since computing the logarithm is often an approximate process, some accuracy loss is inevitable in such designs. However, the area, latency, and power consumption can be significantly improved at the cost of accuracy loss. This paper presents a novel method to approximate log2N that, unlike the existing approaches, rounds N to its nearest power of two instead of the highest power of two smaller than or equal to N. This approximation technique is then used to design two improved 16×16 logarithmic multipliers that use exact and approximate adders (ILM-EA and ILM-AA, respectively). These multipliers achieve up to 24.42% and 9.82% savings in area and power-delay product, respectively, compared to the state-of-the-art design in the literature with similar accuracy. The proposed designs are evaluated in the Joint Photographic Experts Group (JPEG) image compression algorithm and their advantages over other approximate logarithmic multipliers are shown. Mohammad Saeed Ansari, Bruce F. Cockburn, Jie Han 0001 |
DATE | 3 |
| 2019 | Characterizing Approximate Adders and Multipliers Optimized under Different Design ConstraintsabstractTaking advantage of the error resilience in many applications as well as the perceptual limitations of humans, numerous approximate arithmetic circuits have been proposed that trade off accuracy for higher speed or lower power in emerging applications that exploit approximate computing. However, characterizing the various approximate designs for a specific application under certain performance constraints becomes a new challenge. In this paper, approximate adders and multipliers are evaluated and compared for a better understanding of their characteristics when the implementations are optimized for performance or power. Although simple truncation can effectively reduce the hardware of an arithmetic circuit, it is shown that some other designs perform better in speed, power and power-delay product. For instance, many approximate adders have a higher performance than a truncated adder. A truncated multiplier is faster but consumes a higher power than most approximate designs for achieving a similar mean error magnitude. The logarithmic multipliers are very fast and power-efficient at a lower accuracy. Approximate multipliers can also be generated by an automated process to be very efficient while ensuring a sufficiently high accuracy. Honglan Jiang, Francisco J. H. Santiago, Mohammad Saeed Ansari, Leibo Liu, Bruce F. Cockburn, Fabrizio Lombardi, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | A stochastic-computing based deep learning framework using adiabatic quantum-flux-parametron superconducting technologyabstractThe Adiabatic Quantum-Flux-Parametron (AQFP) superconducting technology has been recently developed, which achieves the highest energy efficiency among superconducting logic families, potentially 104--105 gain compared with state-of-the-art CMOS. In 2016, the successful fabrication and testing of AQFP-based circuits with the scale of 83,000 JJs have demonstrated the scalability and potential of implementing large-scale systems using AQFP. As a result, it will be promising for AQFP in high-performance computing and deep space applications, with Deep Neural Network (DNN) inference acceleration as an important example. Ruizhe Cai, Ao Ren, Olivia Chen, Ning Liu 0007, Caiwen Ding, Xuehai Qian, Jie Han 0001, Wenhui Luo, Nobuyuki Yoshikawa, Yanzhi Wang 0001 |
ISCA | 7 |
| 2019 | Low-Power Unsigned Divider and Square Root Circuit Designs Using Adaptive ApproximationabstractIn this paper, an adaptive approximation approach is proposed for the design of a divider and a square root (SQR) circuit. In this design, the division/SQR is computed by using a reduced-width divider/SQR circuit and a shifter by adaptively pruning some insignificant input bits. Specifically, for a $2n/n$ 2 n / n division, $2k$ 2 k and $k$ k ($k< n$ k < n ) consecutive bits are selected starting from the most significant ‘1’ in the dividend and divisor, respectively. At the same time, redundant least significant bits (LSBs) are truncated or if the number of remaining bits after pruning is smaller than the number of bits to be kept, ‘0's are appended to the LSBs of the inputs. To avoid overflow, a $2(k+1)/(k+1)$ 2 ( k + 1 ) / ( k + 1 ) divider is used to compute the $2k/k$ 2 k / k division. Finally, an error correction circuit is proposed to recover the error caused by the shifter using OR gates. For a $2n$ 2 n -bit approximate SQR circuit, similar pruning schemes are used to obtain a $2k$ 2 k -bit radicand. A $2k$ 2 k -bit SQR circuit and a shifter are then utilized to compute the SQR. This adaptive operation leads to very small maximum error distances of the approximate divider and SQR circuits, as shown by a theoretical error analysis. The proposed 16/8 approximate divider using an 8/4 exact array divider is $2.5\times$ 2 . 5 × as fast but only consumes 34.42 percent of the power of the accurate design. Compared to the accurate 16-bit array SQR circuit, the approximate design with a 6-bit radicand is $3.9\times$ 3 . 9 × as fast and consumes 20.66 percent of the power. The approximate SQR circuit using a 6-bit lookup table-based SQR circuit consumes 7.15 percent of the power of its corresponding accurate design. The proposed designs outperform other approximate designs in image processing applications including change detection (for the divider), envelope detection (for the SQR circuit) and image reconstruction (for both designs). Honglan Jiang, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Computers | 4 |
| 2019 | A Lifetime Reliability-Constrained Runtime Mapping for Throughput Optimization in Many-Core SystemsabstractDue to technology scaling, lifetime reliability is becoming one of the major design constraints in the performance optimization of future many-core systems. Given a lifetime reliability constraint, the existing lifetime-constrained runtime mapping schemes often lead to low throughput because of the requirement to map all applications to compact regions. In this paper, we propose a runtime application mapping scheme that exploits a borrowing strategy to improve the throughput of many-core systems given a lifetime constraint. First, we propose using different strategies for mapping communication-intensive applications and computation-intensive applications. The lifetime reliability constraint can be relaxed in the local time scale when the communication requirement is high. The throughput is improved because the communication distance of communication-intensive applications is optimized while the waiting time of computation-intensive application is reduced. Then, we propose a method to effectively classify applications depending on the communication-to-computation ratio. A dynamic threshold is determined according to the current locations of available cores. Finally, we propose an improved neighborhood allocation scheme to reduce the communication cost in the task mapping. The experimental results show that compared to the state-of-the-art lifetime-constrained mapping, the proposed mapping scheme improves the throughput of many-core systems by 26% on average for synthetic task graphs and by 20% on average for realistic task graphs while the lifetime reliability is maintained within a constraint. Liang Wang 0020, Ping Lv, Leibo Liu, Jie Han 0001, Ho-fung Leung, Xiaohang Wang 0001, Shouyi Yin, Shaojun Wei, Terrence S. T. Mak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | An Energy-Efficient and Noise-Tolerant Recurrent Neural Network Using Stochastic ComputingabstractRecurrent neural networks (RNNs) are widely used to solve a large class of recognition problems, including prediction, machine translation, and speech recognition. The hardware implementation of RNNs is, however, challenging due to the high area and energy consumption of these networks. Recently, stochastic computing (SC) has been considered for implementing neural networks and reducing the hardware consumption. In this paper, we propose an energy-efficient and noise-tolerant long short-term memory-based RNN using SC. In this SC-RNN, a hybrid structure is developed by utilizing SC designs and binary circuits to improve the hardware efficiency without significant loss of accuracy. The area and energy consumption of the proposed design are between 1.6%-2.3% and 6.5%-11.2%, respectively, of a 32-bit floating-point (FP) implementation. The SC-RNN requires significantly smaller area and lower energy consumption in most cases compared to an 8-bit fixed point implementation. The proposed design achieves a higher noise tolerance compared to binary implementations. The inference accuracy is from 10% to 13% higher than an FP design when the noise level is high in the computation process. Yidong Liu, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Adaptive approximation in arithmetic circuits: A low-power unsigned divider designabstractMany approximate arithmetic circuits have been proposed for high-performance and low-power applications. However, most designs are either hardware-efficient with a low accuracy or very accurate with a limited hardware saving, mostly due to the use of a static approximation. In this paper, an adaptive approximation approach is proposed for the design of a divider. In this design, division is computed by using a reduced-width divider and a shifter by adaptively pruning the input bits. Specifically, for a 2n/n division 2k/k bits are selected starting from the most significant `1' in the dividend/divisor. At the same time, redundant least significant bits (LSBs) are truncated or if the number of remaining LSBs is smaller than 2k for the dividend or k for the divisor, `0's are appended to the LSBs of the input. To avoid overflow, a 2(k + 1)/(k + 1) divider is used to compute the division of the 2k-bit dividend and the k-bit divisor, both with the most significant bits being `0'. Thus, k <; n is a key variable that determines the size of the divider and the accuracy of the approximate design. Finally, an error correction circuit is proposed to recover the error caused by the shifter by using OR gates. The synthesis results in an industrial 28nm CMOS process show that the proposed 16/8 approximate divider using an 8/4 accurate divider is 2.5χ as fast and consumes 34.42% of the power of the accurate 16/8 design. Compared with the other approximate dividers, the proposed design is significantly more accurate at a similar power-delay product. Moreover, simulation results show that the proposed approximate divider outperforms the other designs in two image processing applications. Honglan Jiang, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
DATE | 4 |
| 2018 | An energy-efficient stochastic computational deep belief networkabstractDeep neural networks (DNNs) are effective machine learning models to solve a large class of recognition problems, including the classification of nonlinearly separable patterns. The applications of DNNs are, however, limited by the large size and high energy consumption of the networks. Recently, stochastic computation (SC) has been considered to implement DNNs to reduce the hardware cost. However, it requires a large number of random number generators (RNGs) that lower the energy efficiency of the network. To overcome these limitations, we propose the design of an energy-efficient deep belief network (DBN) based on stochastic computation. An approximate SC activation unit (A-SCAU) is designed to implement different types of activation functions in the neurons. The A-SCAU is immune to signal correlations, so the RNGs can be shared among all neurons in the same layer with no accuracy loss. The area and energy of the proposed design are 5.27% and 3.31% (or 26.55% and 29.89%) of a 32-bit floating-point (or an 8-bit fixed-point) implementation. It is shown that the proposed SC-DBN design achieves a higher classification accuracy compared to the fixed-point implementation. The accuracy is only lower by 0.12% than the floating-point design at a similar computation speed, but with a significantly lower energy consumption. Yidong Liu, Yanzhi Wang 0001, Fabrizio Lombardi, Jie Han 0001 |
DATE | 4 |
| 2018 | Leveraging Spintronic Devices for Efficient Approximate Logic and Stochastic Neural NetworksabstractITRS has identified nano-magnet based spintronic devices as promising post-CMOS technologies for information processing and data storage due to their ultra-low switching energy, non-volatility, superior endurance, excellent retention time, high integration density and compatibility with CMOS technology. As for data storage, spintronic memory has been widely accepted as a universal high performance next-generation non-volatile memory candidate. As for information processing, spintronic computing remains complementary in its features to CMOS technology. In this paper, we present two innovative spintronic computing primitives, i.e. spintronic approximate logic and spintronic stochastic neural network, which both leverage the intrinsic spintronic device physics to achieve much more compact and efficient designs than CMOS counterparts. In spintronic approximate logic, we employ the intrinsic current-mode thresholding operation to implement an accuracy-configurable adder and further demonstrate its application in approximate DSP applications. In spintronic stochastic neural networks, we leverage the stochastic properties of domain wall devices and magnetic tunnel junction to implement a low-power and robust artificial neural network design. Shaahin Angizi, Zhezhi He, Yu Bai 0004, Jie Han 0001, Mingjie Lin, Ronald F. DeMara, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Design and Application of an Approximate 2-D Convolver with Error CompensationabstractThis paper proposes an error compensation scheme of two-dimensional (2D) convolver in which both approximate circuit- and algorithm-level techniques are utilized in the design. Truncation and voltage scaling are used as circuit techniques, while bit-width reduction is utilized at the algorithm level. These different techniques are related to the configuration of the convolver by which its operation can be configured to meet different and often contrasting figures of merit. An extensive evaluation of different error metrics is performed. An error analysis is also presented to substantiate the simulation results; an error compensation scheme is introduced to remedy a loss of accuracy in computation. Convolution for image processing is treated in detail to show the effectiveness of the proposed approach. The design, the analysis and the simulation results show that the approximate techniques utilized in the inexact convolver can operate in synergy. Ke Chen 0018, Jie Han 0001, Paolo Montuschi, Weiqiang Liu 0001, Fabrizio Lombardi |
ISCAS | 2 |
| 2018 | Approximate Arithmetic Circuits and Their ApplicationsabstractThe demand of higher speed and power efficiency, as well as the feature of error resilience in many applications (e.g., multimedia, recognition and data analytics), have driven the development of approximate computing circuits. Often as the most important arithmetic modules in a processor, adders, multipliers and dividers determine the performance and energy efficiency of many computing tasks. In this talk, a review and classification are presented for the current designs of approximate arithmetic circuits including adders, multipliers and dividers. A comparative evaluation of their error and circuit characteristics is performed for understanding the features of various designs. Image processing and cerebellar models are considered to show the effectiveness of the approximate arithmetic circuits in improving the energy efficiency and performance of computation-intensive applications. Jie Han 0001 |
NOCS | 1 |
| 2018 | A Stochastic Computational Multi-Layer Perceptron with Backward PropagationabstractStochastic computation has recently been proposed for implementing artificial neural networks with reduced hardware and power consumption, but at a decreased accuracy and processing speed. Most existing implementations are based on pre-training such that the weights are predetermined for neurons at different layers, thus these implementations lack the ability to update the values of the network parameters. In this paper, a stochastic computational multi-layer perceptron (SC-MLP) is proposed by implementing the backward propagation algorithm for updating the layer weights. Using extended stochastic logic (ESL), a reconfigurable stochastic computational activation unit (SCAU) is designed to implement different types of activation functions such as the tanh and the rectifier function. A triple modular redundancy (TMR) technique is employed for reducing the random fluctuations in stochastic computation. A probability estimator (PE) and a divider based on the TMR and a binary search algorithm are further proposed with progressive precision for reducing the required stochastic sequence length. Therefore, the latency and energy consumption of the SC-MLP are significantly reduced. The simulation results show that the proposed design is capable of implementing both the training and inference processes. For the classification of nonlinearly separable patterns, at a slight loss of accuracy by 1.32-1.34 percent, the proposed design requires only 28.5-30.1 percent of the area and 18.9-23.9 percent of the energy consumption incurred by a design using floating point arithmetic. Compared to a fixed-point implementation, the SC-MLP consumes a smaller area (40.7-45.5 percent) and a lower energy consumption (38.0-51.0 percent) with a similar processing speed and a slight drop of accuracy by 0.15-0.33 percent. The area and the energy consumption of the proposed design is from 80.7-87.1 percent and from 71.9-93.1 percent, respectively, of a binarized neural network (BNN), with a similar accuracy. Yidong Liu, Siting Liu 0001, Yanzhi Wang 0001, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Computers | 5 |
| 2018 | Gradient Descent Using Stochastic Circuits for Efficient Training of Learning MachinesabstractGradient descent (GD) is a widely used optimization algorithm in machine learning. In this paper, a novel stochastic computing GD circuit (SC-GDC) is proposed by encoding the gradient information in stochastic sequences. Inspired by the structure of a neuron, a stochastic integrator is used to optimize the weights in a learning machine by its “inhibitory” and “excitatory” inputs. Specifically, two AND (or XNOR) gates for the unipolar representation (or the bipolar representation) and one stochastic integrator are, respectively, used to implement the multiplications and accumulations in a GD algorithm. Thus, the SC-GDC is very area- and power-efficient. As per the formulation of the proposed SC-GDC, it provides unbiased estimate of the optimized weights in a learning algorithm. The proposed SC-GDC is then used to implement a least-mean-square algorithm and a softmax regression. With a similar accuracy, the proposed design achieves more than $30 \times $ improvement in throughput per area (TPA) and consumes less than 13% of the energy per training sample, compared with a fixed-point implementation. Moreover, a signed SC-GDC is proposed for training complex neural networks (NNs). It is shown that for a 784-128-128-10 fully connected NN, the signed SC-GDC produces a similar training result with its fixed-point counterpart, while achieving more than 90% energy saving and 82% reduction in training time with more than $50 \times $ improvement in TPA. Siting Liu 0001, Honglan Jiang, Leibo Liu, Jie Han 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Automatic Selection of Process Corner Simulations for Faster Design VerificationabstractIntegrated circuit designs are verified in simulation over a set of process corners, which are combinations of expected transistor properties, power supply voltages, and die temperatures. The simulation time per corner can be long and semiconductor processes can have more than 1000 corners. Simulation is thus a serious bottleneck in design verification. We propose an algorithm that selects the smallest number of process corner simulations that are required to estimate minimum and/or maximum values of the output functions that model circuit behavior. Using our best corner selection algorithm, the required number of process corner simulations is reduced by an average of 79% (a speed-up of 4.71) with respect to a set of 46 output functions from nine industrial benchmark circuits. Michael Shoniker, Oleg Oleynikov, Bruce F. Cockburn, Jie Han 0001, Manish Rana, Witold Pedrycz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Feedback-Based Low-Power Soft-Error-Tolerant Design for Dual-Modular Redundancy
Yufeng Li 0003, Jie Han 0001, Jianhao Hu, Fan Yang 0001, Xuan Zeng 0001, Bruce F. Cockburn, Jie Chen 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Toward Energy-Efficient Stochastic Circuits Using Parallel Sobol Sequences
Siting Liu 0001, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Scalable Construction of Approximate Multipliers With Formally Guaranteed Worst Case Error
Vojtech Mrazek, Zdenek Vasícek, Lukás Sekanina, Honglan Jiang, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | Hardware ODE Solvers using Stochastic CircuitsabstractA novel ordinary differential equation (ODE) solver is proposed by using a stochastic integrator to implement the accumulative function of the Euler method. We show that a stochastic integrator is an unbiased estimator for a Euler numerical solution. Unlike in conventional stochastic circuits, in which long stochastic bit streams are required to produce a result with a high accuracy, the proposed stochastic ODE solver provides an estimate of the solution for every bit in the stochastic bit stream, thus significantly reducing the latency and energy consumption of the circuit. Complex ODE solvers are constructed for solving nonhomogeneous ODEs, systems of ODEs and higher-order ODEs. Experimental results show that the stochastic ODE solvers provide very accurate solutions compared to their binary counterparts, with on average an energy saving of 46% (up to 74%), 8x throughput per area (up to nearly 12x) and a runtime reduction of 72% (up to 82%). Siting Liu 0001, Jie Han 0001 |
DAC | 2 |
| 2017 | Energy efficient stochastic computing with Sobol sequencesabstractEnergy efficiency presents a significant challenge for stochastic computing (SC) due to the long random binary bit streams required for accurate computation. In this paper, a type of low discrepancy (LD) sequences, the Sobol sequence, is considered for energy-efficient implementations of SC circuits. The use of Sobol sequences improves the output accuracy of a stochastic circuit with a reduced sequence length compared to the use of another type of LD sequences, the Halton sequence, and conventional linear feedback shift register (LFSR)-generated pseudorandom sequence. The use of Sobol sequences leads to a similar or higher accuracy than using Halton sequences for basic arithmetic operations. Sobol sequence generators cost less energy than the Halton counterparts when multiple random sequences are required in a circuit, thus the use of Sobol sequences can lead to a higher energy efficiency in an SC circuit than using Halton sequences. Siting Liu 0001, Jie Han 0001 |
DATE | 2 |
| 2017 | A true random number generator based on parallel STT-MTJsabstractRandom number generators are an essential part of cryptographic systems. For the highest level of security, true random number generators (TRNG) are needed instead of pseudorandom number generators. In this paper, the stochastic behavior of the spin transfer torque magnetic tunnel junction (STT-MTJ) is utilized to produce a TRNG design. A parallel structure with multiple MTJs is proposed that minimizes device variation effects. The design is validated in a 28-nm CMOS process with Monte Carlo simulation using a compact model of the MTJ. The National Institute of Standards and Technology (NIST) statistical test suite is used to verify the randomness quality when generating encryption keys for the Transport Layer Security or Secure Sockets Layer (TLS/SSL) cryptographic protocol. This design has a generation speed of 177.8 Mbit/s, and an energy of 0.64 pJ is consumed to set up the state in one MTJ. Yuanzhuo Qu, Jie Han 0001, Bruce F. Cockburn, Witold Pedrycz, Yue Zhang 0010, Weisheng Zhao 0001 |
DATE | 2 |
| 2017 | Design of Approximate High-Radix Dividers by Inexact Binary Signed-Digit AdditionabstractApproximate high radix dividers (HR-AXDs) are proposed and investigated in this paper. High-radix division is reviewed and inexact computing is introduced at different levels. Design parameters such as number of bits (N) and radix (r) are considered in the analysis; the replacement schemes with inexact cells and truncation schemes of exact cells in the binary signed-digit adder array is introduced. Circuit-level performance and the error characteristics of the inexact high radix dividers are analyzed for the proposed designs. The combined assessment of the normal error distance, power dissipation and delay is investigated and applications of approximate high-radix dividers are treated in detail. The simulation results show that the proposed approximate dividers offer extensive saving in terms of power dissipation, circuit complexity and delay, while only incurring in a small degradation in accuracy thus making them possibly suitable and interesting to some applications and domains such as low power/mobile computing. Linbin Chen, Fabrizio Lombardi, Paolo Montuschi, Jie Han 0001, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | Evaluating Data Resilience in CNNs from an Approximate Memory PerspectiveabstractDue to the large volumes of data that need to be processed, efficient memory access and data transmission are crucial for high-performance implementations of convolutional neural networks (CNNs). Approximate memory is a promising technique to achieve efficient memory access and data transmission in CNN hardware implementations. To assess the feasibility of applying approximate memory techniques, we propose a framework for the data resilience evaluation (DRE) of CNNs and verify its effectiveness on a suite of prevalent CNNs. Simulation results show that a high degree of data resilience exists in these networks. By scaling the bit-width of the first five dominant data subsets, the data volume can be reduced by 80.38% on average with a 2.69% loss in relative prediction accuracy. For approximate memory with random errors, all the synaptic weights can be stored in the approximate part when the error rate is less than 10--4, while 3 MSBs must be protected if the error rate is fixed at 10--3. These results indicate a great potential for exploiting approximate memory techniques in CNN hardware design. Yuanchang Chen, Yizhe Zhu, Fei Qiao, Jie Han 0001, Yuansheng Liu, Huazhong Yang |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | A Review, Classification, and Comparative Evaluation of Approximate Arithmetic CircuitsabstractOften as the most important arithmetic modules in a processor, adders, multipliers, and dividers determine the performance and energy efficiency of many computing tasks. The demand of higher speed and power efficiency, as well as the feature of error resilience in many applications (e.g., multimedia, recognition, and data analytics), have driven the development of approximate arithmetic design. In this article, a review and classification are presented for the current designs of approximate arithmetic circuits including adders, multipliers, and dividers. A comprehensive and comparative evaluation of their error and circuit characteristics is performed for understanding the features of various designs. By using approximate multipliers and adders, the circuit for an image processing application consumes as little as 47% of the power and 36% of the power-delay product of an accurate design while achieving similar image processing quality. Improvements in delay, power, and area are obtained for the detection of differences in images by using approximate dividers. Honglan Jiang, Cong Liu 0015, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2017 | Two Approximate Voting Schemes for Reliable ComputingabstractThis paper relies on the principles of inexact computing to alleviate the issues arising in static masking by voting for reliable computing in the nanoscales. Two schemes that utilize in different manners approximate voting, are proposed. The first scheme is referred to as inexact double modular redundancy (IDMR). IDMR does not resort to triplication, thus saving overhead due to modular replication. This scheme is crudely adaptive in its operation, i.e., it allows a threshold to determine the validity of the module outputs. IDMR operates by initially establishing the difference between the values of the outputs of the two modules; only if the difference is below a preset threshold, then the voter calculates the average value of the two module outputs. The second scheme (ITDMR) combines IDMR with TMR (triple modular redundancy) by using novel conditions in the comparison of the outputs of the three modules. Within an inexact framework, the majority is established using different criteria; in ITDMR, adaptive operation is carried further than IDMR to include approximate voting in a pairwise fashion. So, the validity of the three inputs is established and when only two of the three inputs satisfy the threshold condition, the IDMR operation is utilized. An extensive analysis that includes the voting circuits as well as a probabilistic framework is included. The proposed IDMR and ITDMR schemes improve the power dissipation and tolerance to variations compared to a traditional TMR. To further validate the applicability of the proposed schemes, inexact voting has been used in two applications (image processing and FIR filtering); the simulation results show that performance is substantially improved over TMR. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2017 | Design of Approximate Radix-4 Booth Multipliers for Error-Tolerant ComputingabstractApproximate computing is an attractive design methodology to achieve low power, high performance (low delay) and reduced circuit complexity by relaxing the requirement of accuracy. In this paper, approximate Booth multipliers are designed based on approximate radix-4 modified Booth encoding (MBE) algorithms and a regular partial product array that employs an approximate Wallace tree. Two approximate Booth encoders are proposed and analyzed for error-tolerant computing. The error characteristics are analyzed with respect to the so-called approximation factor that is related to the inexact bit width of the Booth multipliers. Simulation results at 45 nm feature size in CMOS for delay, area and power consumption are also provided. The results show that the proposed 16-bit approximate radix-4 Booth multipliers with approximate factors of 12 and 14 are more accurate than existing approximate Booth multipliers with moderate power consumption. The proposed R4ABM2 multiplier with an approximation factor of 14 is the most efficient design when considering both power-delay product and the error metric NMED. Case studies for image processing show the validity of the proposed approximate radix-4 Booth multipliers. Weiqiang Liu 0001, Liangyu Qian, Chenghua Wang, Honglan Jiang, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 5 |
| 2017 | A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing SystemsabstractIn this paper, a multi-objective, i.e., reliability, communication energy, performance, co-optimization model oriented mapping approach is proposed to find optimal mappings when applications are mapped onto network-on-chip (NoC) based reconfigurable architectures. A co-optimization model, defined as reliability efficiency model (REM), is developed to evaluate the overall reliability efficiency of a mapping. In REM, reliability efficiency is defined as the reliability profit at the same energy latency product. Based on REM, a mapping approach, referred to as priority and compensation factor oriented branch and bound (PCBB), is introduced to figure out the best mapping pattern. Two techniques, priority allocation and compensation factor utilization, are adopted to make a tradeoff between search efficiency and accuracy. Experimental results show that the proposed approach has three major contributions compared to state-of-the-art approaches. (1) PCBB is highly efficient in finding best mappings, with a 3x and 720x speedup compared to branch and bound (BB) and simulated annealing (SA). (2) PCBB is able to dynamically remap after the reconfiguration of the architecture. (3) General quantitative evaluation for reliability, communication energy and performance are made respectively before integrated into the unified model REM, whereas other similar models only touch upon two of them quantitatively. Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | A Design of a Non-Volatile PMC-Based (Programmable Metallization Cell) Register FileabstractThis paper presents the design of a non-volatile register file using cells made of a SRAM and a Programmable Metallization Cell (PMC). The proposed cell is a symmetric 8T2P (8-transistors, 2PMC) design; it utilizes three control lines to ensure the correctness in its operations (i.e. Write, Read, Store and Restore). Simulation results using HSPICE are provided for the cell as well as the register file array (both one- and two-dimensional schemes). At cell level, it is shown that the off-state resistance has a limited effect on the Read time, because in the proposed circuit the transistor connecting the PMCs to the SRAM is off. While having no significant effect on the Store time, the time of the Restore operation depends on the value of the off-state resistance, i.e. an increase in off-state PMC resistance causes an increase in Restore time. Comparison between non-volatile register files utilizing either PMCs, or Phase Change Memories (PCMs) is provided. The register file using PMCs has a faster Store and Read times than the PCM-based counterpart; this is mostly caused by the difference in resistance values for these two non-volatile technologies. The lower delay involved in these operations confirms that the proposed PMC-based register file offers significant advantages in terms of delay performance. Salin Junsangsri, Jie Han 0001, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | SRAM memory margin probability failure estimation using Gaussian Process regressionabstractEstimating the failure probabilities of SRAM memory cells using Monte Carlo or Importance Sampling techniques is expensive in the number of SPICE simulations needed. This paper presents a methodology for estimating the dynamic margin failure probabilities by building a surrogate model of the dynamic margin using Gaussian Process regression. Additive kernel functions that can extrapolate the margin values from the simulated samples are presented. These proposed kernel functions decrease the out-of-sample error of the surrogate model for a 6T cell by 32% compared with a six-dimensional universal kernel such as a Radial-Basis-Function kernel (RBF). Finally, the failure probability values predicted by a surrogate model built using 1250 SPICE simulations are reported and compared with Monte Carlo analysis with 106samples. The results show a relative error of 30% at 0.4V (predicted value of 4×10-6for the Monte Carlo estimate of 3×10-6) and a relative error of 172% at 0.3V (predicted value of 3×10-5for the Monte Carlo estimate of 1.1×10-5) for the dynamic read margin. These accuracy numbers are similar to those reported in previous proposals while the reduction in SPICE simulations is between 4× and 23× relative to these proposals and 800× compared to Monte Carlo method. Manish Rana, Ramon Canal, Jie Han 0001, Bruce F. Cockburn |
ICCD | 3 |
| 2016 | Design and evaluation of an approximate Wallace-Booth multiplierabstractApproximate or inexact computing has recently attracted considerable attention due to its potential advantages with respect to high performance and low power consumption. This paper presents the design of an approximate multiplier; this approximate multiplier consists of an approximate Booth encoder, an approximate 4-2 compressor and an approximate tree structure. The approximate design is implemented and verified for 8×8, 16×16 and 32×32-bit signed multiplication schemes targeting applications in embedded systems. Simulation results at 45 nm technology are provided and discussed. Compared with an exact Wallace-Booth multiplier as well as other approximate multipliers found in the technical literature, the proposed approximate scheme achieves significant improvements in power consumption, delay and combined metrics. These results show the viability of the proposed design. Liangyu Qian, Chenghua Wang, Weiqiang Liu 0001, Fabrizio Lombardi, Jie Han 0001 |
ISCAS | 5 |
| 2016 | Introduction to approximate computingabstractApproximate computing has emerged as a new paradigm for energy-efficient design of circuits and systems. This paper presents a brief introduction to approximate computing as well as to the challenges faced by approximate computing with respect to its prospects for applications in energy-efficient and error-resilient computing systems. Jie Han 0001 |
VTS | 1 |
| 2016 | Design of a hybrid non-volatile SRAM cell for concurrent SEU detection and correction
Pilin Junsangsri, Jie Han 0001, Fabrizio Lombardi |
Integr. | 2 |
| 2016 | On the Design of Approximate Restoring Dividers for Error-Tolerant ApplicationsabstractThis paper proposes several designs of approximate restoring dividers; two different levels of approximation (cell and array levels) are employed. Three approximate subtractor cells are utilized for integer subtraction as basic step of division; these cells tend to mitigate accuracy in subtraction with other metrics, such as circuit complexity and power dissipation. At array level, exact cells are either replaced or truncated in the approximate divider designs. A comprehensive evaluation of approximation at both cell- and array (divider) levels is pursued using error analysis and HSPICE simulation; different circuit metrics including complexity and power dissipation are evaluated. Different applications are investigated by utilizing the proposed approximate arithmetic circuits. The simulation results show that with extensive savings for power dissipation and circuit complexity, the proposed designs offer better error tolerant capabilities for quotient oriented applications (image processing) than remainder oriented application (modulo operations). The proposed approximate restoring divider is significantly better than the approximate non-restoring scheme presented in the technical literature. Linbin Chen, Jie Han 0001, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2016 | Approximate Radix-8 Booth Multipliers for Low-Power and High-Performance OperationabstractThe Booth multiplier has been widely used for high performance signed multiplication by encoding and thereby reducing the number of partial products. A multiplier using the radix-$4$(or modified Booth) algorithm is very efficient due to the ease of partial product generation, whereas the radix-$8$Booth multiplier is slow due to the complexity of generating the odd multiples of the multiplicand. In this paper, this issue is alleviated by the application of approximate designs. An approximate$2$-bit adder is deliberately designed for calculating the sum of$1\times$and$2\times$of a binary number. This adder requires a small area, a low power and a short critical path delay. Subsequently, the$2$-bit adder is employed to implement the less significant section of a recoding adder for generating the triple multiplicand with no carry propagation. In the pursuit of a trade-off between accuracy and power consumption, two signed$16\times 16$bit approximate radix-8 Booth multipliers are designed using the approximate recoding adder with and without the truncation of a number of less significant bits in the partial products. The proposed approximate multipliers are faster and more power efficient than the accurate Booth multiplier. The multiplier with 15-bit truncation achieves the best overall performance in terms of hardware and accuracy when compared to other approximate Booth multiplier designs. Finally, the approximate multipliers are applied to the design of a low-pass FIR filter and they show better performance than other approximate Booth multipliers. Honglan Jiang, Jie Han 0001, Fei Qiao, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2016 | Reliability Evaluation of Phased-Mission Systems Using Stochastic ComputationabstractA phased-mission system (PMS) usually consists of several nonoverlapping phases of tasks. All phases are required to be accomplished sequentially for a successful mission. Different features must be considered in the reliability evaluation of a PMS, including the dependence among the phases with respect to a common component and the different system topologies for the phases. To overcome the limitation of existing approaches, a stochastic computational approach is proposed for efficiently analyzing the reliability of a nonrepairable PMS. Stochastic logic models are proposed to analyze the common components in the different phases. In the stochastic analysis, the signal probabilities of the basic components are encoded as non-Bernoulli sequences of random permutations with fixed numbers of 1s and 0s. Thus, the proposed stochastic approach can be used to evaluate a PMS under any distribution. Based on the generated stochastic sequences for the basic components and the system topology, the failure probability of the PMS can be efficiently predicted. Several case studies are evaluated to show the accuracy and efficiency of the stochastic approach. Compared with a combinatorial analysis, the accuracy of the stochastic analysis varies with the length of the stochastic sequences. However, it is shown that the stochastic analysis is more efficient than a Monte Carlo simulation at the same execution complexity in the number of runs. Peican Zhu, Jie Han 0001, Leibo Liu, Fabrizio Lombardi |
IEEE Trans. Reliab. | 2 |
| 2016 | Logic-in-Memory With a Nonvolatile Programmable Metallization CellabstractThis paper introduces two new cells for logic-in-memory (LiM) operation. The first novelty of these cells is the resistive random access memory configuration that utilizes a programmable metallization cell as nonvolatile element. CMOS transistors and ambipolar transistors are used as processing and control elements for the logic operations of the LiM cells. The first cell employs ambipolar transistors and CMOS in its logic circuit (7T2A1P), while the second LiM cell uses only MOSFETs (9T1P) to implement logic functions, such as AND, OR, and XOR. The operational mode of the proposed cells is voltage-based, which is much different from the previous designs in which a LiM cell operates on a current mode. Extensive simulation results using HSPICE are provided for the evaluation of these cells; comparison shows that the proposed two cells outperform previous LiM cells in metrics, such as logic operation delays, power delay product, circuit complexity, write time, and output swing. Pilin Junsangsri, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Stochastic Circuit Design and Performance Evaluation of Vector Quantization for Different Error MeasuresabstractVector quantization (VQ) is a general data compression technique that has a scalable implementation complexity and potentially a high compression ratio. In this paper, a novel implementation of VQ using stochastic circuits is proposed and its performance is evaluated against conventional binary designs. The stochastic and binary designs are compared for the same compression quality, and the circuits are synthesized for an industrial 28-nm cell library. The effects of varying the sequence length of the stochastic representation are studied with respect to throughput per area (TPA) and energy per operation (EPO). The stochastic implementations are shown to have higher EPOs than the conventional binary implementations due to longer latencies. When a shorter encoding sequence with 512 bits is used to obtain a lower quality compression measured by the L1-norm, squared L2-norm, and third-law errors, the TPA ranges from 1.16 to 2.56 times than that of the binary implementation with the same compression quality. Thus, although the stochastic implementation underperforms for a high compression quality, it outperforms the conventional binary design in terms of TPA for a reduced compression quality. By exploiting the progressive precision feature of a stochastic circuit, a readily scalable processing quality can be attained by halting the computation after different numbers of clock cycles. Jie Han 0001, Bruce F. Cockburn, Duncan G. Elliott |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Stochastic circuit design and performance evaluation of vector quantizationabstractVector quantization (VQ) is a general data compression technique that has a scalable implementation complexity and potentially a high compression ratio. In this paper, a novel implementation of VQ using stochastic circuits is proposed and its performance is evaluated. The stochastic and binary designs are compared for the same compression quality and the circuits are synthesized for an industrial 28-nm cell library. The effects of varying the sequence length of the stochastic design are studied with respect to the performance metric of throughput per area (TPA). When a shortened 512-bit encoding sequence is used to obtain a lower quality compression, the TPA is about 2.60 times that of the binary implementation with the same quality as that of the stochastic implementation measured by the L1norm error (i.e., the first-order error). Thus, the stochastic implementation outperforms the conventional binary design in terms of TPA for a relatively low compression quality. By exploiting the progressive precision feature of a stochastic circuit, a readily scalable processing quality can be attained by simply halting the computation after different numbers of clock cycles. Jie Han 0001, Bruce F. Cockburn, Duncan G. Elliott |
ASAP | 2 |
| 2015 | A novel approach using a minimum cost maximum flow algorithm for fault-tolerant topology reconfiguration in NoC architecturesabstractAn approach using a minimum cost maximum flow algorithm is proposed for fault-tolerant topology reconfiguration in a Network-on-Chip system. Topology reconfiguration is converted into a network flow problem by constructing a directed graph with capacity constraints. A cost factor is considered to differentiate between processing elements. This approach maximizes the use of spare cores to repair faulty systems, with minimal impact on area, throughput and delay. It also provides a transparent virtual topology to alleviate the burden for operating systems. Leibo Liu, Chenchen Deng, Shouyi Yin, Shaojun Wei, Jie Han 0001 |
ASP-DAC | 6 |
| 2015 | An approximate voting scheme for reliable computing
Ke Chen 0018, Fabrizio Lombardi, Jie Han 0001 |
DATE | 3 |
| 2015 | Minimizing the number of process corner simulations during design verification
Michael Shoniker, Bruce F. Cockburn, Jie Han 0001, Witold Pedrycz |
DATE | 3 |
| 2015 | Design of Approximate Unsigned Integer Non-restoring Divider for Inexact ComputingabstractThis paper proposes several approximate divider designs; two different levels of approximation (cell and array levels) are investigated for non-restoring division. Three approximate subtractor cells are proposed and designed for the basic subtraction; these cells mitigate accuracy in subtraction with other metrics, such as circuit complexity and power dissipation. At array level, by considering the exact cells, both replacement and truncation schemes are introduced for approximate array divider design. A comprehensive evaluation of approximation at both cell and divider level is pursued. Different circuit metrics including complexity and power dissipation are evaluated by HSPICE simulation. Mean error distance (MED), normalized error distance (NED) and MED-power product (MPP) are provided to substantiate the accuracy and power trade-off of inexact computing. Different applications in image processing are investigated by utilizing the proposed approximate arithmetic circuits. Linbin Chen, Jie Han 0001, Weiqiang Liu 0001, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | A Comparative Review and Evaluation of Approximate AddersabstractAs an important arithmetic module, the adder plays a key role in determining the speed and power consumption of a digital signal processing (DSP) system. The demands of high speed and power efficiency as well as the fault tolerance nature of some applications have promoted the development of approximate adders. This paper reviews current approximate adder designs and provides a comparative evaluation in terms of both error and circuit characteristics. Simulation results show that the equal segmentation adder (ESA) is the most hardware-efficient design, but it has the lowest accuracy in terms of error rate (ER) and mean relative error distance (MRED). The error-tolerant adder type II (ETAII), the speculative carry select adder (SCSA) and the accuracy-configurable approximate adder (ACAA) are equally accurate (provided that the same parameters are used), however ETATII incurs the lowest power-delay-product (PDP) among them. The almost correct adder (ACA) is the most power consuming scheme with a moderate accuracy. The lower-part-OR adder (LOA) is the slowest, but it is highly efficient in power dissipation. Honglan Jiang, Jie Han 0001, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | A Ternary Content Addressable Cell Using a Single Phase Change Memory (PCM)abstractThis paper presents the novel design of a Ternary Content Addressable Memory (TCAM); different from existing designs found in the technical literature, this cell utilizes a single Phase Change Memory (PCM) as storage element and ambipolarity for comparison. A memory core consisting of a CMOS transistor and a PCM is employed (1T1P); for the search operation, the data in the 1T1P memory core is read and its value is established using two differential sense amplifiers. Compared with other non-volatile memory cells using emerging technologies (such as PCM-based, and memristor-based), simulation results show that the proposed non-volatile TCAM cell offer significant advantages in terms of power dissipation, PDP for the search operation, write time and reduced circuit complexity (in terms of lower counts in transistors and storage elements). Pilin Junsangsri, Fabrizio Lombardi, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Reliability-aware mapping for various NoC topologies and routing algorithms under performance constraints
Chenchen Deng, Leibo Liu, Shouyi Yin, Jie Han 0001, Shaojun Wei |
Sci. China Inf. Sci. | 5 |
| 2015 | An Analytical Framework for Evaluating the Error Characteristics of Approximate AddersabstractApproximate adders have been considered as a potential alternative for error-tolerant applications to trade off some accuracy for gains in other circuit-based metrics, such as power, area and delay. Existing approximate adder designs have shown substantial advantages in improving many of these operational features. However, the error characteristics of the approximate adders still remain an issue that is not very well understood. A simulation-based method requires both programming efforts and a time-consuming execution for evaluating the effect of errors. This method becomes particularly expensive when dealing with various sizes and types of approximate adders. In this paper, a framework based on analytical models is proposed for evaluating the error characteristics of approximate adders. Error features such as the error rate and the mean error distance are obtained using this framework without developing functional models of the approximate adders for time-consuming simulation. As an example, the estimate of peak signal-to-noise ratios (PSNRs) in image processing is considered to show the potential application of the proposed analysis. This analytical framework provides an efficient method to evaluate various designs of approximate adders for meeting different figures of merit in error-tolerant applications. Cong Liu 0015, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2015 | Design and Analysis of Approximate Compressors for MultiplicationabstractInexact (or approximate) computing is an attractive paradigm for digital processing at nanometric scales. Inexact computing is particularly interesting for computer arithmetic designs. This paper deals with the analysis and design of two new approximate 4-2 compressors for utilization in a multiplier. These designs rely on different features of compression, such that imprecision in computation (as measured by the error rate and the so-called normalized error distance) can meet with respect to circuit-based figures of merit of a design (number of transistors, delay and power consumption). Four different schemes for utilizing the proposed approximate compressors are proposed and analyzed for a Dadda multiplier. Extensive simulation results are provided and an application of the approximate multipliers to image processing is presented. The results show that the proposed designs accomplish significant reductions in power dissipation, delay and transistor count compared to an exact design; moreover, two of the proposed multiplier designs provide excellent capabilities for image multiplication with respect to average normalized error distance and peak signal-to-noise ratio (more than 50 dB for the considered image examples). Amir Momeni, Jie Han 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2015 | An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC ArchitecturesabstractIn this paper, an efficient application mapping approach is proposed for the co-optimization of reliability, communication energy, and performance (CoREP) in network-on-chip (NoC)-based reconfigurable architectures. A cost model for the CoREP is developed to evaluate the overall cost of a mapping. In this model, communication energy and latency (as a measure of performance) are first considered in energy latency product (ELP), and then ELP is co-optimized with reliability by a weight parameter that defines the optimization priority. Both transient and intermittent errors in NoC are modeled in CoREP. Based on CoREP, a mapping approach, referred to as priority and ratio oriented branch and bound (PRBB), is proposed to derive the best mapping by enumerating all the candidate mappings organized in a search tree. Two techniques, branch node priority recognition and partial cost ratio utilization, are adopted to improve the search efficiency. Experimental results show that the proposed approach achieves significant improvements in reliability, energy, and performance. Compared with the state-of-the-art methods in the same scope, the proposed approach has the following distinctive advantages: 1) CoREP is highly flexible to address various NoC topologies and routing algorithms while others are limited to some specific topologies and/or routing algorithms; 2) general quantitative evaluation for reliability, energy, and performance are made, respectively, before being integrated into unified cost model in general context while other similar models only touch upon two of them; and 3) CoREP-based PRBB attains a competitive processing speed, which is faster than other mapping approaches. Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | A Stochastic Approach for the Analysis of Dynamic Fault Trees With Spare Gates Under Probabilistic Common Cause FailuresabstractA redundant system usually consists of primary and standby modules. The so-called spare gate is extensively used to model the dynamic behavior of redundant systems in the application of dynamic fault trees (DFTs). Several methodologies have been proposed to evaluate the reliability of DFTs containing spare gates by computing the failure probability. However, either a complex analysis or significant simulation time are usually required by such an approach. Moreover, it is difficult to compute the failure probability of a system with component failures that are not exponentially distributed. Additionally, probabilistic common cause failures (PCCFs) have been widely reported, usually occurring in a statistically dependent manner. Failure to account for the effect of PCCFs overestimates the reliability of a DFT. In this paper, stochastic computational models are proposed for an efficient analysis of spare gates and PCCFs in a DFT. Using these models, a DFT with spare gates under PCCFs can be efficiently evaluated. In the proposed stochastic approach, a signal probability is encoded as a non-Bernoulli sequence of random permutations of fixed numbers of ones and zeros. The component's failure probability is not limited to an exponential distribution, thus this approach is applicable to a DFT analysis in a general case. Several case studies are evaluated to show the accuracy and efficiency of the proposed approach, compared to both an analytical approach and Monte Carlo (MC) simulation. Peican Zhu, Jie Han 0001, Leibo Liu, Fabrizio Lombardi |
IEEE Trans. Reliab. | 2 |
| 2015 | Efficient Fault-Tolerant Topology Reconfiguration Using a Maximum Flow AlgorithmabstractWith an increasing number of processing elements (PEs) integrated on a single chip, fault-tolerant techniques are critical to ensure the reliability of such complex systems. In current reconfigurable architectures, redundant PEs are utilized for fault tolerance. In the presence of faulty PEs, the physical topologies of various chips may be different, so the concept of virtual topology from network embedding problem has been used to alleviate the burden for the operating systems. With limited hardware resources, how to reconfigure a system into the most effective virtual topology such that the maximum repair rate can be reached presents a significant challenge. In this article, a new approach using a maximum flow (MF) algorithm is proposed for an efficient topology reconfiguration in reconfigurable architectures. In this approach, topology reconfiguration is converted into a network flow problem by constructing a directed graph; the solution is then found by using the MF algorithm. This approach optimizes the use of spare PEs with minimal impacts on area, throughput, and delay, and thus it significantly improves the repair rate of faulty PEs. In addition, it achieves a polynomial reconfiguration time. Experimental results show that compared to previous methods, the MF approach increases the probability to repair faulty PEs by up to 50% using the same redundant resources. Compared to a fault-free system, the throughput only decreases by less than 2.5% and latency increases by less than 4%. To consider various types of PEs in a practical application, a cost factor is introduced into the MF algorithm. An enhanced approach using a minimum-cost MF algorithm is further shown to be efficient in the fault-tolerant reconfiguration of heterogeneous reconfigurable architectures. Leibo Liu, Shouyi Yin, Jie Han 0001, Shaojun Wei |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2015 | On the Nonvolatile Performance of Flip-Flop/SRAM Cells With a Single MTJabstractIn this brief, three nonvolatile flip-flop (FF)/SRAM cells that utilize a single magnetic tunneling junction (MTJ) as nonvolatile resistive element are proposed. These cells have the same core (i.e., 6T) but they employ different numbers of MOSFETs to implement the so-called instantly ON, normally OFF mode of operation. The additional transistors are utilized for the restore operation to ensure that the data stored in the nonvolatile circuitry can be written back into the FF core once the power is made available. These three cells (7T, 9T, and 11T) are extensively analyzed in terms of their operations in 32 nm technology, such as operational delays (for the write, read, and restore operations), the static noise margin (SNM), critical charge and process variations (in both the MOSFETs and the resistive element). Simulation results show that an increase in the number of MOSFETs in the cells causes improvements in critical charge and tolerance to process variations at the expense of an increase in power dissipation. The SNM and the delay of the restore operation, however, do not necessarily increase with the number of MOSFETs in the cell, but rather on the control of access to the storage nodes from the single MTJ. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | On the Restore Operation in MTJ-Based Nonvolatile SRAM CellsabstractThis brief investigates the Restore mechanism of a nonvolatile static random access memory (NVSRAM) cell that utilizes two magnetic tunneling junctions (MTJs) as nonvolatile resistive elements and a 6T SRAM core. Two cells are proposed by employing different mechanisms for the Restore operation once the power is reestablished. The proposed cells use the bitline and supply as mechanisms to initiate the Restore operation, so connecting the two MTJs to different nodes of the NVSRAM circuitry. The cells are extensively analyzed in terms of their operations with respect to different figures of merit, such as operational delays (for the Write, Read, and Restore operations), the static noise margin, power consumption, critical charge, and process variations (in both the MOSFETs and the resistive elements). Simulation results show that the cell with the MTJs connected to the supply offers the best performance in terms of power for the Read/Restore operations; it also achieves the best Read delay, but the worst Restore delay. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Fault-Tolerant Technique Using Quadded Logic and Quadded TransistorsabstractAdvances in CMOS technology have made digital circuits and systems very sensitive to manufacturing variations, aging, and/or soft errors. Fault-tolerant techniques using hardware redundancy have been extensively investigated for improving reliability. Quadded logic (QL) is an interwoven redundant logic technique that corrects errors by switching them from critical to subcritical status; however, QL cannot correct errors in the last one or two layers of a circuit. In contrast to QL, quadded transistor (QT) corrects errors while performing the function of a circuit. In this brief, a technique that combines QL with QT is proposed to take advantage of both techniques. The proposed quadded logic with quadded transistor (QLQT) technique is evaluated and compared with other fault-tolerant techniques, such as triple modular redundancy and triple interwoven redundancy, using stochastic computational models. Simulation results show that QLQT has a better reliability than the other fault-tolerant techniques (except in the very restrictive case of small circuits with low gate error rates and very short paths from primary inputs to primary outputs). These results provide a new insight for implementing efficient fault-tolerant techniques in the design of reliable circuits and systems. Jie Han 0001, Eugene Leung, Leibo Liu, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A Flexible Energy- and Reliability-Aware Application Mapping for NoC-Based Reconfigurable ArchitecturesabstractThis paper proposes a flexible energy- and reliability-aware application mapping approach for network-on-chip (NoC)-based reconfigurable architecture. A parameterized cost model is first developed by combining energy and reliability with a weight parameter that defines the optimization priority. Using this model, the overall mapping cost could be evaluated. Subsequently, a mapping method using branch and bound with a partial cost ratio is employed to find the best mapping by enumerating all the possible patterns organized in a search tree. To improve the search efficiency, nonoptimal mappings are discarded at early stages using the partial cost ratio. Using the proposed approach, applications can be mapped onto most NoC topologies and running with various routing algorithms when considering both energy and reliability. Other state-of-the-art works have also done substantial research for the same topic but only limited to a specific topology or routing algorithm. Even for the same topology and routing algorithm, the proposed approach still shows considerable advantages in many aspects. Experiments show that this approach gains not only significant reduction in energy but also improvement in reliability. It also outperforms other approaches in throughput and latency with competitive run time. Leibo Liu, Chenchen Deng, Shouyi Yin, Jie Han 0001, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2014 | A hybrid non-volatile SRAM cell with concurrent SEU detection and correctionabstractThis paper presents a hybrid non-volatile (NV) SRAM cell with a new scheme for SEU tolerance. The proposed NVSRAM cell consists of a 6T SRAM core and a Resistive RAM (RRAM), made of a 1T and a Programmable Metallization Cell (PMC). The proposed cell has concurrent error detection (CED) and correction capabilities; CED is accomplished using a dual-rail checker, while correction is accomplished by utilizing the restore operation; data from the non-volatile memory element is copied back to the SRAM core. The dual-rail checker utilizes two XOR gates each made of 2 inverters and 2 ambipolar transistors, hence, it has a hybrid nature. Extensive simulation results are provided. The simulation results show that the proposed scheme is very efficient in terms of numerous figures of merit such as delay and circuit complexity and thus applicable to integrated circuits such as FPGAs requiring secure on-chip non-volatile storage (i.e. LUTs) for multi-context configurability. Pilin Junsangsri, Fabrizio Lombardi, Jie Han 0001 |
DATE | 3 |
| 2014 | A low-power, high-performance approximate multiplier with configurable partial error recoveryabstractApproximate circuits have been considered for error-tolerant applications that can tolerate some loss of accuracy with improved performance and energy efficiency. Multipliers are key arithmetic circuits in many such applications such as digital signal processing (DSP). In this paper, a novel approximate multiplier with a lower power consumption and a shorter critical path than traditional multipliers is proposed for high-performance DSP applications. This multiplier leverages a newly-designed approximate adder that limits its carry propagation to the nearest neighbors for fast partial product accumulation. Different levels of accuracy can be achieved through a configurable error recovery by using different numbers of most significant bits (MSBs) for error reduction. The approximate multiplier has a low mean error distance, i.e., most of the errors are not significant in magnitude. Compared to the Wallace multiplier, a 16-bit approximate multiplier implemented in a 28nm CMOS process shows a reduction in delay and power of 20% and up to 69%, respectively. It is shown that by utilizing an appropriate error recovery, the proposed approximate multiplier achieves similar processing accuracy as traditional exact multipliers but with significant improvements in power and performance. Cong Liu 0015, Jie Han 0001, Fabrizio Lombardi |
DATE | 2 |
| 2014 | A Stochastic Computational Approach for Accurate and Efficient Reliability EvaluationabstractReliability is fast becoming a major concern due to the nanometric scaling of CMOS technology. Accurate analytical approaches for the reliability evaluation of logic circuits, however, have a computational complexity that generally increases exponentially with circuit size. This makes intractable the reliability analysis of large circuits. This paper initially presents novel computational models based on stochastic computation; using these stochastic computational models (SCMs), a simulation-based analytical approach is then proposed for the reliability evaluation of logic circuits. In this approach, signal probabilities are encoded in the statistics of random binary bit streams and non-Bernoulli sequences of random permutations of binary bits are used for initial input and gate error probabilities. By leveraging the bit-wise dependencies of random binary streams, the proposed approach takes into account signal correlations and evaluates the joint reliability of multiple outputs. Therefore, it accurately determines the reliability of a circuit; its precision is only limited by the random fluctuations inherent in the stochastic sequences. Based on both simulation and analysis, the SCM approach takes advantages of ease in implementation and accuracy in evaluation. The use of non-Bernoulli sequences as initial inputs further increases the evaluation efficiency and accuracy compared to the conventional use of Bernoulli sequences, so the proposed stochastic approach is scalable for analyzing large circuits. It can further account for various fault models as well as calculating the soft error rate (SER). These results are supported by extensive simulations and detailed comparison with existing approaches. Jie Han 0001, Jinghang Liang, Peican Zhu, Zhixi Yang, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2014 | A Stochastic Approach for the Analysis of Fault Trees With Priority AND GatesabstractDynamic fault tree (DFT) analysis has been used to account for dynamic behaviors such as the sequence-dependent, functional-dependent, and priority relationships among the failures of basic events. Various methodologies have been developed to analyze a DFT; however, most methods require a complex analytical procedure or a significant simulation time for an accurate analysis. In this paper, a stochastic computational approach is proposed for an efficient analysis of the top event's failure probability in a DFT with priority AND (PAND) gates. A stochastic model is initially proposed for a two-input PAND gate, and a successive cascading model is then presented for a general multiple-input PAND gate. A stochastic approach using the proposed models provides an efficient analysis of a DFT compared to an accurate analysis or algebraic approach. The accuracy of a stochastic analysis increases with the length of random binary bit streams in stochastic computation. The use of non-Bernoulli sequences of random permutations of fixed counts of 1s and 0s as initial input events' probabilities makes the stochastic approach more efficient, and more accurate than Monte Carlo simulation. Non-exponential failure distributions and repeated events are readily handled by the stochastic approach. The accuracy, efficiency, and scalability of the stochastic approach are shown by several case studies of DFT analysis. Peican Zhu, Jie Han 0001, Leibo Liu, Mingjian Zuo |
IEEE Trans. Reliab. | 2 |
| 2013 | Approximate computing: An emerging paradigm for energy-efficient designabstractApproximate computing has recently emerged as a promising approach to energy-efficient design of digital systems. Approximate computing relies on the ability of many systems and applications to tolerate some loss of quality or optimality in the computed result. By relaxing the need for fully precise or completely deterministic operations, approximate computing techniques allow substantially improved energy efficiency. This paper reviews recent progress in the area, including design of approximate arithmetic blocks, pertinent error and quality measures, and algorithm-level techniques for approximate computing. Jie Han 0001, Michael Orshansky |
ETS | 1 |
| 2013 | A VLSI architecture for enhancing the fault tolerance of NoC using quad-spare mesh topology and dynamic reconfigurationabstractEffective fault tolerant techniques are crucial for a Network-on-Chip (NoC) to achieve reliable communication. In this paper, a novel VLSI architecture employing redundant routers is proposed to enhance the fault tolerance of an NoC. The NoC mesh is divided into blocks of 2×2 routers with a spare router placed in the center. The proposed fault-tolerant architecture, referred to as a quad-spare mesh, can be dynamically reconfigured by changing control signals without altering the underlying topology. This dynamic reconfiguration and its corresponding routing algorithm are demonstrated in detail. Experimental results show that the proposed design achieves significant improvements on reliability compared with those reported in the literature. Leibo Liu, Shouyi Yin, Shaojun Wei, Jie Han 0001 |
ISCAS | 6 |
| 2013 | A fault tolerant NoC architecture using quad-spare mesh topology and dynamic reconfiguration
Leibo Liu, Shouyi Yin, Jie Han 0001, Shaojun Wei |
J. Syst. Archit. | 4 |
| 2013 | Analysis of Error Masking and Restoring Properties of Sequential CircuitsabstractScaling of CMOS technology into nanometric feature sizes has raised concerns for the reliable operation of logic circuits, such as in the presence of soft errors. This paper deals with the analysis of the operation of sequential circuits. As the feedback signals in a sequential circuit can be logically masked by specific combinations of primary inputs, the cumulative effects of soft errors can be eliminated. This phenomenon, referred to as error masking, is related to the presence of so-called restoring inputs and/or the consecutive presence of specific inputs in multiple clock cycles (equivalent to a synchronizing sequence in switching theory). In this paper, error masking is extensively analyzed using the operations of state transition matrices (STMs) and binary decision diagrams (BDDs) of a finite state machine (FSM) model. The characteristics of state transitions with respect to correlations between the restoring inputs and time sequence are mathematically established using STMs; although the applicability of the STM analysis is restricted due to its complexity, the BDD approach is more efficient and scalable for use in the analysis of large circuits. These results are supported by simulations of benchmark circuits and may provide a basis for further devising efficient and robust implementations when designing FSMs. Jinghang Liang, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2013 | New Metrics for the Reliability of Approximate and Probabilistic AddersabstractAddition is a fundamental function in arithmetic operation; several adder designs have been proposed for implementations in inexact computing. These adders show different operational profiles; some of them are approximate in nature while others rely on probabilistic features of nanoscale circuits. However, there has been a lack of appropriate metrics to evaluate the efficacy of various inexact designs. In this paper, new metrics are proposed for evaluating the reliability as well as the power efficiency of approximate and probabilistic adders. Reliability is analyzed using the so-called sequential probability transition matrices (SPTMs). Error distance (ED) is initially defined as the arithmetic distance between an erroneous output and the correct output for a given input. The mean error distance (MED) and normalized error distance (NED) are then proposed as unified figures that consider the averaging effect of multiple inputs and the normalization of multiple-bit adders. It is shown that the MED is an effective metric for measuring the implementation accuracy of a multiple-bit adder and that the NED is a nearly invariant metric independent of the size of an adder. The MED is, therefore, useful in assessing the effectiveness of an approximate or probabilistic adder implementation, while the NED is useful in characterizing the reliability of a specific design. Since inexact adders are often used for saving power, the product of power and NED is further utilized for evaluating the tradeoffs between power consumption and precision. Although illustrated using adders, the proposed metrics are potentially useful in assessing other arithmetic circuit designs for applications of inexact computing. Jinghang Liang, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2012 | Modeling a single electron turnstile in HSPICEabstractThis paper presents a novel HSPICE circuit model for designing and simulating a Single-Electron (SE) turnstile, as applicable at the nanometric feature sizes. The proposed SE model consists of two nearly similar parts whose operation is independent of each other; this disjoint feature permits to accurately model the sequential transfer of electrons through the turnstile in the storage node (modeled on a voltage level basis). It therefore avoids the transient (current-based) nature of a previous model. The model has been simulated and results show that it can correctly operate at 32 nm with excellent stability in its operation. Extensive simulation results are presented to substantiate the advantages of using the proposed model with respect to changes in the circuit model parameter. Fabrizio Lombardi, Wei Wei 0034, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | Stochastic computational models for accurate reliability evaluation of logic circuitsabstractAs reliability becomes a major concern with the continuous scaling of CMOS technology, several computational methodologies have been developed for the reliability evaluation of logic circuits. Previous accurate analytical approaches, however, have a computational complexity that generally increases exponentially with the size of a circuit, making the evaluation of large circuits intractable. This paper presents novel computational models based on stochastic computation, in which probabilities are encoded in the statistics of random binary bit streams, for the reliability evaluation of logic circuits. A computational approach using the stochastic computational models (SCMs) accurately determines the reliability of a circuit with its precision only limited by the random fluctuations inherent in the representation of random binary bit streams. The SCM approach has a linear computational complexity and is therefore scalable for use for any large circuits. Our simulation results demonstrate the accuracy and scalability of the SCM approach, and suggest its possible applications in VLSI design. Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2005 | Faults, Error Bounds and Reliability of Nanoelectronic CircuitsabstractThis paper is concerned with faults, error bounds and reliability modeling of nanotechnology-based circuits. First, we briefly review failure mechanisms and fault models in nanoelectronics. Second, reliability functions based on probabilistic models are developed for unreliable logic gates. We then show that fundamental gate error bounds for general probabilistic computation can be derived using the nonlinear mapping functions constructed from the gate models. Finally, an analytical approach is proposed for estimating reliabilities of nanoelectronic circuits. This approach is based on the probabilistic modeling of unreliable logic gates and interconnects. In spite of the approximations used in probabilistic modeling, our study suggests that the proposed approach provides a simple and efficient way to model the reliability of nanoelectronic circuits. Jie Han 0001, Erin Taylor 0001, José A. B. Fortes |
ASAP | 1 |