EDBT 2026 Demo / reviewers in the wild / expert
Honglan Jiang
dblp:163/0012
· DBLP profile ↗
41ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0003-3705-4240ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 7 first-author · 28 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-ANT: Efficient Low-Bit Group-Wise Quantization for Large Language Models via Sign-Asymmetric Adaptive Numeric TypeabstractLarge language models (LLMs) have demonstrated remarkable potential across diverse domains; meanwhile, their large parameter sizes pose substantial inference costs, motivating the need for efficient low-bit quantization. Group-wise quantization, which adopts finer granularity, has been widely used to improve low-bit quantization performance. Several adaptive numeric types have been proposed to further enhance low-bit group-wise quantization; however, they construct quantization grids based on symmetric numeric types, which limits their ability to model asymmetric distributions. To address this limitation, we propose SA-ANT, a sign-asymmetric adaptive numeric type for efficient low-bit group-wise quantization. SA-ANT constructs quantization grids separately on the positive and negative sides, enabling adaptive support for asymmetric and non-uniform distributions. Furthermore, the carefully designed SA-ANT not only reduces quantization errors but also ensures a unified computing across different sub numeric types, thereby facilitating hardware efficiency. To accelerate LLM inference, we develop (1) a quantization framework that transforms LLM weights into the SA-ANT and adaptively selects the sub numeric type for each group, and (2) an accelerator that maps SA-ANT inference to low-bit INT operations. Experimental results show that SA-ANT delivers 3.92%– 5.57% higher accuracy than state-of-the-art adaptive numeric types under 3-bit weight quantization, while also enabling 7.84%– 44.65% area savings and 7.80%–43.88% power reductions. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 5 |
| 2026 | Bridging the Power Estimation Gap: A GNN-Based Prediction Model for Approximate Logic SynthesisabstractApproximate computing is an application-related paradigm that trades limited accuracy for improvements in hardware cost. As a key technique of approximate computing, approximate logic synthesis (ALS) automatically generates approximate circuits with reduced area, power, and delay while satisfying predefined quality-of-result (QoR) constraints. However, in typical gate-level ALS workflows, synthesis tools are invoked at the final stage for optimization, leading to a discrepancy between the circuit in design space exploration (DSE) and the final obtained circuit. Thus, the power estimation for a candidate circuit during DSE may exhibit a significant gap from the actual power consumed by its post-synthesis circuit. This gap may mislead the DSE to a sub-optimal design. To address this issue, we propose a graph neural network (GNN)-based power prediction model that operates on gate-level circuits. The model incorporates multi-head channel attention, which extracts high-level topological and functional features that correlate with power dissipation and implicitly captures the optimization behavior of synthesis tools. Thus, it enables a direct prediction of post-synthesis power from pre-synthesis gate-level circuits. Experimental results show that the proposed model improves the concordance index (C-index) for power ranking by up to 14.0% over traditional methods. Furthermore, we construct an ALS framework by integrating the proposed model with Cartesian genetic programming (CGP). Compared to state-of-the-art ALS approaches, our GNN-CGP framework generates circuits with up to 26.8% power savings under the same error constraints. Fuxuan Li, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 6 |
| 2026 | HAP: Accelerating DNNs with Resolution-Preserved Quantization by Harnessing Adaptive-PrecisionabstractReducing the precision in post-training quantization can cause catastrophic accuracy loss in Deep Neural Networks, especially when compressing the activations. To address this problem, we present a novel adaptive-precision quantization (APQ) and accelerator design that achieves lossless activation compression by exploiting the inherent coding redundancy. Compared to existing APQ methods, this design can be generalized to implement asymmetric quantization, making it particularly suitable for activations. The accelerator offers a practical solution to mitigate the computational workload imbalance problem incurred by variable precision. A dual-precision quantization scheme further provides the flexibility to trade off accuracy and performance. Erjing Luo, Xinkuang Geng, Honglan Jiang, Leibo Liu, Jie Han 0001 |
DATE | 3 |
| 2026 | ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic Computing
Kang You, Chen Nie, Lee Jun Yan, Ziling Wei, Yu Feng 0007, Honglan Jiang, Zhezhi He |
ISCA | 8 |
| 2026 | A 10-Bit Successive Reference Approximation ADC Enabled By Switched-Capacitor Dynamic Reference Generation With Pipelined Pre-charging And Voltage Ripple Cancellation
Hanlin Xu, Runtao Huo, Honglan Jiang, Minmin You, Hui Wang 0023 |
ISCAS | 4 |
| 2026 | Approximate Signed Multiplier Designs for Efficient CNN Inference
Mengshuo Zhang, Xiaolu Hu, Xinkuang Geng, Honglan Jiang |
ISCAS | 5 |
| 2026 | LUT-ALMs: Trading Off Accuracy and Power for Approximate Logarithmic Multipliers via LUT OptimizationabstractLogarithmic multiplier (LM) converts fixed-point (FxP) input operands to logarithmic numbers and performs multiplication with simple shift and addition operations, which achieves distinct power reduction, yet with significant single-sided errors. This paper proposes to fuse error compensation with logarithmic conversion by using customized look-up tables (LUTs). To avoid the use of large LUTs, partition strategies are designed for the optimization of LUTs. In addition, to effectively balance the accuracy and hardware costs, two iterative algorithms are proposed for generating precision-configurable LUTs. Based on the optimized LUTs, high-accuracy and low-power approximate LMs (LUT-ALMs) are constructed for 8-bit and 16-bit multiplications. Furthermore, to enhance the flexibility of data type and bit width, a mixed-mode LM (MM-ALM) supporting eight multiplication modes is devised. Compared with the exact 8-bit signed multiplier from Synopsys DesignWare (DW) library (DW-Exact), LUT-ALMs show up to 30.96% and 23.99% reductions in the power-delay product (PDP) and area-delay product (ADP), respectively, with a mean relative error distance (MRED) of 3.45%. Compared with state-of-the-art approximate 8- bit signed multipliers, LUT-ALMs form the Pareto front in terms of PDP and MRED. For 16-bit multiplication, LUT-ALMs obtain up to 70.90% and 62.13% savings in PDP and ADP with a MRED of 2.75%, compared with the corresponding DW-Exact. Compared with the corresponding mixed-mode exact multiplier constructed of DW multipliers, MM-ALM performing 8-bit multiplication achieves up to 57.44% and 47.99% reductions in PDP and ADP, respectively. When performing 4-bit multiplications, MM-ALM can save up to 37.16% savings in PDP. With lower hardware over-heads, LUT-ALMs and MM-ALM present comparable accuracy to the corresponding exact designs in the considered convolutional neural networks (CNNs) and image processing applications. The hardware description of the devised LMs is open-sourced athttps://anonymous.4open.science/r/LUT-ALM-5FD1. Xinkuang Geng, Xiaolu Hu, Hui Wang 0023, Jianfei Jiang 0001, Qin Wang 0009, Siting Liu 0001, Jie Han 0001, Honglan Jiang |
IEEE Trans. Computers | 9 |
| 2026 | Glitch-Aware Optimization of 4-2 Compressor-Based Approximate MultipliersabstractApproximate multipliers have widely been used in error-resilient applications with hardware-efficient and relaxed precision requirements, such as multimedia signal processing and deep learning. Aiming to achieve improvements in power efficiency and performance, approximate multipliers are generally designed by simplifying the implementation circuits. However, the redundant switching activities (also known as glitches) due to unbalanced signal paths are seldom considered, leading to significant dynamic power. Moreover, prior approximate designs often overlook transistor sizing optimization, which limits their potential for power–delay efficiency. This article proposes a custom design and optimization framework for approximately 4-2 compressors at the structural and circuit levels, which effectively reduces glitches by balancing output delays. Based on the devised approximate 4-2 compressors, the constructed multipliers can then achieve significant reductions in the generation and propagation of spurious activities. In addition, to further reduce the dynamic power consumption, a delay-aware signal routing (DASR) strategy is introduced for interconnecting approximate compressors. The stability and efficiency of the proposed designs under varying conditions are verified by extensive simulations. The experimental results on HLMC 28-nm CMOS technology show that the approximate 4-2 compressors obtained by the proposed framework achieve 10.4%–52.1% power–delay product (PDP) reductions compared to existing approximate designs with the same accuracy, resulting in 5.20%–33.1% PDP improvements for multipliers. Moreover, the proposed optimization framework is generalizable to arbitrary approximate 4-2 compressor designs. Tongjing Wu, Honglan Jiang, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | EPIC: Error PredIction and Correction for Power-Efficient Voltage Underscaling Multiply-Accumulate UnitabstractMatrix multiplication dominates the power consumption in compute-intensive applications such as deep neural networks (DNNs), spurring intensive investigations into power-efficient multiply-accumulate (MAC) units. Among the mainstream low-power design methodologies, voltage underscaling can achieve effective power savings yet induce timing errors that may lead to catastrophic accuracy loss. In this paper, we propose an error prediction and correction framework (denoted as EPIC) for arbitrary MAC unit under voltage underscaling, which predicts the timing errors and samples the correct output by using a delay-tunable clock. A prediction bits searching algorithm is proposed to enhance the prediction accuracy with low hardware cost, resulting in up to 100% accuracy. While preserving the accuracy, EPIC achieves up to 52% power savings over the corresponding MAC operating at nominal voltage. With transistor-level optimizations, EPIC incurs only 8% area and 1% power overheads, achieving 100% error correction under a voltage underscaling ratio of $\mathbf{0. 7 4}$. Compared to state-of-the-art error resilient circuit designs, EPIC consumes 60%-88% less area. Additionally, to achieve the accuracy performance of EPIC in error-resilient applications, we propose a simulation workflow involving precise timing features, enabling an accurate simulation of voltage underscaling MAC in large-scale applications. The experimental results show that, under voltage-underscaling, the MAC with EPIC consumes 11% less power than the one without EPIC, when a same accuracy as exact implementation is required in multi-layer perceptron (MLP). Tongjing Wu, Xiaolu Hu, Siting Liu 0001, Hui Wang 0023, Weifeng He, Zhigang Mao, Honglan Jiang |
DAC | 8 |
| 2025 | Lookup Table Refactoring: Towards Efficient Logarithmic Number System Addition for Large Language ModelsabstractCompared to integer quantization, logarithmic quantization aligns more effectively with the long-tailed distribution of data in large language models (LLMs), resulting in lower quantization errors. Moreover, the logarithmic number system (LNS) employs a fixed-point adder to perform multiplication, indicating a potential reduction in computational complexity for LLM accelerators that require extensive multiply-accumulate (MAC) operations. However, a key bottleneck is that LNS addition requires complex nonlinear functions, which are typically approximated using lookup tables (LUTs). This study aims to reduce the hardware resources needed for LUTs in LNS addition while maintaining high precision. Specifically, we investigate the specific nature of addition operations within LLMs; the relationship between the hardware parameters of the LUT and the computing errors is then mathematically derived. Based on these insights, we propose LUT refactoring to optimize the LUT for enhanced efficiency in LNS addition. With 10.93% and 19.78% reductions in area-delay product (ADP) and power-delay product (PDP), respectively, LUT refactoring results in an accuracy improvement of up to 33.5% in LLM benchmarks compared to the naive design. When compared to integer quantization, our method achieves higher accuracy while reducing area by 18.27% and power by 42.61%. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 5 |
| 2025 | Segment-Wise Accumulation: Low-Error Logarithmic Domain Computing for Efficient Large Language Model InferenceabstractLogarithmic domain computing (LDC) has great potential for reducing quantization errors and computational complexity in Large Language Models (LLMs). While logarithmic multiplication can be efficiently implemented using fixed-point addition, the primary challenge in multiply-accumulate (MAC) operations is balancing the precision of logarithmic adders with their hardware overhead. Through a detailed analysis of the errors inherent in LDC-based LLMs, we propose segment-wise accumulation (SWA) to mitigate these errors. In addition, a processing element (PE) is introduced to enable SWA in the systolic array architecture. Compared with the accumulation scheme devised for enhancing floating-point computing, the proposed SWA facilitates the integration into existing accelerator architectures, resulting in lower hardware overhead. The experimental results show that SWA allows LDC under low-precision configurations to achieve remarkable accuracy in LLMs, demonstrating higher hardware efficiency than merely increasing the precision of individual computations. Our method, while maintaining a lower hardware overhead than traditional LDC, achieves more than 13.9% improvement in average accuracy across multiple zero-shot benchmarks in LLAMA-2-7B. Furthermore, compared to integer domain computing, a logarithmic processing element array based on the proposed SWA yields reductions of 24.6% in area and 42.3% in power, while achieving higher accuracy. Xinkuang Geng, Yunjie Lu, Hui Wang 0023, Honglan Jiang |
DATE | 4 |
| 2025 | A Low-Power Mixed-Precision Integrated Multiply-Accumulate Architecture for Quantized Deep Neural NetworksabstractAs mixed-precision quantization techniques have been widely considered for balancing computational efficiency and flexibility in quantized deep neural networks (DNNs), mixed-precision multiply-accumulate (MAC) units are increasingly important in DNN accelerators. However, conventional mixed-precision MAC architectures support either signed × signed or unsigned ×unsigned multiplications. The signed ×unsigned multiplication enhancing the computing efficiency of DNNs with ReLU activations has never been considered in the design of mixed-precision MAC. Thus, this work proposes a mixed-precision MAC architecture supporting six operation modes, int8 × int8, int8 × uint8, two int4 × int4, two int4 × uint4, four int2 × int2, and four int2 × uint2. In this design, to balance the power and delay of different modes, the multiplication is implemented based on four precision-split 4×4 multipliers (PS4Ms). The accumulation is integrated into the partial product accumulation of the multiplication to eliminate redundant switching activities in separate compression. With 10% area reduction, the proposed MAC denoted as PS4MAC, reduces the power by over 35%, 42%, and 56% for 8-bit, 4-bit, and 2-bit operations, respectively, compared with the design based on the Synopsys DesignWare (DW) multipliers. Additionally, it achieves over 23% power savings for 8-bit operations compared to state-of-the-art (SotA) mixed-precision MAC designs. To save more power, an approximate computing mode for 8-bit multiplication is further designed, resulting in a MAC unit enabling eight operation modes, referred to as PS4MAC_AP. Finally, output-stationary systolic arrays (SAs) are explored using the above-mentioned MAC designs to implement DNNs operating under a 1 GHz clock. Our designs show the highest energy efficiency and outstanding area efficiency in all 8-bit, 4-bit, and 2-bit operation modes. Compared with the traditional SA with high-precision-split multipliers, PS4MAC_AP improves the energy efficiency for 8-bit operations by 0.6 TOPS/W, and PS4MAC achieves 0.4 TOPS/W - 0.7 TOPS/W improvement for all operation modes. Xiaolu Hu, Xinkuang Geng, Zhigang Mao, Jie Han 0001, Honglan Jiang |
DATE | 5 |
| 2025 | Low-Power Multiplier Designs by Leveraging Correlations of 2$\times$×2 Encoded Partial ProductsabstractMultipliers, particularly those with small bit widths, are essential for modern neural network (NN) applications. In addition, multiple-precision multipliers are in high demand for efficient NN accelerators; therefore, recursive multipliers used in low-precision fusion schemes are gaining increasing attention. In this work, we design exact recursive multipliers based on customized approximate full adders (AFAs) for low-power purposes. Initially, the partial products (PPs) encoded by 2×2 multiplications are analyzed, which reveals the correlations among adjacent PPs. Based on these correlations, we propose 4×4 recursive multiplier architectures where certain full adders (FAs) can be simplified without affecting the correctness of the multiplication. Manually and synthesis tool-based FA simplifications are performed separately. The obtained 4×4 multipliers are then used to construct 8×8 multipliers based on a low-power recursive architecture. Finally, the proposed signed and unsigned 4×4 and 8×8 multipliers are evaluated using a 28nm CMOS technology. Compared with DesignWare (DW) multipliers, the proposed signed and unsigned 4×4 multipliers achieve power reductions of 16.5% and 11.6%, respectively, without compromising area or delay; alternatively, the delay can be reduced by 20.9% and 39.4%, respectively, without compromising power or area. For signed and unsigned 8×8 multipliers, the maximum power reductions are 9.7% and 13.7%, respectively, albeit with a trade-off in area. Siting Liu 0001, Hui Wang 0023, Qin Wang 0009, Fabrizio Lombardi, Zhigang Mao, Honglan Jiang |
IEEE Trans. Computers | 7 |
| 2025 | MACS: A Multidomain Collaborative Adaptive Clock Scheme for Large-Scale Reconfigurable Dataflow AcceleratorsabstractTo guarantee reliability and correctness, VLSI circuits are designed with conservative margins to maintain timing and power integrity against process, voltage, and temperature (PVT) variations across diverse workloads. However, worst-case PVT and workload conditions rarely occur in practice, resulting in significant timing slack and hence performance and energy loss, especially in reconfigurable dataflow accelerator RDA due to their large-scale and configurable features. Previous studies have attempted to exploit workload or PVT slack, yet achieving limited benefits for reconfigurable dataflow accelerator (RDAs) with large-scale processing element PE arrays. The key issues come from restricted scaling ranges for the clock, insufficient representations for the workload, and unbalanced workloads within processing elementss (PEs). To address these challenges, this article proposes the first multidomain collaborative adaptive clock scheme (MACS) to efficiently exploit both the workload and PVT timing slack for large-scale reconfigurable dataflow acceleratorss (RDAs). MACS partitions the RDA into several clock domains and allows constrained clock domain crossing, which enhances the hardware efficiency with minimal overhead and supports timing validation using conventional static timing analysis (STA) tools. In each domain, an operand-aware workload detection unit is developed, using both static configurations and dynamic operands to assess workload. The detected workload, combined with the monitored PVT conditions, determines the subsequent clock period. Additionally, to enable the exploration of timing slack over a broader range, the period range of the adaptive clock is extended. Experimental results show that MACS achieves a performance improvement of 76.3% or an energy saving of 36.6% with a hardware cost of 3.5%. Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Honglan Jiang, Zhigang Mao, Qin Wang 0009 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | A Reference Oversampling PLL With a FoMREF of -240.1 dB Enabled By a Capacitive Parasitic-Proof Ring Oscillator and a Time-Multiplexed Gm StageabstractThis paper presents a compact ring-oscillator (RO)-based phase-locked loop (PLL) implemented upon the principle that reference oversampling essentially boosts the reference frequency and thus extends the achievable bandwidth to such an extent that an area-efficient RO can be used without significantly sacrificing the phase noise (PN) and jitter performance when compared to conventional LC-based PLLs. By employing an analog reference oversampling PLL structure, RO noise is greatly suppressed by taking advantage of such an extended maximum PLL bandwidth. In conjunction with power- and spur-reduction techniques including a low-power time-multiplexed Gm stage and a capacitive parasitic-proof RO, this work implements a compact and low PN PLL without requiring complicated calibration or additional power/area penalties. Fabricated in a standard$0.18~\mu $m CMOS technology, the proposed PLL occupies an active area of 0.41 mm2. When operating at 1.6 GHz, the proposed PLL achieves an rms jitter of 585 fs with 5.7 mW power consumption, yielding a FoMREFof -240.1 dB. Xueke Cai, Tong Zhang 0030, Jianjun Zhou 0002, Howard Yang, Honglan Jiang, Yongfu Li 0002, Hui Wang 0023 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceabstractWhile exhibiting superior performance in many tasks, vision transformers (ViTs) face challenges in quantization. Some existing low-bit-width quantization techniques cannot effectively cover the whole inference process of ViTs, leading to an additional memory overhead (22.3%-172.6%) compared with corresponding fully quantized models. To address this issue, we propose quadruplet uniform quantization (QUQ) to deal with data of various distributions in ViT. QUQ divides the entire data range into at most four subranges that are uniformly quantized with different scale factors. To determine the partition scheme and quantization parameters, an efficient relaxation algorithm is proposed accordingly. Moreover, dedicated encoding and decoding strategies are devised to facilitate the design of an efficient accelerator. Experimental results show that QUQ surpasses state-of-the-art quantization techniques; it is the first viable scheme that can fully quantize ViTs to 6-bit with acceptable accuracy. Compared with conventional uniform quantization, QUQ leads to not only a higher accuracy but also an accelerator with lower area and power. Xinkuang Geng, Siting Liu 0001, Leibo Liu, Jie Han 0001, Honglan Jiang |
DAC | 5 |
| 2024 | Compact Powers-of-Two: An Efficient Non-Uniform Quantization for Deep Neural NetworksabstractTo reduce the demands for computation and memory of deep neural networks (DNNs), various quantization techniques have been extensively investigated. However, conventional methods cannot effectively capture the intrinsic data characteristics in DNNs, leading to a high accuracy degradation when employing low-bit-width quantization. In order to better align with the bell-shaped distribution, we propose an efficient non-uniform quantization scheme, denoted as compact powers-of-two (CPoT). Aiming to avoid the rigid resolution inherent in powers-of-two (PoT) without introducing new issues, we add a fractional part to its encoding, followed by a biasing operation to eliminate the unrepresentable region around O. This approach effectively balances the grid resolution in both the vicinity of 0 and the edge region. To facilitate the hardware implementation, we optimize the dot product for CPoT based on the computational characteristics of the quantized DNNs, where the precomputable terms are extracted and incorporated into bias. Consequently, a multiply-accumulate (MAC) unit is designed for CPoT using shifters and look-up tables (LUTs). The experimental results show that, even with a certain level of approximation, our proposed CPoT outperforms state-of-the-art methods in data-free quantization (DFQ), a post-training quantization (PTQ) technique focusing on data privacy and computational efficiency. Furthermore, CPoT demonstrates superior efficiency in area and power compared to other methods in hardware implementation. Xinkuang Geng, Siting Liu 0001, Jianfei Jiang 0001, Honglan Jiang |
DATE | 5 |
| 2024 | A Configurable Approximate Multiplier for CNNs Using Partial Product SpeculationabstractTo improve the performance and energy efficiency of the compute-intensive convolutional neural networks (CNNs), approximate multipliers have widely been investigated, taking advantage of the inherent error tolerance in CNNs. However, as per their divergencies in the capability of error tolerance, different CNN models and datasets may require various accuracy in multiplication. Thus, in this paper, we propose an energy-efficient approximate multiplier with configurable accuracy to satisfy the continuously evolving requirements of CNNs. In this design, the approximation level is configured by changing the processing scheme for inputs due to their significance to accuracy. The correlations between partial products (PPs) are utilized to eliminate the generation and accumulation of some less significant PPs that are speculated by their adjacent more significant ones. Consequently, four approximate multiplier configurations are devised for 8x8 unsigned multiplication, denoted as AMPPS_S2, AMPPS_S3, AMPPS_S4, and AMPPS_S6. Compared with existing approximate multipliers, the proposed designs show significantly higher accuracy. Compared with an exact carry-save array multiplier, AMPPS_S2 can reduce the power dissipation, delay, and area by 21.7%, 35.1 %, and 24.4%, respectively. Moreover, to enhance the efficiency of the proposed approximate multiplier in systolic array-based hardware architectures for CNNs, a novel encoding strategy is proposed for storing the pre-trained weights. Obtaining a similar classification accuracy to the accurate implementation (tested in ResNet18 and ResNet50 on ImageNet), the 32x32 systolic array using AMPPS_S2 shows a 13.5% reduction in power consumption and a 19.2% reduction in area. Overall, the experimental results demonstrate that the proposed approximate multiplier results in higher accuracy in CNN-based image classification, with lower hardware overhead, compared with state-of-the-art approximate multipliers. Xiaolu Hu, Xinkuang Geng, Zizhong Wei, Honglan Jiang |
DATE | 6 |
| 2024 | A Low-Power and High-Accuracy Approximate Adder for Logarithmic Number SystemabstractThe Logarithmic Number System (LNS) exploits the non-uniform distribution of data in convolutional neural networks (CNNs), so it leads to a high accuracy for image classification. An LNS provides an easier way to implement complex operations such as multiplication and division. However, addition and subtraction in the LNS require huge hardware resources due to the involved nonlinear operations. To mitigate this problem, we design a low-power approximate logarithmic adder with high-accuracy. Initially, a compact piecewise linear approximation (CPLA) algorithm is proposed to approximately compute the binary exponentiation and logarithm. Implemented by using simple circuits, the CPLA algorithm results in higher accuracy than the classical Mitchell’s algorithm. Consequently, three approximate logarithmic adders are devised, denoted as LA_CPLA1, LA_CPLA2, and LA_CPLA3. Compared with the logarithmic adder design based on lookup tables, the proposed LA_CPLA3 with a configuration of (e, f, n) = (7, 6, 3) achieves 35.05% and 39.80% reductions in area and power dissipation respectively, with a 0.01% mean relative error distance (MRED). We define (e, f) as the bit width of the logarithmic adder, where e and f are the bit widths of the integer and fractional parts, respectively. n is the approximate LSBs in the proposed LA_CPLAs processed by using OR gates. Compared with the multiply and accumulate (MAC) unit in a conventional system using fixed-point numbers, the MAC in the LNS using the proposed LA_CPLAs achieve a lower power by 5.96% to 32.02%, and a smaller area by 6.48% to 32.40%. To assess the efficiency of the proposed approximate adders, they are applied to the implementations of two image processing and CNN applications. The simulation results show that LA_CPLAs result in marginal accuracy loss compared to the corresponding accurate implementations. Xinkuang Geng, Qin Wang 0009, Jie Han 0001, Honglan Jiang |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Learning the Error Features of Approximate Multipliers for Neural Network ApplicationsabstractApproximate multipliers (AMs) have widely been investigated to pursue high-performance and energy-efficient hardware designs for error-tolerant applications, such as neural networks (NNs). The computing accuracy of an AM has been evaluated by using statistical error features; however, it is difficult to estimate the quality of a specific application using AMs. Thus, it is a great challenge to select or design appropriate AMs for an accuracy-constrained application. This paper proposes an application-oriented error evaluation framework for AMs with the aim of exploring the correlation between statistical error features of AMs and the accuracy degradation in AM-based NN applications. Specifically, based on the Dropout Feature Ranking technique, statistical error features of AMs are extensively studied and ranked by their importance to the accuracy of AM-based NN applications. The three most informative features are obtained to construct error models to predict the accuracy loss of AM-based NN applications. The constructed classification models show a probability higher than 96% for correctly classifying the AMs into three categories in accordance with the induced accuracy loss in AM-based NN applications. Furthermore, regression models can predict the accuracy of NN applications using an AM with a deviation as low as 6%. These results show that the proposed error evaluation framework can guide an efficient selection of AMs for NN applications by using just several AM error features, instead of running time-consuming and complicated hardware simulation. The obtained statistical error features can also provide a guidance for the design or generation of application-oriented AMs. Moreover, the proposed framework is applicable for quickly analyzing and selecting other approximate circuits for error-tolerant applications. Hai Mo, Yong Wu 0009, Honglan Jiang, Zining Ma, Fabrizio Lombardi, Jie Han 0001, Leibo Liu |
IEEE Trans. Computers | 3 |
| 2024 | Hardware-Efficient Logarithmic Floating-Point Multipliers for Error-Tolerant ApplicationsabstractThe increasing computational intensity of important new applications poses a challenge for their use in resource-restricted devices. Approximate computing using power-efficient arithmetic circuits is one of the emerging strategies to reach this objective. In this article, five hardware-efficient logarithmic floating-point (FP) multipliers are proposed, which all use simple operators, such as adders and multiplexers, to replace complex and more costly conventional FP multipliers. Radix-4 logarithms are used to further reduce the hardware complexity. These designs produce double-sided error distributions to mitigate error accumulation in complex computations. The proposed multipliers provide superior trade-offs between accuracy and hardware, with up to 30.8% higher accuracy than a recent logarithmic FP design or up to$68\times $less energy than the conventional FP multiplier. Using the proposed FP logarithmic multipliers in JPEG image compression achieves higher image quality than a recent logarithmic multiplier design with up to 4.7 dB larger peak signal-to-noise ratio. For training in benchmark NN applications, the proposed FP multipliers can slightly improve the classification accuracy while achieving$4.2\times $less energy and$2.2\times $smaller area than the state-of-the-art design. Zijing Niu, Honglan Jiang, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion MethodologiesabstractAttention mechanism acceleration is becoming increasingly vital to achieve superior performance in deep learning tasks. Existing accelerators are commonly devised dedicatedly by exploring the potential sparsity in neural network (NN) models, which suffer from complicated training, tuning processes, and accuracy degradation. By systematically analyzing the inherent dataflow characteristics of attention mechanism, we propose the Co-Operative Systolic Array (COSA) to pursue higher computational efficiency for its acceleration. In COSA, two systolic arrays that can be dynamically configured into weight or output stationary modes are cascaded to enable efficient attention operation. Thus, hybrid dataflows are simultaneously supported in COSA. Furthermore, various fusion methodologies and an advanced softmax unit are designed. Experimental results show that the COSA-based accelerator can achieve 2.95-28.82× speedup compared with the existing designs, with up to 97.4% PE utilization rate and less memory access. Zhican Wang, Gang Wang 0063, Honglan Jiang, Ningyi Xu, Guanghui He 0002 |
DAC | 3 |
| 2023 | Feature-Embedding Triplet Networks with a Separately Constrained Loss FunctionabstractFeature-embedding triplet networks (TNs) with three symmetric subchannels are very promising for similarity-measuring applications. This paper proposes a novel separately constrained triple loss (SCTL) function that applies to TNs for classification. Through minimizing the intra-class distance and maximizing the inter-class distance, SCTL eliminates possible false solutions and provides insight into the dependency of training based on these two terms. Based on this dependency, the strategy of selecting hyperparameters in SCTL is also analyzed to further improve performance. The effectiveness of the proposed SCTL is evaluated based on TNs with multi-layer perceptrons; the results show that compared to all existing loss functions, the use of SCTL offers the best classification accuracy for the TNs, while incurring in negligible hardware overhead (e.g., only a 0.0002% area overhead of the subnetworks). Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Honglan Jiang, Siting Liu 0001, Pedro Reviriego, Fabrizio Lombardi |
ISCAS | 4 |
| 2023 | Approximate Processing Element Design and Analysis for the Implementation of CNN Accelerators
Honglan Jiang, Hai Mo, Jie Han 0001, Leibo Liu, Zhigang Mao |
J. Comput. Sci. Technol. | 2 |
| 2022 | Upward Packet Popup for Deadlock Freedom in Modular Chiplet-Based SystemsabstractMonolithic SoCs can be decomposed into disparate chiplets that support integration with advanced pack-aging technologies. This concept is promising in reducing the manufacturing cost of large scale SoCs due to the higher yield rate and reusability of chiplets. The chiplets should be designed in a modular manner without holistic system knowledge so that they can be reused in different SoCs. However, the design modularity is a major challenge to the networks-on-chip (NoCs) of chiplets.New deadlocks may occur across both the chiplets and the interposer due to the integration, even if the NoC of each individually designed chiplet is deadlock free. However, conventional deadlock freedom approaches are unsuitable to handle such deadlocks because they require holistic knowledge and violate the modularity. Although there are several modular approaches that specifically target at integration-induced deadlocks, their routing is overly restricted and the injection control incurs additional latency. They also lack flexibility in dynamically changing topologies due to their complex software algorithm and the hard-wired components.In this paper, a key insight on the chiplet integration-induced deadlocks is gained, inspired by which a deadlock recovery framework (named UPP) is proposed. Specifically, it is verified that an integration-induced deadlock always involves a stalled upward packet moving from the interposer to the connected chiplet via the vertical link. Thus, UPP detects a deadlock by discovering the upward packet and recovers the system from deadlock by transmitting the upward packet to its destination. Hybrid flow control mechanisms are proposed to enable the upward packet to bypass the buffers and be transmitted via the normal router datapath. To guarantee the ejection of the upward packet after transmission, a lightweight protocol is proposed to reserve ejection queue entries of the network interface. Experimental results show that while adhering to design modularity, UPP provides an average runtime speedup of 3.1%∼10.3% with an area overhead of less than 4%. Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Jianfeng Zhu 0001, Honglan Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu |
HPCA | 6 |
| 2022 | Characterizing Approximate Adders and Multipliers for Mitigating Aging and Temperature DegradationsabstractThe performance of nanoscale semiconductor technologies has become susceptible to high temperatures and aging phenomena. While guard-bands have conventionally been used to combat degradation-induced timing violations, approximations have recently been leveraged to compensate for degradations in lieu of adding timing guard-bands, without a loss in performance. However, only simple approximation techniques such as truncation have been considered in prior work. In this paper, a wide range of approximate arithmetic circuits including adders and multipliers using various sophisticated approximation techniques are investigated to cope with aging- and temperature-induced degradations. To this end, approximate circuits are first characterized for their delay increase under degradations. With this, we then determine the approximation level required to compensate for guard-bands under different degradations. Degradation-aware logic synthesis results show that the simple use of truncated arithmetic circuits leads to a higher quality loss compared to using other approximate circuits. However, a truncated multiplier has the lowest error distance towards a reliable operation in 10 years. The approximate multipliers with configurable error recovery are most suitable when the level of degradation is higher, e.g., at a temperature of 70 °C. The characterization of degradation at the circuit level is then used for design exploration at the architecture level without the need for further gate-level simulations. For three different image processing applications, experimental results show that guard-bands can be mitigated while maintaining an output result with a high visual quality. Francisco J. H. Santiago, Honglan Jiang, Hussam Amrouch, Andreas Gerstlauer, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | An Energy-Efficient Approximate Divider Based on Logarithmic Conversion and Piecewise Constant ApproximationabstractApproximate computing (AC) has been considered as a promising paradigm to improve the energy-efficiency of computing hardware for error-tolerant applications, with negligible quality degradation to the output. Dividers frequently limit the performance of a computing system; however, they have not received as much attention as multipliers and adders in AC. In this paper, an energy-efficient and high-performance approximate divider is proposed based on logarithmic conversion and piecewise constant approximation. In this design, the range for the conversion between binary and logarithmic numbers is first expanded from$\mathbf {[{0,1}]}$to$\mathbf {[-0.5,1]}$. A heuristic search algorithm is then devised to find the most accurate constant set to approximate the reciprocal of the divisor, by minimizing a statistical error. The hardware implementation is presented for both floating-point (FP) and integer dividers. With a high configurability, the proposed divider results in a mean relative error distance (MRED) from 2.78% to 0.046%, indicating a high accuracy among state-of-the-art approximate dividers. Compared to the half-precision FP divider, the proposed divider with a MRED of 0.74% can achieve nearly$\mathbf {90\times }$improvement in PDP. Moreover, compared to state-of-the-art approximate dividers, the proposed design is in the Pareto Frontier in terms of power delay product (PDP) and MRED. The three image processing application results demonstrate that the proposed divider can result in the highest peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) even with truncation. Yong Wu 0009, Honglan Jiang, Zining Ma, Pengfei Gou, Jie Han 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | A Logarithmic Floating-Point Multiplier for the Efficient Training of Neural NetworksabstractThe development of important applications of increasingly large neural networks (NNs) is spurring research that aims to increase the power efficiency of the arithmetic circuits that perform the huge amount of computation in NNs. The floating-point (FP) representation with a large dynamic range is usually used for training. In this paper, it is shown that the FP representation is naturally suited for the binary logarithm of numbers. Thus, it favors a design based on logarithmic arithmetic. Specifically, we propose an efficient hardware implementation of logarithmic FP multiplication that uses simpler operations to replace complex multipliers for the training of NNs. This design produces a double-sided error distribution that mitigates the accumulative effect of errors in iterative operations, so it is up to 45% more accurate than a recent logarithmic FP design. The proposed multiplier also consumes up to 23.5x less energy and 10.7x smaller area compared to exact FP multipliers. Benchmark NN applications, including a 922-neuron model for the MNIST dataset, show that the classification accuracy can be slightly improved using the proposed multiplier, while achieving up to 2.4x less energy and 2.8x smaller area with a better performance. Zijing Niu, Honglan Jiang, Mohammad Saeed Ansari, Bruce F. Cockburn, Leibo Liu, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | Non-Volatile Approximate Arithmetic Circuits Using Scalable Hybrid Spin-CMOS Majority GatesabstractIn the nanoscale era, leakage/static power dissipation has become an inevitable and important issue for CMOS devices. To alleviate this issue, we propose to use spintronic devices with near-zero leakage power and non-volatility as key components in arithmetic circuits for error-resilient applications. To this end, spintronic threshold devices are first utilized to construct highly-scalable majority gates (MGs) based on spin-CMOS technology. These MGs are then used in the design of compressors for constructing multipliers and accumulators. For an MG-based compressor, the truth table of a conventional compressor is transformed to ensure that the outputs depend only on the number of input “1”s. To synthesize and optimize the MG-based circuits, a heuristic majority-inverter graph (HMIG) is further proposed for the design of an accurate and two approximate non-volatile 4-2 compressors (denoted as MG-EC, MG-AC1 and MG-AC2). Due to the high scalability of the MGs, approximate compressors with a larger number of inputs can be devised using the same method. Compared to previous designs, the proposed 4-2 compressors show shorter critical path delays and lower energy consumption; MG-AC1 and MG-AC2 also achieve a higher accuracy than state-of-the-art approximate designs. For achieving a similar image quality in image compression, the multiplier implementations using MG-AC1 and MG-AC2 result in more significant reductions in delay and energy than those using other approximate designs. Honglan Jiang, Shaahin Angizi, Deliang Fan, Jie Han 0001, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Power-Efficient Approximate Multiplier Using Adaptive Error CompensationabstractIn this paper, a design framework is proposed for a power-efficient approximate multiplier using adaptive error compensation and the optimal error compensation values are determined by using probability theory, resulting in a minimal mean-squared-error (MSE). To further pursue the viability of adaptive accuracy, adaptive error compensation scheme using different levels of quantization is utilized for the predetermined compensation values. The simulation results show that the approximate multipliers based on the proposed design framework outperform state-of-the-art designs in both accuracy and circuit measurements. Specifically, with a higher accuracy, the proposed designs save up to 36.84% and 21.61% in power consumption compared to state-of-the-art unsigned and signed approximate $16\times16$ multipliers, respectively. In terms of power-delay-product (PDP), the improvements are up to 42.3% and 21.93% for the unsigned and signed multiplier designs, respectively. Finally, the approximate multipliers are further assessed in the implementation of an FIR filter. It shows that the proposed approximate multiplier achieves a similar filtering quality to the accurate design, with more than 50% reduction in power dissipation. Zhixi Yang, Honglan Jiang, Jun Yang 0006 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | Interplay Bitwise Operation in Emerging MRAM for Efficient In-memory Computing
Hao Cai 0001, Honglan Jiang, Yongliang Zhou, Menglin Han, Bo Liu 0019 |
CCF Trans. High Perform. Comput. | 2 |
| 2020 | Approximate Arithmetic Circuits: A Survey, Characterization, and Recent ApplicationsabstractApproximate computing has emerged as a new paradigm for high-performance and energy-efficient design of circuits and systems. For the many approximate arithmetic circuits proposed, it has become critical to understand a design or approximation technique for a specific application to improve performance and energy efficiency with a minimal loss in accuracy. This article aims to provide a comprehensive survey and a comparative evaluation of recently developed approximate arithmetic circuits under different design constraints. Specifically, approximate adders, multipliers, and dividers are synthesized and characterized under optimizations for performance and area. The error and circuit characteristics are then generalized for different classes of designs. The applications of these circuits in image processing and deep neural networks indicate that the circuits with lower error rates or error biases perform better in simple computations, such as the sum of products, whereas more complex accumulative computations that involve multiple matrix multiplications and convolutions are vulnerable to single-sided errors that lead to a large error bias in the computed result. Such complex computations are more sensitive to errors in addition than those in multiplication, so a larger approximation can be tolerated in multipliers than in adders. The use of approximate arithmetic circuits can improve the quality of image processing and deep learning in addition to the benefits in performance and power consumption for these applications. Honglan Jiang, Francisco J. H. Santiago, Hai Mo, Leibo Liu, Jie Han 0001 |
Proc. IEEE | 1 |
| 2019 | Characterizing Approximate Adders and Multipliers Optimized under Different Design ConstraintsabstractTaking advantage of the error resilience in many applications as well as the perceptual limitations of humans, numerous approximate arithmetic circuits have been proposed that trade off accuracy for higher speed or lower power in emerging applications that exploit approximate computing. However, characterizing the various approximate designs for a specific application under certain performance constraints becomes a new challenge. In this paper, approximate adders and multipliers are evaluated and compared for a better understanding of their characteristics when the implementations are optimized for performance or power. Although simple truncation can effectively reduce the hardware of an arithmetic circuit, it is shown that some other designs perform better in speed, power and power-delay product. For instance, many approximate adders have a higher performance than a truncated adder. A truncated multiplier is faster but consumes a higher power than most approximate designs for achieving a similar mean error magnitude. The logarithmic multipliers are very fast and power-efficient at a lower accuracy. Approximate multipliers can also be generated by an automated process to be very efficient while ensuring a sufficiently high accuracy. Honglan Jiang, Francisco J. H. Santiago, Mohammad Saeed Ansari, Leibo Liu, Bruce F. Cockburn, Fabrizio Lombardi, Jie Han 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Low-Power Unsigned Divider and Square Root Circuit Designs Using Adaptive ApproximationabstractIn this paper, an adaptive approximation approach is proposed for the design of a divider and a square root (SQR) circuit. In this design, the division/SQR is computed by using a reduced-width divider/SQR circuit and a shifter by adaptively pruning some insignificant input bits. Specifically, for a $2n/n$ 2 n / n division, $2k$ 2 k and $k$ k ($k< n$ k < n ) consecutive bits are selected starting from the most significant ‘1’ in the dividend and divisor, respectively. At the same time, redundant least significant bits (LSBs) are truncated or if the number of remaining bits after pruning is smaller than the number of bits to be kept, ‘0's are appended to the LSBs of the inputs. To avoid overflow, a $2(k+1)/(k+1)$ 2 ( k + 1 ) / ( k + 1 ) divider is used to compute the $2k/k$ 2 k / k division. Finally, an error correction circuit is proposed to recover the error caused by the shifter using OR gates. For a $2n$ 2 n -bit approximate SQR circuit, similar pruning schemes are used to obtain a $2k$ 2 k -bit radicand. A $2k$ 2 k -bit SQR circuit and a shifter are then utilized to compute the SQR. This adaptive operation leads to very small maximum error distances of the approximate divider and SQR circuits, as shown by a theoretical error analysis. The proposed 16/8 approximate divider using an 8/4 exact array divider is $2.5\times$ 2 . 5 × as fast but only consumes 34.42 percent of the power of the accurate design. Compared to the accurate 16-bit array SQR circuit, the approximate design with a 6-bit radicand is $3.9\times$ 3 . 9 × as fast and consumes 20.66 percent of the power. The approximate SQR circuit using a 6-bit lookup table-based SQR circuit consumes 7.15 percent of the power of its corresponding accurate design. The proposed designs outperform other approximate designs in image processing applications including change detection (for the divider), envelope detection (for the SQR circuit) and image reconstruction (for both designs). Honglan Jiang, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Computers | 1 |
| 2018 | Adaptive approximation in arithmetic circuits: A low-power unsigned divider designabstractMany approximate arithmetic circuits have been proposed for high-performance and low-power applications. However, most designs are either hardware-efficient with a low accuracy or very accurate with a limited hardware saving, mostly due to the use of a static approximation. In this paper, an adaptive approximation approach is proposed for the design of a divider. In this design, division is computed by using a reduced-width divider and a shifter by adaptively pruning the input bits. Specifically, for a 2n/n division 2k/k bits are selected starting from the most significant `1' in the dividend/divisor. At the same time, redundant least significant bits (LSBs) are truncated or if the number of remaining LSBs is smaller than 2k for the dividend or k for the divisor, `0's are appended to the LSBs of the input. To avoid overflow, a 2(k + 1)/(k + 1) divider is used to compute the division of the 2k-bit dividend and the k-bit divisor, both with the most significant bits being `0'. Thus, k <; n is a key variable that determines the size of the divider and the accuracy of the approximate design. Finally, an error correction circuit is proposed to recover the error caused by the shifter by using OR gates. The synthesis results in an industrial 28nm CMOS process show that the proposed 16/8 approximate divider using an 8/4 accurate divider is 2.5χ as fast and consumes 34.42% of the power of the accurate 16/8 design. Compared with the other approximate dividers, the proposed design is significantly more accurate at a similar power-delay product. Moreover, simulation results show that the proposed approximate divider outperforms the other designs in two image processing applications. Honglan Jiang, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
DATE | 1 |
| 2018 | Gradient Descent Using Stochastic Circuits for Efficient Training of Learning MachinesabstractGradient descent (GD) is a widely used optimization algorithm in machine learning. In this paper, a novel stochastic computing GD circuit (SC-GDC) is proposed by encoding the gradient information in stochastic sequences. Inspired by the structure of a neuron, a stochastic integrator is used to optimize the weights in a learning machine by its “inhibitory” and “excitatory” inputs. Specifically, two AND (or XNOR) gates for the unipolar representation (or the bipolar representation) and one stochastic integrator are, respectively, used to implement the multiplications and accumulations in a GD algorithm. Thus, the SC-GDC is very area- and power-efficient. As per the formulation of the proposed SC-GDC, it provides unbiased estimate of the optimized weights in a learning algorithm. The proposed SC-GDC is then used to implement a least-mean-square algorithm and a softmax regression. With a similar accuracy, the proposed design achieves more than $30 \times $ improvement in throughput per area (TPA) and consumes less than 13% of the energy per training sample, compared with a fixed-point implementation. Moreover, a signed SC-GDC is proposed for training complex neural networks (NNs). It is shown that for a 784-128-128-10 fully connected NN, the signed SC-GDC produces a similar training result with its fixed-point counterpart, while achieving more than 90% energy saving and 82% reduction in training time with more than $50 \times $ improvement in TPA. Siting Liu 0001, Honglan Jiang, Leibo Liu, Jie Han 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Scalable Construction of Approximate Multipliers With Formally Guaranteed Worst Case Error
Vojtech Mrazek, Zdenek Vasícek, Lukás Sekanina, Honglan Jiang, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | A Review, Classification, and Comparative Evaluation of Approximate Arithmetic CircuitsabstractOften as the most important arithmetic modules in a processor, adders, multipliers, and dividers determine the performance and energy efficiency of many computing tasks. The demand of higher speed and power efficiency, as well as the feature of error resilience in many applications (e.g., multimedia, recognition, and data analytics), have driven the development of approximate arithmetic design. In this article, a review and classification are presented for the current designs of approximate arithmetic circuits including adders, multipliers, and dividers. A comprehensive and comparative evaluation of their error and circuit characteristics is performed for understanding the features of various designs. By using approximate multipliers and adders, the circuit for an image processing application consumes as little as 47% of the power and 36% of the power-delay product of an accurate design while achieving similar image processing quality. Improvements in delay, power, and area are obtained for the detection of differences in images by using approximate dividers. Honglan Jiang, Cong Liu 0015, Leibo Liu, Fabrizio Lombardi, Jie Han 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2017 | Design of Approximate Radix-4 Booth Multipliers for Error-Tolerant ComputingabstractApproximate computing is an attractive design methodology to achieve low power, high performance (low delay) and reduced circuit complexity by relaxing the requirement of accuracy. In this paper, approximate Booth multipliers are designed based on approximate radix-4 modified Booth encoding (MBE) algorithms and a regular partial product array that employs an approximate Wallace tree. Two approximate Booth encoders are proposed and analyzed for error-tolerant computing. The error characteristics are analyzed with respect to the so-called approximation factor that is related to the inexact bit width of the Booth multipliers. Simulation results at 45 nm feature size in CMOS for delay, area and power consumption are also provided. The results show that the proposed 16-bit approximate radix-4 Booth multipliers with approximate factors of 12 and 14 are more accurate than existing approximate Booth multipliers with moderate power consumption. The proposed R4ABM2 multiplier with an approximation factor of 14 is the most efficient design when considering both power-delay product and the error metric NMED. Case studies for image processing show the validity of the proposed approximate radix-4 Booth multipliers. Weiqiang Liu 0001, Liangyu Qian, Chenghua Wang, Honglan Jiang, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 4 |
| 2016 | Approximate Radix-8 Booth Multipliers for Low-Power and High-Performance OperationabstractThe Booth multiplier has been widely used for high performance signed multiplication by encoding and thereby reducing the number of partial products. A multiplier using the radix-$4$(or modified Booth) algorithm is very efficient due to the ease of partial product generation, whereas the radix-$8$Booth multiplier is slow due to the complexity of generating the odd multiples of the multiplicand. In this paper, this issue is alleviated by the application of approximate designs. An approximate$2$-bit adder is deliberately designed for calculating the sum of$1\times$and$2\times$of a binary number. This adder requires a small area, a low power and a short critical path delay. Subsequently, the$2$-bit adder is employed to implement the less significant section of a recoding adder for generating the triple multiplicand with no carry propagation. In the pursuit of a trade-off between accuracy and power consumption, two signed$16\times 16$bit approximate radix-8 Booth multipliers are designed using the approximate recoding adder with and without the truncation of a number of less significant bits in the partial products. The proposed approximate multipliers are faster and more power efficient than the accurate Booth multiplier. The multiplier with 15-bit truncation achieves the best overall performance in terms of hardware and accuracy when compared to other approximate Booth multiplier designs. Finally, the approximate multipliers are applied to the design of a low-pass FIR filter and they show better performance than other approximate Booth multipliers. Honglan Jiang, Jie Han 0001, Fei Qiao, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2015 | A Comparative Review and Evaluation of Approximate AddersabstractAs an important arithmetic module, the adder plays a key role in determining the speed and power consumption of a digital signal processing (DSP) system. The demands of high speed and power efficiency as well as the fault tolerance nature of some applications have promoted the development of approximate adders. This paper reviews current approximate adder designs and provides a comparative evaluation in terms of both error and circuit characteristics. Simulation results show that the equal segmentation adder (ESA) is the most hardware-efficient design, but it has the lowest accuracy in terms of error rate (ER) and mean relative error distance (MRED). The error-tolerant adder type II (ETAII), the speculative carry select adder (SCSA) and the accuracy-configurable approximate adder (ACAA) are equally accurate (provided that the same parameters are used), however ETATII incurs the lowest power-delay-product (PDP) among them. The almost correct adder (ACA) is the most power consuming scheme with a moderate accuracy. The lower-part-OR adder (LOA) is the slowest, but it is highly efficient in power dissipation. Honglan Jiang, Jie Han 0001, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 1 |