EDBT 2026 Demo / reviewers in the wild / expert
Siting Liu 0001
dblp:199/8619-1
· DBLP profile ↗
24ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-0505-8183ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 5 first-author · 17 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-ANT: Efficient Low-Bit Group-Wise Quantization for Large Language Models via Sign-Asymmetric Adaptive Numeric TypeabstractLarge language models (LLMs) have demonstrated remarkable potential across diverse domains; meanwhile, their large parameter sizes pose substantial inference costs, motivating the need for efficient low-bit quantization. Group-wise quantization, which adopts finer granularity, has been widely used to improve low-bit quantization performance. Several adaptive numeric types have been proposed to further enhance low-bit group-wise quantization; however, they construct quantization grids based on symmetric numeric types, which limits their ability to model asymmetric distributions. To address this limitation, we propose SA-ANT, a sign-asymmetric adaptive numeric type for efficient low-bit group-wise quantization. SA-ANT constructs quantization grids separately on the positive and negative sides, enabling adaptive support for asymmetric and non-uniform distributions. Furthermore, the carefully designed SA-ANT not only reduces quantization errors but also ensures a unified computing across different sub numeric types, thereby facilitating hardware efficiency. To accelerate LLM inference, we develop (1) a quantization framework that transforms LLM weights into the SA-ANT and adaptively selects the sub numeric type for each group, and (2) an accelerator that maps SA-ANT inference to low-bit INT operations. Experimental results show that SA-ANT delivers 3.92%– 5.57% higher accuracy than state-of-the-art adaptive numeric types under 3-bit weight quantization, while also enabling 7.84%– 44.65% area savings and 7.80%–43.88% power reductions. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 2 |
| 2026 | Bridging the Power Estimation Gap: A GNN-Based Prediction Model for Approximate Logic SynthesisabstractApproximate computing is an application-related paradigm that trades limited accuracy for improvements in hardware cost. As a key technique of approximate computing, approximate logic synthesis (ALS) automatically generates approximate circuits with reduced area, power, and delay while satisfying predefined quality-of-result (QoR) constraints. However, in typical gate-level ALS workflows, synthesis tools are invoked at the final stage for optimization, leading to a discrepancy between the circuit in design space exploration (DSE) and the final obtained circuit. Thus, the power estimation for a candidate circuit during DSE may exhibit a significant gap from the actual power consumed by its post-synthesis circuit. This gap may mislead the DSE to a sub-optimal design. To address this issue, we propose a graph neural network (GNN)-based power prediction model that operates on gate-level circuits. The model incorporates multi-head channel attention, which extracts high-level topological and functional features that correlate with power dissipation and implicitly captures the optimization behavior of synthesis tools. Thus, it enables a direct prediction of post-synthesis power from pre-synthesis gate-level circuits. Experimental results show that the proposed model improves the concordance index (C-index) for power ranking by up to 14.0% over traditional methods. Furthermore, we construct an ALS framework by integrating the proposed model with Cartesian genetic programming (CGP). Compared to state-of-the-art ALS approaches, our GNN-CGP framework generates circuits with up to 26.8% power savings under the same error constraints. Fuxuan Li, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 3 |
| 2026 | LUT-ALMs: Trading Off Accuracy and Power for Approximate Logarithmic Multipliers via LUT OptimizationabstractLogarithmic multiplier (LM) converts fixed-point (FxP) input operands to logarithmic numbers and performs multiplication with simple shift and addition operations, which achieves distinct power reduction, yet with significant single-sided errors. This paper proposes to fuse error compensation with logarithmic conversion by using customized look-up tables (LUTs). To avoid the use of large LUTs, partition strategies are designed for the optimization of LUTs. In addition, to effectively balance the accuracy and hardware costs, two iterative algorithms are proposed for generating precision-configurable LUTs. Based on the optimized LUTs, high-accuracy and low-power approximate LMs (LUT-ALMs) are constructed for 8-bit and 16-bit multiplications. Furthermore, to enhance the flexibility of data type and bit width, a mixed-mode LM (MM-ALM) supporting eight multiplication modes is devised. Compared with the exact 8-bit signed multiplier from Synopsys DesignWare (DW) library (DW-Exact), LUT-ALMs show up to 30.96% and 23.99% reductions in the power-delay product (PDP) and area-delay product (ADP), respectively, with a mean relative error distance (MRED) of 3.45%. Compared with state-of-the-art approximate 8- bit signed multipliers, LUT-ALMs form the Pareto front in terms of PDP and MRED. For 16-bit multiplication, LUT-ALMs obtain up to 70.90% and 62.13% savings in PDP and ADP with a MRED of 2.75%, compared with the corresponding DW-Exact. Compared with the corresponding mixed-mode exact multiplier constructed of DW multipliers, MM-ALM performing 8-bit multiplication achieves up to 57.44% and 47.99% reductions in PDP and ADP, respectively. When performing 4-bit multiplications, MM-ALM can save up to 37.16% savings in PDP. With lower hardware over-heads, LUT-ALMs and MM-ALM present comparable accuracy to the corresponding exact designs in the considered convolutional neural networks (CNNs) and image processing applications. The hardware description of the devised LMs is open-sourced athttps://anonymous.4open.science/r/LUT-ALM-5FD1. Xinkuang Geng, Xiaolu Hu, Hui Wang 0023, Jianfei Jiang 0001, Qin Wang 0009, Siting Liu 0001, Jie Han 0001, Honglan Jiang |
IEEE Trans. Computers | 7 |
| 2025 | EPIC: Error PredIction and Correction for Power-Efficient Voltage Underscaling Multiply-Accumulate UnitabstractMatrix multiplication dominates the power consumption in compute-intensive applications such as deep neural networks (DNNs), spurring intensive investigations into power-efficient multiply-accumulate (MAC) units. Among the mainstream low-power design methodologies, voltage underscaling can achieve effective power savings yet induce timing errors that may lead to catastrophic accuracy loss. In this paper, we propose an error prediction and correction framework (denoted as EPIC) for arbitrary MAC unit under voltage underscaling, which predicts the timing errors and samples the correct output by using a delay-tunable clock. A prediction bits searching algorithm is proposed to enhance the prediction accuracy with low hardware cost, resulting in up to 100% accuracy. While preserving the accuracy, EPIC achieves up to 52% power savings over the corresponding MAC operating at nominal voltage. With transistor-level optimizations, EPIC incurs only 8% area and 1% power overheads, achieving 100% error correction under a voltage underscaling ratio of $\mathbf{0. 7 4}$. Compared to state-of-the-art error resilient circuit designs, EPIC consumes 60%-88% less area. Additionally, to achieve the accuracy performance of EPIC in error-resilient applications, we propose a simulation workflow involving precise timing features, enabling an accurate simulation of voltage underscaling MAC in large-scale applications. The experimental results show that, under voltage-underscaling, the MAC with EPIC consumes 11% less power than the one without EPIC, when a same accuracy as exact implementation is required in multi-layer perceptron (MLP). Tongjing Wu, Xiaolu Hu, Siting Liu 0001, Hui Wang 0023, Weifeng He, Zhigang Mao, Honglan Jiang |
DAC | 4 |
| 2025 | Lookup Table Refactoring: Towards Efficient Logarithmic Number System Addition for Large Language ModelsabstractCompared to integer quantization, logarithmic quantization aligns more effectively with the long-tailed distribution of data in large language models (LLMs), resulting in lower quantization errors. Moreover, the logarithmic number system (LNS) employs a fixed-point adder to perform multiplication, indicating a potential reduction in computational complexity for LLM accelerators that require extensive multiply-accumulate (MAC) operations. However, a key bottleneck is that LNS addition requires complex nonlinear functions, which are typically approximated using lookup tables (LUTs). This study aims to reduce the hardware resources needed for LUTs in LNS addition while maintaining high precision. Specifically, we investigate the specific nature of addition operations within LLMs; the relationship between the hardware parameters of the LUT and the computing errors is then mathematically derived. Based on these insights, we propose LUT refactoring to optimize the LUT for enhanced efficiency in LNS addition. With 10.93% and 19.78% reductions in area-delay product (ADP) and power-delay product (PDP), respectively, LUT refactoring results in an accuracy improvement of up to 33.5% in LLM benchmarks compared to the naive design. When compared to integer quantization, our method achieves higher accuracy while reducing area by 18.27% and power by 42.61%. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 2 |
| 2025 | A Hardware Prototype of an MRAM-based Stochastic Computing System
Zhengkun Yu, Tianhan Fei, Hao Cai 0001, Heng Shi 0003, Siting Liu 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | Layer Fusion-Accelerated Online Scheduling for Multi-Tenancy on Heterogeneous DNN AcceleratorsabstractToday, hardware accelerators are being deployed in cloud and edge computing to serve DNN inference jobs that multiple tenants keep issuing. The use of heterogeneous multi-core accelerator systems has been considered. The intricate nature of one such system and the dynamicity of multi-tenant jobs over time yet make the scheduling a complex problem. In this paper, we study layer fusion techniques and propose a new scheduling algorithm named Lucas. Lucas aims to maximally meet the Quality of Service (QoS) requirements for all tenants when mapping their inference jobs onto heterogeneous accelerators. After breaking DNN layers into fine-grained units, Lucas online decides whether to perform multi-core layer fusion, single-core layer fusion, or layer-by-layer execution regarding factors such as the memory bandwidth consumption, the layers awaiting execution, and the cost of layer fusion. Evaluation shows that, compared to state-of-the-art schedulers, Lucas achieves significantly higher Service Level Agreement (SLA) compliance across various workloads. Zhaojun Ni, Yutong Wang 0011, Siting Liu 0001, Chundong Wang 0001 |
ICPADS | 4 |
| 2025 | An SRAM-based Stochastic Number Generator for Stochastic ComputingabstractStochastic computing (SC) features a unique number representation, where real values are encoded by the probability of "1"s in a random binary bit stream or a stochastic sequence. It enables hardware-efficient arithmetic circuit designs with simple logic gates. However, stochastic number generators (SNGs) are required to produce stochastic sequences. The high hardware cost of an SNG offsets the advantage of SC. To reduce the hardware cost of an SNG, we propose an SRAM-based SNG using voltage under-scaling. It generates random bits by leveraging the access instability of selected SRAM cells, induced by a reduced supply voltage. It is suitable for energy-efficient SC. We implemented the SRAM-based SNG on a Xilinx ZC702 FPGA using block RAMs and evaluated its performance across multiple SC applications, including finite-state machine-based tanh function generation and an Ising machine that solves max-cut problems (MCPs). For the tanh function, our design achieves a comparable mean-squared error (MSE) (8.7 × 10−3) compared to the use of traditional SNGs, such as Sobol- (4.67 × 10−2) and linear feedback shift register (LFSR)-based (1.53 × 10−2) ones. For MCPs, a maximum cut value comparable to that of a cutting-edge design is achieved. Compared with LFSR- and Sobol-based designs, the proposed design consumes 84.3% and 92.5% less energy, respectively. Heng Shi 0003, Zhengkun Yu, Jie Han 0001, Siting Liu 0001 |
ISCAS | 6 |
| 2025 | Low-Power Multiplier Designs by Leveraging Correlations of 2$\times$×2 Encoded Partial ProductsabstractMultipliers, particularly those with small bit widths, are essential for modern neural network (NN) applications. In addition, multiple-precision multipliers are in high demand for efficient NN accelerators; therefore, recursive multipliers used in low-precision fusion schemes are gaining increasing attention. In this work, we design exact recursive multipliers based on customized approximate full adders (AFAs) for low-power purposes. Initially, the partial products (PPs) encoded by 2×2 multiplications are analyzed, which reveals the correlations among adjacent PPs. Based on these correlations, we propose 4×4 recursive multiplier architectures where certain full adders (FAs) can be simplified without affecting the correctness of the multiplication. Manually and synthesis tool-based FA simplifications are performed separately. The obtained 4×4 multipliers are then used to construct 8×8 multipliers based on a low-power recursive architecture. Finally, the proposed signed and unsigned 4×4 and 8×8 multipliers are evaluated using a 28nm CMOS technology. Compared with DesignWare (DW) multipliers, the proposed signed and unsigned 4×4 multipliers achieve power reductions of 16.5% and 11.6%, respectively, without compromising area or delay; alternatively, the delay can be reduced by 20.9% and 39.4%, respectively, without compromising power or area. For signed and unsigned 8×8 multipliers, the maximum power reductions are 9.7% and 13.7%, respectively, albeit with a trade-off in area. Siting Liu 0001, Hui Wang 0023, Qin Wang 0009, Fabrizio Lombardi, Zhigang Mao, Honglan Jiang |
IEEE Trans. Computers | 2 |
| 2024 | QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceabstractWhile exhibiting superior performance in many tasks, vision transformers (ViTs) face challenges in quantization. Some existing low-bit-width quantization techniques cannot effectively cover the whole inference process of ViTs, leading to an additional memory overhead (22.3%-172.6%) compared with corresponding fully quantized models. To address this issue, we propose quadruplet uniform quantization (QUQ) to deal with data of various distributions in ViT. QUQ divides the entire data range into at most four subranges that are uniformly quantized with different scale factors. To determine the partition scheme and quantization parameters, an efficient relaxation algorithm is proposed accordingly. Moreover, dedicated encoding and decoding strategies are devised to facilitate the design of an efficient accelerator. Experimental results show that QUQ surpasses state-of-the-art quantization techniques; it is the first viable scheme that can fully quantize ViTs to 6-bit with acceptable accuracy. Compared with conventional uniform quantization, QUQ leads to not only a higher accuracy but also an accelerator with lower area and power. Xinkuang Geng, Siting Liu 0001, Leibo Liu, Jie Han 0001, Honglan Jiang |
DAC | 2 |
| 2024 | A High-Performance Stochastic Simulated Bifurcation Ising MachineabstractIsing model-based computers, or Ising machines, have recently emerged as high-performance solvers for combinatorial optimization problems (COPs). A simulated bifurcation (SB) Ising machine searches for the solution by solving pairs of differential equations related to the oscillator positions and momenta. It benefits from massive parallelism but suffers from high energy. As an unconventional computing paradigm, dynamic stochastic computing implements accumulation-based operations efficiently. By exploiting the advantages in algorithm and hardware codesign, this article proposes a high-performance stochastic SB machine (SSBM) with efficient hardware. To this end, we develop a stochastic SB (sSB) algorithm such that the multiply-and-accumulate (MAC) operation is converted to multiplexing and addition while the numerical integration is implemented by using signed stochastic integrators (SSIs). Specifically, the sSB stochastically ternarizes position values used for the MAC operation. Two types of SB cells are constructed. A stochastic computing SB cell contains two SSIs with a high area efficiency, while a binary-stochastic computing SB cell contains one binary integrator and one SSI with a reduced delay. Based on sSB, an SSBM is then built by using the proposed SB cells as the basic building block. The designs and syntheses of two SSBMs with 2000 fully connected spins require at least 10.62% smaller area than the state-of-the-art designs. It shows the potential of stochastic computing for SB to efficiently solve COPs. Hongqiao Zhang, Zhengkun Yu, Siting Liu 0001, Jie Han 0001 |
DAC | 4 |
| 2024 | Compact Powers-of-Two: An Efficient Non-Uniform Quantization for Deep Neural NetworksabstractTo reduce the demands for computation and memory of deep neural networks (DNNs), various quantization techniques have been extensively investigated. However, conventional methods cannot effectively capture the intrinsic data characteristics in DNNs, leading to a high accuracy degradation when employing low-bit-width quantization. In order to better align with the bell-shaped distribution, we propose an efficient non-uniform quantization scheme, denoted as compact powers-of-two (CPoT). Aiming to avoid the rigid resolution inherent in powers-of-two (PoT) without introducing new issues, we add a fractional part to its encoding, followed by a biasing operation to eliminate the unrepresentable region around O. This approach effectively balances the grid resolution in both the vicinity of 0 and the edge region. To facilitate the hardware implementation, we optimize the dot product for CPoT based on the computational characteristics of the quantized DNNs, where the precomputable terms are extracted and incorporated into bias. Consequently, a multiply-accumulate (MAC) unit is designed for CPoT using shifters and look-up tables (LUTs). The experimental results show that, even with a certain level of approximation, our proposed CPoT outperforms state-of-the-art methods in data-free quantization (DFQ), a post-training quantization (PTQ) technique focusing on data privacy and computational efficiency. Furthermore, CPoT demonstrates superior efficiency in area and power compared to other methods in hardware implementation. Xinkuang Geng, Siting Liu 0001, Jianfei Jiang 0001, Honglan Jiang |
DATE | 2 |
| 2024 | Can Stochastic Computing Truly Tolerate Bit Flips?abstractTransient bit flips may cause unexpected behaviors in digital arithmetic circuits, and thus computation errors. It is generally believed that stochastic computing (SC) can tolerate bit flips due to its unique number representation format. In SC, a value is encoded by a random bit stream using the probability of 1s in the bit stream. Therefore, one bit flip changes the probability by a small number given a long enough bit stream, and the computation results might not change much. This is verified in previous works mostly through empirical studies. However, we found that bit flips actually change the probability of a stochastic bit stream by a predictable value. This is formulated by considering the bit flips as Bernoulli events. The distorted probability of the bit stream considering bit flips is derived and the SC bit flip models for various stochastic circuits are proposed. The error recovery based on the formulation is also proposed. The SC bit flip model and error recovery scheme are verified on basic SC elements such as the SC multipliers and an FSM-based stochastic circuit. They are then applied to a complex SC system computing a neural network. The results show that the accuracy of MNIST digit recognition can be recovered even with a high bit flip rate of 10%. Yutong Wang 0011, Zhaojun Ni, Siting Liu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | LDL-SCA: Linearized Deep Learning Side-Channel Attack Targeting Multi-tenant FPGAs✱abstractIn recent years, deep-learning side-channel attacks (DL-SCA) have gained increasing attention due to their enhanced efficacy against cryptographic modules. This paper explores that traditional non-profiled DL-SCA is unable to discern correct cipher keys in multi-tenant Field Programmable Gate Array (FPGA) scenarios due to the low correlation between power traces and cipher keys. To address this challenge, we propose Linearized Deep Learning Side-Channel Attack (LDL-SCA). Through modifying the output layer and integrating K-means clustering, LDL-SCA is capable of capturing linear features regardless of the low correlation between input and label. Moreover, we introduce new evaluation metrics derived from R2 and Cohen-kappa score. Our experiments show that LDL-SCA generate results with improved distinguishability, which has about ten times smaller standard deviation and ten times larger peak differences compared with Correlation Power Analysis (CPA) and Linear Regression Analysis (LRA). Yankun Zhu, Siting Liu 0001, Liyu Yang, Pingqiang Zhou |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Feature-Embedding Triplet Networks with a Separately Constrained Loss FunctionabstractFeature-embedding triplet networks (TNs) with three symmetric subchannels are very promising for similarity-measuring applications. This paper proposes a novel separately constrained triple loss (SCTL) function that applies to TNs for classification. Through minimizing the intra-class distance and maximizing the inter-class distance, SCTL eliminates possible false solutions and provides insight into the dependency of training based on these two terms. Based on this dependency, the strategy of selecting hyperparameters in SCTL is also analyzed to further improve performance. The effectiveness of the proposed SCTL is evaluated based on TNs with multi-layer perceptrons; the results show that compared to all existing loss functions, the use of SCTL offers the best classification accuracy for the TNs, while incurring in negligible hardware overhead (e.g., only a 0.0002% area overhead of the subnetworks). Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Honglan Jiang, Siting Liu 0001, Pedro Reviriego, Fabrizio Lombardi |
ISCAS | 5 |
| 2023 | An Energy-Efficient Binary-Interfaced Stochastic Multiplier Using Parallel DatapathsabstractStochastic computing (SC) typically requires a low design complexity compared with weighted binary computing, so it has been successfully applied in neural networks (NNs). Usually, SC utilizes random bitstreams as its medium, which makes it suffer from a long delay that offsets its advantages. This drawback can be alleviated by utilizing parallel datapaths, which, however, will significantly increase the hardware cost due to the requirement of multiple parallel computing units. In this article, a hybrid bit-splitting generator (HBSG) is proposed to efficiently produce parallel bitstreams in a single clock cycle to reduce delay. The HBSG uniformly splits binary numbers into R segments, each of which is encoded in parallel by using hardwired connections according to the weight of each bit. A binary-interfaced parallel stochastic multiplier (BipSMul) using the HBSG is then proposed to accelerate the multiplication in SC. Experimental results show that the BipSMul is more energy efficient than the state-of-the-art parallel and serial stochastic designs, as well as their binary and Booth counterparts, in delay, power-delay product (PDP), and area-delay product (ADP). Yongqiang Zhang 0006, Siting Liu 0001, Jie Han 0001, Zhendong Lin, Xin Cheng 0001, Guangjun Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Special Session: Fault-Tolerant Deep Learning: A Hierarchical PerspectiveabstractWith the rapid advancements of deep learning in the past decade, it can be foreseen that deep learning will be continuously deployed in more and more safety-critical applications such as autonomous driving and robotics. In this context, reliability turns out to be critical to the deployment of deep learning in these applications and gradually becomes a first-class citizen among the major design metrics like performance and energy efficiency. Nevertheless, the back-box deep learning models combined with the diverse underlying hardware faults make resilient deep learning extremely challenging. In this special session, we conduct a comprehensive survey of fault-tolerant deep learning design approaches with a hierarchical perspective and investigate these approaches from model layer, architecture layer, circuit layer, and cross layer respectively. Cheng Liu 0008, Zhen Gao 0005, Siting Liu 0001, Xuefei Ning, Huawei Li 0001, Xiaowei Li 0001 |
VTS | 3 |
| 2021 | A Survey of Stochastic Computing Neural Networks for Machine Learning ApplicationsabstractNeural networks (NNs) are effective machine learning models that require significant hardware and energy consumption in their computing process. To implement NNs, stochastic computing (SC) has been proposed to achieve a tradeoff between hardware efficiency and computing performance. In an SC NN, hardware requirements and power consumption are significantly reduced by moderately sacrificing the inference accuracy and computation speed. With recent developments in SC techniques, however, the performance of SC NNs has substantially been improved, making it comparable with conventional binary designs yet by utilizing less hardware. In this article, we begin with the design of a basic SC neuron and then survey different types of SC NNs, including multilayer perceptrons, deep belief networks, convolutional NNs, and recurrent NNs. Recent progress in SC designs that further improve the hardware efficiency and performance of NNs is subsequently discussed. The generality and versatility of SC NNs are illustrated for both the training and inference processes. Finally, the advantages and challenges of SC NNs are discussed with respect to binary counterparts. Yidong Liu, Siting Liu 0001, Yanzhi Wang 0001, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Dynamic Stochastic Computing for Digital Signal Processing ApplicationsabstractStochastic computing (SC) utilizes a random binary bit stream to encode a number by counting the frequency of 1's in the stream (or sequence). Typically, a small circuit is used to perform a bit-wise logic operation on the stochastic sequences, which leads to significant hardware and power savings. Energy efficiency, however, is a challenge for SC due to the long sequences required for accurately encoding numbers. To overcome this challenge, we consider to use a stochastic sequence to encode a continuously variable signal instead of a number to achieve higher accuracy, higher energy efficiency and greater flexibility. Specifically, one single bit is used to encode a sample from a signal for efficient processing. This type of sequences encodes constantly variable values, so it is referred to as dynamic stochastic sequences (DSS's). The DSS enables the use of SC circuits to efficiently perform tasks such as frequency mixing and function estimation. It is shown that such a dynamic SC (DSC) system achieves savings up to 98.4% in energy and up to 96.8% in time with a slightly higher accuracy compared to conventional SC. It also achieves energy and time savings of up to 60% compared to a fixed-width binary implementation. Siting Liu 0001, Jie Han 0001 |
DATE | 1 |
| 2018 | A Stochastic Computational Multi-Layer Perceptron with Backward PropagationabstractStochastic computation has recently been proposed for implementing artificial neural networks with reduced hardware and power consumption, but at a decreased accuracy and processing speed. Most existing implementations are based on pre-training such that the weights are predetermined for neurons at different layers, thus these implementations lack the ability to update the values of the network parameters. In this paper, a stochastic computational multi-layer perceptron (SC-MLP) is proposed by implementing the backward propagation algorithm for updating the layer weights. Using extended stochastic logic (ESL), a reconfigurable stochastic computational activation unit (SCAU) is designed to implement different types of activation functions such as the tanh and the rectifier function. A triple modular redundancy (TMR) technique is employed for reducing the random fluctuations in stochastic computation. A probability estimator (PE) and a divider based on the TMR and a binary search algorithm are further proposed with progressive precision for reducing the required stochastic sequence length. Therefore, the latency and energy consumption of the SC-MLP are significantly reduced. The simulation results show that the proposed design is capable of implementing both the training and inference processes. For the classification of nonlinearly separable patterns, at a slight loss of accuracy by 1.32-1.34 percent, the proposed design requires only 28.5-30.1 percent of the area and 18.9-23.9 percent of the energy consumption incurred by a design using floating point arithmetic. Compared to a fixed-point implementation, the SC-MLP consumes a smaller area (40.7-45.5 percent) and a lower energy consumption (38.0-51.0 percent) with a similar processing speed and a slight drop of accuracy by 0.15-0.33 percent. The area and the energy consumption of the proposed design is from 80.7-87.1 percent and from 71.9-93.1 percent, respectively, of a binarized neural network (BNN), with a similar accuracy. Yidong Liu, Siting Liu 0001, Yanzhi Wang 0001, Fabrizio Lombardi, Jie Han 0001 |
IEEE Trans. Computers | 2 |
| 2018 | Gradient Descent Using Stochastic Circuits for Efficient Training of Learning MachinesabstractGradient descent (GD) is a widely used optimization algorithm in machine learning. In this paper, a novel stochastic computing GD circuit (SC-GDC) is proposed by encoding the gradient information in stochastic sequences. Inspired by the structure of a neuron, a stochastic integrator is used to optimize the weights in a learning machine by its “inhibitory” and “excitatory” inputs. Specifically, two AND (or XNOR) gates for the unipolar representation (or the bipolar representation) and one stochastic integrator are, respectively, used to implement the multiplications and accumulations in a GD algorithm. Thus, the SC-GDC is very area- and power-efficient. As per the formulation of the proposed SC-GDC, it provides unbiased estimate of the optimized weights in a learning algorithm. The proposed SC-GDC is then used to implement a least-mean-square algorithm and a softmax regression. With a similar accuracy, the proposed design achieves more than $30 \times $ improvement in throughput per area (TPA) and consumes less than 13% of the energy per training sample, compared with a fixed-point implementation. Moreover, a signed SC-GDC is proposed for training complex neural networks (NNs). It is shown that for a 784-128-128-10 fully connected NN, the signed SC-GDC produces a similar training result with its fixed-point counterpart, while achieving more than 90% energy saving and 82% reduction in training time with more than $50 \times $ improvement in TPA. Siting Liu 0001, Honglan Jiang, Leibo Liu, Jie Han 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Toward Energy-Efficient Stochastic Circuits Using Parallel Sobol Sequences
Siting Liu 0001, Jie Han 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Hardware ODE Solvers using Stochastic CircuitsabstractA novel ordinary differential equation (ODE) solver is proposed by using a stochastic integrator to implement the accumulative function of the Euler method. We show that a stochastic integrator is an unbiased estimator for a Euler numerical solution. Unlike in conventional stochastic circuits, in which long stochastic bit streams are required to produce a result with a high accuracy, the proposed stochastic ODE solver provides an estimate of the solution for every bit in the stochastic bit stream, thus significantly reducing the latency and energy consumption of the circuit. Complex ODE solvers are constructed for solving nonhomogeneous ODEs, systems of ODEs and higher-order ODEs. Experimental results show that the stochastic ODE solvers provide very accurate solutions compared to their binary counterparts, with on average an energy saving of 46% (up to 74%), 8x throughput per area (up to nearly 12x) and a runtime reduction of 72% (up to 82%). Siting Liu 0001, Jie Han 0001 |
DAC | 1 |
| 2017 | Energy efficient stochastic computing with Sobol sequencesabstractEnergy efficiency presents a significant challenge for stochastic computing (SC) due to the long random binary bit streams required for accurate computation. In this paper, a type of low discrepancy (LD) sequences, the Sobol sequence, is considered for energy-efficient implementations of SC circuits. The use of Sobol sequences improves the output accuracy of a stochastic circuit with a reduced sequence length compared to the use of another type of LD sequences, the Halton sequence, and conventional linear feedback shift register (LFSR)-generated pseudorandom sequence. The use of Sobol sequences leads to a similar or higher accuracy than using Halton sequences for basic arithmetic operations. Sobol sequence generators cost less energy than the Halton counterparts when multiple random sequences are required in a circuit, thus the use of Sobol sequences can lead to a higher energy efficiency in an SC circuit than using Halton sequences. Siting Liu 0001, Jie Han 0001 |
DATE | 1 |