Ziheng Wang 0005

dblp:79/10743-5 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-9668-7318ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Perturbation-based error detection and correction (PBEDC) in dependable large-scale machine learning systems
Ziheng Wang 0005, Pedro Reviriego, Shanshan Liu 0001, Farzad Niknia, Xiaochen Tang, Zhen Gao 0005, Fabrizio Lombardi
Future Gener. Comput. Syst.1
2025 Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)
abstract
Stochastic computing (SC) has emerged as an efficient low-power alternative for deploying neural networks (NNs) in resource-limited scenarios, such as the Internet of Things (IoT). By encoding values as serial bitstreams, SC significantly reduces energy dissipation compared to conventional floating-point (FP) designs; however, further improvement of layer-wise mixed-precision implementation for SC remains unexplored. This paper introduces Adjustable Sequence Length (ASL), a novel scheme that applies mixedprecision concepts specifically to SC NNs. By introducing an operator-norm – based theoretical model, this paper shows that truncation noise can cumulatively propagate through the layers by the estimated amplification factors. An extended sensitivity analysis is presented, using Random Forest (RF) regression to evaluate multi-layer truncation effects and validate the alignment of theoretical predictions with practical network behaviors. To accommodate different application scenarios, this paper proposes two truncation strategies (coarse-grained and fine-grained), which apply diverse sequence length configurations at each layer. Evaluations on a pipelined SC MLP synthesized at 32 nm demonstrate that ASL can reduce energy and latency overheads by up to over 60% with negligible accuracy loss. It confirms the feasibility of the ASL scheme for IoT applications and highlights the distinct advantages of mixed-precision truncation in SC designs.
Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Zhen Gao 0005, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi
IEEE Internet Things J.1
2024 Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of Things
abstract
The implementation of Machine Learning (ML) in Internet of Things (IoT) devices poses significant operational challenges due to limited energy and computation resources. In recent years, significant efforts have been made to implement simplified ML models that can achieve reasonable performance while reducing computation and energy, for example by pruning weights in neural networks, or using reduced precision for the parameters and arithmetic operations. However, this type of approach is limited by the performance of the ML implementation, i.e., by the loss for example in accuracy due to the model simplification. In this paper, we present Adaptive Resolution Inference (ARI), a novel approach that enables to evaluate new trade-offs between energy dissipation and model performance in ML implementations. The main principle of the proposed approach is to run inferences with reduced precision (quantization) and use the margin over the decision threshold to determine if either the result is reliable, or the inference must run with the full model. The rationale is that quantization only introduces small deviations in the inference scores, such that if the scores have a sufficient margin over the decision threshold, it is very unlikely that the full model would have a different result. Therefore, we can run the quantized model first, and only when the scores do not have a sufficient margin, the full model is run. This enables most inferences to run with the reduced precision model and only a small fraction requires the full model, so significantly reducing computation and energy while not affecting model performance. The proposed ARI approach is presented, analyzed in detail, and evaluated using different datasets both for floating-point and stochastic computing implementations. The results show that ARI can significantly reduce the energy for inference in different configurations with savings between 40% and 85%.
Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi
IEEE Internet Things J.1
2024 Concurrent Classifier Error Detection (CCED) in Large Scale Machine Learning Systems
abstract
The complexity of machine learning (ML) systems increases each year. As these systems are widely utilized, ensuring their reliable operation is becoming a design requirement. Traditional error detection mechanisms introduce circuit or time redundancy that significantly impacts system performance. An alternative is the use of concurrent error detection (CED) schemes that operate in parallel with the system and exploit their properties to detect errors. CED is attractive for large ML systems because it can potentially reduce the cost of error detection. In this article, we introduce concurrent classifier error detection (CCED), a scheme to implement CED in ML systems using a concurrent ML classifier to detect errors. CCED identifies a set of check signals in the main ML system and feed them to the concurrent ML classifier that is trained to detect errors. The proposed CCED scheme has been implemented and evaluated on two widely used large-scale ML models: Contrastive language-image pretraining (CLIP) used for image classification and bidirectional encoder representations from transformers (BERT) used for natural language applications. The results show that more than 95% of the errors are detected when using a simple Random Forest classifier that is orders of magnitude simpler than CLIP or BERT.
Pedro Reviriego, Ziheng Wang 0005, Zhen Gao 0005, Farzad Niknia, Shanshan Liu 0001, Fabrizio Lombardi
IEEE Trans. Reliab.2
2023 Feature-Embedding Triplet Networks with a Separately Constrained Loss Function
abstract
Feature-embedding triplet networks (TNs) with three symmetric subchannels are very promising for similarity-measuring applications. This paper proposes a novel separately constrained triple loss (SCTL) function that applies to TNs for classification. Through minimizing the intra-class distance and maximizing the inter-class distance, SCTL eliminates possible false solutions and provides insight into the dependency of training based on these two terms. Based on this dependency, the strategy of selecting hyperparameters in SCTL is also analyzed to further improve performance. The effectiveness of the proposed SCTL is evaluated based on TNs with multi-layer perceptrons; the results show that compared to all existing loss functions, the use of SCTL offers the best classification accuracy for the TNs, while incurring in negligible hardware overhead (e.g., only a 0.0002% area overhead of the subnetworks).
Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Honglan Jiang, Siting Liu 0001, Pedro Reviriego, Fabrizio Lombardi
ISCAS1
2023 Tolerance of Siamese Networks (SNs) to Memory Errors: Analysis and Design
abstract
This article considers memory errors in a Siamese Network (SN) through an extensive analysis and proposes two schemes (using a weight filter and a code) to provide efficient hardware solutions for error tolerance. Initially the impact of memory errors on the weights of the SN (stored as floating-point (FP) numbers) is analyzed; this shows that the degradation is mostly caused by outliers in weights. Two schemes are subsequently proposed. An analysis is pursued to establish the filter's bounds selection by the maximum/minimum values of the weight distributions, by which outliers can be removed from the operation of the SN. A code scheme for protecting the sign and exponent bits of each weight in an FP number, is also proposed; this code incurs in no memory overhead by utilizing the 4 least significant bits (LSB) to store parity bits. Simulation shows that the filter has a better performance for multi-bit errors correction (a reduction of 95.288% in changed predictions), while the code achieves superior results in single-bit errors correction (a reduction of 99.775% in changed predictions). The combined method that uses the two proposed schemes, retains their advantages, so adaptive to all scenarios; The ASIC-based FP designs of the SN using serial and hybrid implementations are also presented; these pipelined designs utilize a novel multi-layer perceptron (MLP) (as branch networks of the SN) that operates at a frequency of 681.2 MHz (at a 32nm technology node), so significantly higher than existing designs found in the technical literature. The proposed error-tolerant approaches also show advantages in overheads comparing with for example traditional error correction code (ECC). These error-tolerant MLP-based designs are well suited to hardware/power-constrained platforms.
Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi
IEEE Trans. Computers1
2022 A Delta Sigma Modulator-Based Stochastic Divider
abstract
The divider is one of the most complex hardware units in Stochastic Computing (SC); even though several new designs have been presented to reduce the computation latency of the conventional divider, all of them still require a considerable number of clock cycles. Moreover, they incur in low performance due to the employed arithmetic computational scheme. In this paper, a Delta Sigma Modulator (DSM) based stochastic divider is proposed. As an entirely digital circuit, the proposed divider offers the best computation latency and accuracy over all existing stochastic dividers found in the technical literature (with a typical reduction between 66.8% and 96.9% in the number of clock cycles and a reduction from$10^{\mathrm {-3.4}}$to$10^{\mathrm {-3.9}}$in the average mean square error for a 10-bit resolution). An SC-based Neural Network (NN) is considered as an initial case study to evaluate the advantages of the proposed design in an emerging application; results show that the proposed divider enables an SC-based NN to achieve a higher classification accuracy and hardware efficiency than existing designs. To show the flexibility of the proposed divider design, its application to Sobol-based sequences is also presented; also in this case, its superiority over other designs is confirmed. These features make the proposed design very attractive for hardware-constrained platforms; moreover, such a novel design approach that incorporates ideas from analog/mixed signal circuit design into a digital circuit design, can motivate other researchers to design efficient SC designs using similar schemes.
Xiaochen Tang, Shanshan Liu 0001, Farzad Niknia, Pedro Reviriego, Ziheng Wang 0005, Wei Tang 0002, Ahmed Louri, Fabrizio Lombardi
IEEE Trans. Circuits Syst. I Regul. Pap.5