VLDB 2026 Research / reviewers in the wild / expert
Farzad Niknia
dblp:277/5369
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-4062-3638ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perturbation-based error detection and correction (PBEDC) in dependable large-scale machine learning systems
Ziheng Wang 0005, Pedro Reviriego, Shanshan Liu 0001, Farzad Niknia, Xiaochen Tang, Zhen Gao 0005, Fabrizio Lombardi |
Future Gener. Comput. Syst. | 4 |
| 2025 | Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)abstractStochastic computing (SC) has emerged as an efficient low-power alternative for deploying neural networks (NNs) in resource-limited scenarios, such as the Internet of Things (IoT). By encoding values as serial bitstreams, SC significantly reduces energy dissipation compared to conventional floating-point (FP) designs; however, further improvement of layer-wise mixed-precision implementation for SC remains unexplored. This paper introduces Adjustable Sequence Length (ASL), a novel scheme that applies mixedprecision concepts specifically to SC NNs. By introducing an operator-norm – based theoretical model, this paper shows that truncation noise can cumulatively propagate through the layers by the estimated amplification factors. An extended sensitivity analysis is presented, using Random Forest (RF) regression to evaluate multi-layer truncation effects and validate the alignment of theoretical predictions with practical network behaviors. To accommodate different application scenarios, this paper proposes two truncation strategies (coarse-grained and fine-grained), which apply diverse sequence length configurations at each layer. Evaluations on a pipelined SC MLP synthesized at 32 nm demonstrate that ASL can reduce energy and latency overheads by up to over 60% with negligible accuracy loss. It confirms the feasibility of the ASL scheme for IoT applications and highlights the distinct advantages of mixed-precision truncation in SC designs. Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Zhen Gao 0005, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Internet Things J. | 3 |
| 2024 | Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of ThingsabstractThe implementation of Machine Learning (ML) in Internet of Things (IoT) devices poses significant operational challenges due to limited energy and computation resources. In recent years, significant efforts have been made to implement simplified ML models that can achieve reasonable performance while reducing computation and energy, for example by pruning weights in neural networks, or using reduced precision for the parameters and arithmetic operations. However, this type of approach is limited by the performance of the ML implementation, i.e., by the loss for example in accuracy due to the model simplification. In this paper, we present Adaptive Resolution Inference (ARI), a novel approach that enables to evaluate new trade-offs between energy dissipation and model performance in ML implementations. The main principle of the proposed approach is to run inferences with reduced precision (quantization) and use the margin over the decision threshold to determine if either the result is reliable, or the inference must run with the full model. The rationale is that quantization only introduces small deviations in the inference scores, such that if the scores have a sufficient margin over the decision threshold, it is very unlikely that the full model would have a different result. Therefore, we can run the quantized model first, and only when the scores do not have a sufficient margin, the full model is run. This enables most inferences to run with the reduced precision model and only a small fraction requires the full model, so significantly reducing computation and energy while not affecting model performance. The proposed ARI approach is presented, analyzed in detail, and evaluated using different datasets both for floating-point and stochastic computing implementations. The results show that ARI can significantly reduce the energy for inference in different configurations with savings between 40% and 85%. Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Internet Things J. | 3 |
| 2024 | Concurrent Classifier Error Detection (CCED) in Large Scale Machine Learning SystemsabstractThe complexity of machine learning (ML) systems increases each year. As these systems are widely utilized, ensuring their reliable operation is becoming a design requirement. Traditional error detection mechanisms introduce circuit or time redundancy that significantly impacts system performance. An alternative is the use of concurrent error detection (CED) schemes that operate in parallel with the system and exploit their properties to detect errors. CED is attractive for large ML systems because it can potentially reduce the cost of error detection. In this article, we introduce concurrent classifier error detection (CCED), a scheme to implement CED in ML systems using a concurrent ML classifier to detect errors. CCED identifies a set of check signals in the main ML system and feed them to the concurrent ML classifier that is trained to detect errors. The proposed CCED scheme has been implemented and evaluated on two widely used large-scale ML models: Contrastive language-image pretraining (CLIP) used for image classification and bidirectional encoder representations from transformers (BERT) used for natural language applications. The results show that more than 95% of the errors are detected when using a simple Random Forest classifier that is orders of magnitude simpler than CLIP or BERT. Pedro Reviriego, Ziheng Wang 0005, Zhen Gao 0005, Farzad Niknia, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Reliab. | 5 |
| 2023 | Feature-Embedding Triplet Networks with a Separately Constrained Loss FunctionabstractFeature-embedding triplet networks (TNs) with three symmetric subchannels are very promising for similarity-measuring applications. This paper proposes a novel separately constrained triple loss (SCTL) function that applies to TNs for classification. Through minimizing the intra-class distance and maximizing the inter-class distance, SCTL eliminates possible false solutions and provides insight into the dependency of training based on these two terms. Based on this dependency, the strategy of selecting hyperparameters in SCTL is also analyzed to further improve performance. The effectiveness of the proposed SCTL is evaluated based on TNs with multi-layer perceptrons; the results show that compared to all existing loss functions, the use of SCTL offers the best classification accuracy for the TNs, while incurring in negligible hardware overhead (e.g., only a 0.0002% area overhead of the subnetworks). Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Honglan Jiang, Siting Liu 0001, Pedro Reviriego, Fabrizio Lombardi |
ISCAS | 2 |
| 2023 | Tolerance of Siamese Networks (SNs) to Memory Errors: Analysis and DesignabstractThis article considers memory errors in a Siamese Network (SN) through an extensive analysis and proposes two schemes (using a weight filter and a code) to provide efficient hardware solutions for error tolerance. Initially the impact of memory errors on the weights of the SN (stored as floating-point (FP) numbers) is analyzed; this shows that the degradation is mostly caused by outliers in weights. Two schemes are subsequently proposed. An analysis is pursued to establish the filter's bounds selection by the maximum/minimum values of the weight distributions, by which outliers can be removed from the operation of the SN. A code scheme for protecting the sign and exponent bits of each weight in an FP number, is also proposed; this code incurs in no memory overhead by utilizing the 4 least significant bits (LSB) to store parity bits. Simulation shows that the filter has a better performance for multi-bit errors correction (a reduction of 95.288% in changed predictions), while the code achieves superior results in single-bit errors correction (a reduction of 99.775% in changed predictions). The combined method that uses the two proposed schemes, retains their advantages, so adaptive to all scenarios; The ASIC-based FP designs of the SN using serial and hybrid implementations are also presented; these pipelined designs utilize a novel multi-layer perceptron (MLP) (as branch networks of the SN) that operates at a frequency of 681.2 MHz (at a 32nm technology node), so significantly higher than existing designs found in the technical literature. The proposed error-tolerant approaches also show advantages in overheads comparing with for example traditional error correction code (ECC). These error-tolerant MLP-based designs are well suited to hardware/power-constrained platforms. Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2022 | Aging Effects on Template Attacks Launched on Dual-Rail Protected ChipsabstractProfiling side-channel attacks in which an adversary creates a “profile” of a sensitive device and uses such a profile to model a target device with similar implementation has received the lion’s share of attention in the recent years. In particular, template attacks are known to be the most powerful profiling side-channel attacks from an information theoretic point of view. When launching such an attack, the adversary first builds a model based on the leakage of the profiling (training) device in his disposal, which is then exploited in the second phase of the attack (i.e., matching) to extract the key from the target device. Discrepancies between the device used for modeling and the target device affect the attack success. The effect of process variation and temperature misalignment between the profiling and target devices in the template attack’s success has been studied extensively in the literature, while the impact of device aging on the template attack’s success is yet to be investigated thoroughly. This article moves one step forward and studies the impact of device aging, mainly bias temperature instability (BTI) and hot carrier injection (HCI), in the devices that have been protected against power analysis attacks via dual rail logics. In particular, we focus on the wave dynamic differential logic (WDDL) circuits, and via extensive transistor-level simulations, we will show how device aging misalignments between the profiling and target devices can hinder template attacks for both unprotected and WDDL protected counterparts. We mounted several attacks on the PRESENT cipher, with and without WDDL protection, at different temperatures and aging times. Our results show that the attack is more difficult if there is an aging-duration mismatch between the training and target devices, and the attack-efficiency decrease is especially significant for mismatches of few weeks. Farzad Niknia, Jean-Luc Danger, Sylvain Guilley, Naghmeh Karimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | A Delta Sigma Modulator-Based Stochastic DividerabstractThe divider is one of the most complex hardware units in Stochastic Computing (SC); even though several new designs have been presented to reduce the computation latency of the conventional divider, all of them still require a considerable number of clock cycles. Moreover, they incur in low performance due to the employed arithmetic computational scheme. In this paper, a Delta Sigma Modulator (DSM) based stochastic divider is proposed. As an entirely digital circuit, the proposed divider offers the best computation latency and accuracy over all existing stochastic dividers found in the technical literature (with a typical reduction between 66.8% and 96.9% in the number of clock cycles and a reduction from$10^{\mathrm {-3.4}}$to$10^{\mathrm {-3.9}}$in the average mean square error for a 10-bit resolution). An SC-based Neural Network (NN) is considered as an initial case study to evaluate the advantages of the proposed design in an emerging application; results show that the proposed divider enables an SC-based NN to achieve a higher classification accuracy and hardware efficiency than existing designs. To show the flexibility of the proposed divider design, its application to Sobol-based sequences is also presented; also in this case, its superiority over other designs is confirmed. These features make the proposed design very attractive for hardware-constrained platforms; moreover, such a novel design approach that incorporates ideas from analog/mixed signal circuit design into a digital circuit design, can motivate other researchers to design efficient SC designs using similar schemes. Xiaochen Tang, Shanshan Liu 0001, Farzad Niknia, Pedro Reviriego, Ziheng Wang 0005, Wei Tang 0002, Ahmed Louri, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Learning Assisted Side Channel Delay Test for Detection of Recycled ICsabstractWith the outsourcing of design flow, ensuring the security and trustworthiness of integrated circuits has become more challenging. Among the security threats, IC counterfeiting and recycled ICs have received a lot of attention due to their inferior quality, and in turn, their negative impact on the reliability and security of the underlying devices. Detecting recycled ICs is challenging due to the effect of process variations and process drift occurring during the chip fabrication. Moreover, relying on a golden chip as a basis for comparison is not always feasible. Accordingly, this paper presents a recycled IC detection scheme based on delay side-channel testing. The proposed method relies on the features extracted during the design flow and the sample delays extracted from the target chip to build a Neural Network model using which the target chip can be truly identified as new or recycled. The proposed method classifies the timing paths of the target chip into two groups based on their vulnerability to aging using the information collected from the design and detects the recycled ICs based on the deviation of the delay of these two sets from each other. Ashkan Vakil, Farzad Niknia, Ali Mirzaeian, Avesta Sasan, Naghmeh Karimi |
ASP-DAC | 2 |
| 2021 | Stochastic Dividers for Low Latency Neural NetworksabstractDue to the low complexity in arithmetic unit design, stochastic computing (SC) has attracted considerable interest to implement Artificial Neural Networks (ANNs) for resources-limited applications, because ANNs must usually perform a large number of arithmetic operations. To attain a high computation accuracy in an SC-based ANN, extended stochastic logic is utilized together with standard SC units and thus, a stochastic divider is required to perform the conversion between these logic representations. However, the conventional divider incurs in a large computation latency, so limits an SC implementation for ANNs used in applications needing high performance. Therefore, there is a need to design fast stochastic dividers for SC-based ANNs. Recent works (e.g., a binary searching and triple modular redundancy (BS-TMR) based stochastic divider) are targeting a reduction in computation latency, while keeping the same accuracy compared with the traditional design. However, this divider still requires$N$iterations to deal with$2^{N}$-bit stochastic sequences, and thus the latency increases in proportion to the sequence length. In this paper, a decimal searching and TMR (DS-TMR) based stochastic divider is initially proposed to further reduce the computation latency; it only requires two iterations to calculate the quotient, so regardless of the sequence length. Moreover, a trade-off design between accuracy and hardware is also presented. An SC-based Multi-Layer Perceptron (MLP) is then considered to show the effectiveness of the proposed dividers over current designs. Results show that when utilizing the proposed dividers, the MLP achieves the lowest computation latency while keeping the same classification accuracy; although incurring in an area increase, the overhead due to the proposed dividers is low over the entire MLP. When using as combined metric for both hardware design and computation complexity the product of the implementation area, latency, power and number of clock cycles, the proposed designs are also shown to be superior to the SC-based MLPs (at the same level of accuracy) employing other dividers found in the technical literature as well as the commonly used 32-bit floating point implementation. Shanshan Liu 0001, Xiaochen Tang, Farzad Niknia, Pedro Reviriego, Weiqiang Liu 0001, Ahmed Louri, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |