EDBT 2026 Demo / reviewers in the wild / expert
Siavash Bayat Sarmadi
dblp:45/3186
· DBLP profile ↗
27ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-3294-2505ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 6 first-author · 5 since 2021Computer networks · 7 · 5 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iMAC: Toward Accelerating Deep Neural Networks by Predicting and Removing Ineffectual MAC OperationsabstractConvolutional Neural Networks (CNNs) are employed in a broad range of classification tasks. Despite CNNs have reached high accuracy, they have high computational complexity, so employing this method in many applications such as internet of things (IoT) is challenging. To overcome this challenge, CNN hardware accelerators are proposed. As the main activation function employed in the CNNs is ReLU, a large portion of the output becomes zero. Therefore, sparse CNN accelerators are proposed to reduce the computational complexity. Recent studies show that, rather than employing sparse CNN accelerators, it would be ideal to predict the zero outputs and remove the corresponding computations. In this paper, we propose the iMAC accelerator to bypass ineffectual computations (a computation leads to zero). For this purpose, we propose a prediction unit that determines ineffectual computations without accuracy loss. Then we propose a dataflow to bypass ineffectual computations. The results show that by increasing 12% of area and 18% power overhead, iMAC achieves 2.31× speedup and 1.95× energy efficiency without any accuracy loss. With 3% accuracy loss, it further improves the speed and the energy by 2.94× and 2.49×, respectively. Moreover, compared to the previous prediction method, our method achieves higher speedup and improvement. Farhad Taheri, Siavash Bayat Sarmadi, Alireza Tabatabaeian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | Post-Quantum Authentication and Communication Security for Drone NetworksabstractModern societies increasingly rely on interconnected devices and individuals for information exchange and monitoring. In this manner, drones, as flying smart nodes in the wireless network ecosystem, facilitate diverse applications, such as industrial, civilian, and disaster response operations. However, the widespread adoption of drone networks introduces significant challenges regarding interoperability, privacy, and security. Advances in quantum computing exacerbate these issues by rendering classical cryptographic measures inadequate, necessitating the adoption of quantum-safe approaches. Existing solutions fail to address all potential threats or lack post-quantum security. This work sets out to propose a robust lattice-based authentication and communication protocol for drone networks, addressing these challenges. The protocol integrates dynamic credentials, timestamps, multi-factor authentication, and context information to enhance security. The scheme's security is analyzed using the DY and BPR models and formally verified via the AVISPA tool. Performance evaluations and comparisons with existing solutions demonstrate the protocol's functionality, showcasing a reasonable trade-off between security and performance. By leveraging modern cryptographic techniques, the proposed scheme ensures reliable and robust drone operations, effectively mitigating various attacks beyond session establishment while safeguarding privacy, and supporting dynamic network events. Parya Derakhshan Roodsari, Siavash Bayat Sarmadi, Hatameh Mosanaei-Boorani |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | PartialHD: Toward Efficient Hyperdimensional Computing by Partial ProcessingabstractHyperdimensional (HD) computing is a brain-inspired learning method that is widely employed in resource-constrained applications, such as the Internet of Things (IoT) regarding its lightweight computation. Although HD computing has high efficiency in the IoT applications, it suffers from high computation due to the large vector size. Thus, several studies are proposed to speedup and increase the efficiency of HD computing. In this work, we propose a method to process HD computing called PartialHD. Our method divides a long hypervector into multiple partial vectors and processes each vector separately. In the retraining phase, this method improves accuracy up to 1.93%. Moreover, by employing our proposed method in the retraining and inference, only a few partial vectors participate and this causes the computational overhead reduces in these phases. The evaluation shows that our method processes on average 22.2% and 36.23% of the entire hypervectors with negligible accuracy loss in inference and retraining, respectively. Furthermore, we propose two general architectures (lightweight and high-speed), which accelerate partial vector computations on different FPGA platforms. The result shows that the lightweight architecture can accelerate HD computing on resource-constrained FPGAs such as Artix while the high-speed architecture has$4.72\times $higher throughput than the lightweight architecture. Farhad Taheri, Siavash Bayat Sarmadi, Hamed Rastaghi |
IEEE Internet Things J. | 2 |
| 2024 | A Digital Signature Architecture Suitable for V2V ApplicationsabstractThe elliptic curve digital signature algorithm (ECDSA) is widely used for guaranteeing data integrity and user authentication in internet of things (IoT) applications such as intelligent transport systems (ITS). In ITS, vehicles, infrastructures, and data networks communicate using vehicle-to-everything (V2X) protocols. In V2X message broadcasting, the ECDSA guarantees data security and privacy. During traffic congestion, the signature generation/verification latency becomes crucial. Hence, this paper proposes a high-throughput and efficient ECDSA architecture for vehicle-to-vehicle (V2V) applications. We investigate the double point multiplication (DPM) method for reducing computational overhead and propose a new finite field multiplier architecture for latency improvement. Our implementation results over$p_{256}$on Virtex-7 field programmable gate array (FPGA) show that our design’s throughput and efficiency are improved compared to previous works by at least$4.9\times $and$7.2\times $, respectively. This unit generates a signature in$167 {\mathrm {\mu \text { s} }}$and verifies a message in$188 {\mathrm {\mu \text { s} }}$. Also, our application-specific integrated circuit (ASIC) synthesis on$45 {\mathrm { \text {n} \text { m} }}$and$180 {\mathrm { \text {n} \text { m} }}$technologies can verify 4416 messages per second by consuming$1.3 {\mathrm {\mu \text { J} }}$and$9.4 {\mathrm {\mu \text { J} }}$energy, respectively. This advantage makes the design affordable for other IoT applications. Hatameh Mosanaei-Boorani, Siavash Bayat Sarmadi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | High-Speed Post-Quantum Cryptoprocessor Based on RISC-V Architecture for IoTabstractPublic-key plays a significant role in today’s communication over the network. However, current state-of-the-art public-key encryption (PKE) schemes are too complex to be efficiently employed in resource-constrained devices. Moreover, they are vulnerable to quantum attacks and soon will not have the required security. In the last decade, lattice-based cryptography has been a progenitor platform of the post-quantum cryptography (PQC) due to its lower complexity, which makes it more suitable for Internet of Things applications. In this article, we propose an efficient implementation of the binary learning with errors over ring (Ring-BinLWE) on the reduced instruction set computer-five (RISC-V) platform. Our field-programmable gate array (FPGA) implementations improve the speed of scheme operations by more than 51% in terms of CPU cycles compared to previous work. The proposed hardware module has low complexity and only imposes around 6%–9% overhead to the original core. Moreover, it has constant-time operations and is resistant to timing attacks. Besides, a more reliable fault-resilient variant of the architecture with 1% area overhead is proposed. According to the application-specific integrated circuits (ASICs) implementations, the proposed architecture achieves at least 32%, 79%, and 50% lower power, energy, and area consumption compared to the previous work, respectively. Shahriar Hadayeghparast, Siavash Bayat Sarmadi, Shahriar Ebrahimi |
IEEE Internet Things J. | 2 |
| 2022 | RISC-HD: Lightweight RISC-V Processor for Efficient Hyperdimensional Computing InferenceabstractHyperdimensional (HD) computing is a lightweight machine learning method widely used in Internet of Things applications for classification tasks. Although many hardware accelerators are proposed to improve the performance of HD, they suffer from low flexibility that makes them not practical in most real-life scenarios. To improve the flexibility, an open-source instruction set architecture (ISA) called RISC-V has been employed and extended for a specific application such as machine learning. This article aims to improve the efficiency and flexibility of HD computing for resource-constrained applications. To this end, we extend a RISC-V core (RI5CY) for HD computing called RISC-HD. First, to reduce the computational overhead at the HD inference phase, we introduce a pruning method to remove the ineffectual dimensions. The proposed pruning method can reduce the dimension from 10k to 1k with negligible accuracy loss. Second, an ISA extension for RI5CY is proposed to compute the HD inference efficiently. Experimental results indicate that RISC-HD adds$1.42\times $area overhead to the RI5CY core; however, it consumes only 2932 slices on the Artix-7 FPGA, which is suitable for resource-constrained devices. Additionally, RISC-HD improves the total clock cycle by$7.48\times $compared to the RI5CY core and$6.17\times $compared to ARM Cortex-M4 in the ISOLET data set. Moreover, RISC-HD achieves$7.22\times $energy efficiency compared to the RI5CY core. Farhad Taheri, Siavash Bayat Sarmadi, Shahriar Hadayeghparast |
IEEE Internet Things J. | 2 |
| 2022 | Fast Supersingular Isogeny Diffie-Hellman and Key Encapsulation Using a Customized Pipelined Montgomery MultiplierabstractWe present a pipelined Montgomery multiplier tailored for SIKE primes. The latency of this multiplier is far shorter than that of the previous work while its frequency competes with the highest-rated ones. The implementation results on a Virtex-7 FPGA show that this multiplier improves the time, the area-time product (AT), and the throughput of computing modular multiplication by at least 2.30, 1.60, and 1.36 times over SIKE primes respectively. We have also developed a CPU-like architecture to perform SIDH and SIKE using several instances of our modular multiplier. Using four multipliers on a Virtex-7 FPGA, the encapsulation and the decapsulation of SIKE can be performed at least 1.45 times faster while improving the AT by at least 1.35 times over all SIKE primes. We have also evaluated our implementation on two other FPGAs. The implementation on Artix-7 improves the time and the AT of performing these two steps of SIKE by at least 1.90 and 1.80 times, respectively. On Kintex UltraScale+, these improvement factors are 2.05 and 2.08, respectively. On this device, these two steps take 3.11, 3.52, 4.66, and 6.59 milliseconds on$p_{434}$,$p_{503}$,$p_{610}$, and$p_{751}$, respectively. Mohammad Hossein Farzam, Siavash Bayat Sarmadi, Hatameh Mosanaei-Boorani, Armin Alivand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Efficient Hardware Implementations of Legendre Symbol Suitable for MPC ApplicationsabstractMulti-party computation (MPC) allows each peer to take part in the execution of a common function with their private share of data without the need to expose it to other participants. The Legendre symbol is a pseudo-random function (PRF) that is suitable for MPC protocols due to their efficient evaluation process compared to other symmetric primitives. Recently, Legendre-based PRFs have also been employed in the construction of a post-quantum signature scheme, namely LegRoast. In this paper, we propose, to the best of our knowledge, the first hardware implementations for the Legendre symbol by three approaches: 1) low-area, 2) high-speed, and 3) high-frequency. The high-speed architecture outperforms state-of-the-art software implementations, which run on Intel’s Core-i5. Our evaluation results on FPGA show that this architecture reduces the Legendre calculation time by$2.56\times $compared to software implementations on Core-i5. On the other hand, the low-area architecture consumes only 5489 slices on the Artix-7 FPGA and is suitable for resource-constrained devices. Moreover, our ASIC implementation results indicate that the low-area architecture consumes 97.56K gates to implement and requires$4.01~mW$to operate on 50 MHz. The high-frequency architecture increases the frequency by$1.72\times $over the high-speed architecture and achieves 200 MHz frequency on FPGA. Farhad Taheri, Siavash Bayat Sarmadi, Shahriar Ebrahimi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Lightweight Fuzzy Extractor Based on LPN for Device and Biometric Authentication in IoTabstractUser and device biometrics are proven to be a reliable source for authentication, especially for the Internet-of-Things (IoT) applications. One of the methods to employ biometric data in authentication are fuzzy extractors (FE) that can extract cryptographically secure and reproducible keys from noisy biometric sources with some entropy loss. It has been shown that one can reliably build an FE based on the learning parity with noise (LPN) problem with higher error-tolerance than previous FE schemes. However, the only available LPN-based FE implementation suffers from extreme resource demands that are not practical for IoT devices. This article proposes a lightweight hardware/software (HW/SW) co-design for implementing LPN-based FE. We provide different optimizations on architecture to decrease the resource requirements of the scheme. The proposed architecture is resistant against simple side-channel analysis and improves area and area-time product (AT) by more than 89% and 83%, respectively, compared to previous work. Our experimental results indicate that the proposed architecture can be implemented on off-the-shelf resource-constrained SoC-FPGA boards from different vendors such as Xilinx, Digilent, and Trenz. Moreover, we provide the first implementation results of LPN-based FE on an application-specific integrated circuit (ASIC) platform using HW/SW co-design. Shahriar Ebrahimi, Siavash Bayat Sarmadi |
IEEE Internet Things J. | 2 |
| 2021 | PLCDefender: Improving Remote Attestation Techniques for PLCs Using Physical ModelabstractIn order to guarantee the security of industrial control system (ICS) processes, the proper functioning of the programmable logic controllers (PLCs) must be ensured. In particular, cyberattacks can manipulate the PLC control logic program and cause terrible damage that jeopardize people's life when bringing the state of the critical system into an unreliable state. Unfortunately, no remote attestation technique has yet been proposed that can validate the PLC control logic program using a physics-based model that demonstrates device behavior. In this article, we propose PLCDefender, a mitigation method that combines hybrid remote attestation technique with a physics-based model to preserve the control behavior integrity of ICS. We implemented PLCDefender and evaluated its effectiveness against a wide range of attacks on a secure water treatment facility. As our evaluation shows, we can model PLC physical behavior with accuracy as high as 98%. The evaluation results show that by determining the different threshold values, PLCDefender can accurately detect a wide range of attack scenarios on PLCs. Mohsen Salehi, Siavash Bayat Sarmadi |
IEEE Internet Things J. | 2 |
| 2021 | Hardware Architecture for Supersingular Isogeny Diffie-Hellman and Key Encapsulation Using a Fast Montgomery MultiplierabstractPublic key cryptography lies among the most important bases of security protocols. The classic instances of these cryptosystems are no longer secure when a large-scale quantum computer emerges. These cryptosystems must be replaced by post-quantum ones, such as isogeny-based cryptographic schemes. Supersingular isogeny Diffie-Hellman (SIDH) and key encapsulation (SIKE) are two of the most important such schemes. To improve the performance of these protocols, we have designed several modular multipliers. These multipliers have been implemented for all the prime fields used in SIKE round 3, on a Virtex-7 FPGA, showing a time and area-time product improvement of up to 60.1% and 64.5%, respectively. These multipliers are also suitable for applications such as RSA, as shown by implementations for 512-bit, 1024-bit, and 2048-bit generic moduli on a Virtex-7 FPGA. Our fastest multiplier has been used in the implementation of SIDH and SIKE round 3. Employing six instances of this multiplier, SIDH completes after 7.33, 8.93, 13.39, and 18.67 milliseconds and the encapsulation and the decapsulation of SIKE is performed in 7.13, 8.68, 13.08, and 18.16 milliseconds over p434, p503, p610, p751, respectively, which yields a least improvement factor of 1.23. Mohammad Hossein Farzam, Siavash Bayat Sarmadi, Hatameh Mosanaei-Boorani, Armin Alivand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | Lightweight and Fault-Resilient Implementations of Binary Ring-LWE for IoT DevicesabstractWhile the Internet of Things (IoT) shapes the future of the Internet, communications among nodes must be secured by employing cryptographic schemes such as public-key encryption (PKE). However, classic PKE schemes, such as RSA and elliptic curve cryptography (ECC) suffer from both high complexity and vulnerability to quantum attacks. During the past decade, post-quantum schemes based on the learning with errors (LWEs) problem have gained high attention due to the lower complexity among PKE schemes. In addition to resistance against theoretical (quantum and classic) attacks, every practical implementation of any cryptosystem must also be evaluated against different side-channel attacks such as power analysis or fault injection ones. In this article, we analyze the vulnerability of binary ring learning with error (Ring-LWE) scheme regarding (first-order) fault attacks, such as randomization, zeroing, and skipping faults. We show that previous implementations can be easily broken by employing such fault attacks. Moreover, we propose fault-resilient software implementations of binary Ring-LWE on 8- and 32-b lightweight microcontrollers, namely, AVR ATxmega128A1 and ARM Cortex-M0 that are ideal for IoT devices. Furthermore, we formally prove the resilience of the proposed implementations against different fault attacks. To the best of our knowledge, this article is the first one to propose fault-resilient binary Ring-LWE implementations on resource-constrained microcontrollers. Our implementations on AVR ATxmega128A1 require only 80 and 120 ms for encryption and decryption, respectively. Shahriar Ebrahimi, Siavash Bayat Sarmadi |
IEEE Internet Things J. | 2 |
| 2019 | A Unified Approach to Detect and Distinguish Hardware Trojans and Faults in SRAM-based FPGAs
Omid Ranjbar, Siavash Bayat Sarmadi, Fatemeh Pooyan, Hossein Asadi 0001 |
J. Electron. Test. | 2 |
| 2019 | Post-Quantum Cryptoprocessors Optimized for Edge and Resource-Constrained Devices in IoTabstractBy exponential increase in applications of the Internet of Things (IoT), such as smart ecosystems or e-health, more security threats have been introduced. In order to resist known attacks for IoT networks, multiple security protocols must be established among nodes. Thus, IoT devices are required to execute various cryptographic operations, such as public key encryption/decryption. However, classic public key cryptosystems, such as Rivest-Shammir-Adlemon and elliptic curve cryptography are computationally more complex to be efficiently implemented on IoT devices and are vulnerable regarding quantum attacks. Therefore, after complete development of quantum computing, these cryptosystems will not be secure and practical. In this paper, we propose InvRBLWE, an optimized variant for binary learning with errors over the ring (Ring-LWE) scheme that is proven to be secure against quantum attacks and is highly efficient for hardware implementations. We propose two architectures for InvRBLWE: 1) a high-speed architecture targeting edge and powerful IoT devices and 2) an ultralightweight architecture, which can be implemented on resource-constrained nodes in IoT. The proposed architectures are scalable regarding security levels and we provide experimental results for two versions of the InvRBLWE scheme providing 84 and 190 bits of classic security. Our implementation results on field programmable gate array dominate the best of the classic and post-quantum previous implementations. Moreover, our two different application specific integrated circuit (ASIC) implementations show improvement in terms of speed, area, power, and/or energy. To the best of our knowledge, we are the first to implement learning with error-based cryptosystems on ASIC platform. Shahriar Ebrahimi, Siavash Bayat Sarmadi, Hatameh Mosanaei-Boorani |
IEEE Internet Things J. | 2 |
| 2019 | Toward On-chip Network Security Using Runtime Isolation MappingabstractMany-cores execute a large number of diverse applications concurrently. Inter-application interference can lead to a security threat as timing channel attack in the on-chip network. A non-interference communication in the shared on-chip network is a dominant necessity for secure many-core platforms to leverage the concepts of the cloud and embedded system-on-chip. The current non-interference techniques are limited to static scheduling and need router modification at micro-architecture level. Mapping of applications can effectively determine the interference among applications in on-chip network. In this work, we explore non-interference approaches through run-time mapping at software and application level. We map the same group of applications in isolated domain(s) to meet non-interference flows. Through run-time mapping, we can maximize utilization of the system without leaking information. The proposed run-time mapping policy requires no router modification in contrast to the best known competing schemes, and the performance degradation is, on average, 16% compared to the state-of-the-art baselines. Mohammad Sadegh Sadeghi, Siavash Bayat Sarmadi, Shaahin Hessabi |
ACM Trans. Archit. Code Optim. | 2 |
| 2018 | Reliable hardware architectures for efficient secure hash functions ECHO and fugueabstractIn cryptographic engineering, extensive attention has been devoted to ameliorating the performance and security of the algorithms within. Nonetheless, in the state-of-the-art, the approaches for increasing the reliability of the efficient hash functions ECHO and Fugue have not been presented to date. We propose efficient fault detection schemes by presenting closed formulations for the predicted signatures of different transformations in these algorithms. These signatures are derived to achieve low overhead for the specific transformations and can be tailored to include byte/word-wide predicted signatures. Through simulations, we show that the proposed fault detection schemes are highly-capable of detecting natural hardware failures and are capable of deteriorating the effectiveness of malicious fault attacks. The proposed reliable hardware architectures are implemented on the application-specific integrated circuit (ASIC) platform using a 65-nm standard technology to benchmark their hardware and timing characteristics. The results of our simulations and implementations show very high error coverage with acceptable overhead for the proposed schemes. Mehran Mozaffari Kermani, Reza Azarderakhsh, Siavash Bayat Sarmadi |
CF | 3 |
| 2015 | On Constrained Implementation of Lattice-Based Cryptographic Primitives and Schemes on Smart CardsabstractMost lattice-based cryptographic schemes with a security proof suffer from large key sizes and heavy computations. This is also true for the simpler case of authentication protocols that are used on smart cards as a very-constrained computing environment. Recent progress on ideal lattices has significantly improved the efficiency and made it possible to implement practical lattice-based cryptography on constrained devices. However, to the best of our knowledge, no previous attempts have been made to implement lattice-based schemes on smart cards. In this article, we provide the results of our implementation of several state-of-the-art lattice-based authentication protocols on smart cards and a microcontroller widely used in smart cards. Our results show that only a few of the proposed lattice-based authentication protocols can be implemented using limited resources of such constrained devices; however, cutting-edge ones are suitably efficient to be used practically on smart cards. Moreover, we have implemented fast Fourier transform (FFT) and discrete Gaussian sampling with different typical parameter sets, as well as versatile lattice-based public-key encryptions. These results have noticeable points that help to design or optimize lattice-based schemes for constrained devices. Ahmad Boorghany, Siavash Bayat Sarmadi, Rasool Jalili |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Systolic Gaussian Normal Basis Multiplier Architectures Suitable for High-Performance ApplicationsabstractNormal basis multiplication in finite fields is vastly utilized in different applications, including error control coding and the like due to its advantageous characteristics and the fact that squaring of elements can be obtained without hardware complexity. In this brief, we present decomposition algorithms to develop novel systolic structures for digit-level Gaussian normal basis multiplication over GF(2m). The proposed architectures are suitable for high-performance applications, which require fast computations in finite fields with high throughputs. We also present the results of our application-specific integrated circuit synthesis using a 65-nm standard-cell library to benchmark the effectiveness of the proposed systolic architectures. The presented architectures for multiplication can result in more efficient and high-performance VLSI systems. Reza Azarderakhsh, Mehran Mozaffari Kermani, Siavash Bayat Sarmadi, Chiou-Yng Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Efficient and Concurrent Reliable Realization of the Secure Cryptographic SHA-3 AlgorithmabstractThe secure hash algorithm (SHA)-3 has been selected in 2012 and will be used to provide security to any application which requires hashing, pseudo-random number generation, and integrity checking. This algorithm has been selected based on various benchmarks such as security, performance, and complexity. In this paper, in order to provide reliable architectures for this algorithm, an efficient concurrent error detection scheme for the selected SHA-3 algorithm, i.e., Keccak, is proposed. To the best of our knowledge, effective countermeasures for potential reliability issues in the hardware implementations of this algorithm have not been presented to date. In proposing the error detection approach, our aim is to have acceptable complexity and performance overheads while maintaining high error coverage. In this regard, we present a low-complexity recomputing with rotated operands-based scheme which is a step-forward toward reducing the hardware overhead of the proposed error detection approach. Moreover, we perform injection-based fault simulations and show that the error coverage of close to 100% is derived. Furthermore, we have designed the proposed scheme and through ASIC analysis, it is shown that acceptable complexity and performance overheads are reached. By utilizing the proposed high-performance concurrent error detection scheme, more reliable and robust hardware implementations for the newly-standardized SHA-3 are realized. Siavash Bayat Sarmadi, Mehran Mozaffari Kermani, Arash Reyhani-Masoleh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Reliable Concurrent Error Detection Architectures for Extended Euclidean-Based Division Over GF(2m)abstractThe extended Euclidean algorithm (EEA) is an important scheme for performing the division operation in finite fields. Many sensitive and security-constrained applications such as those using the elliptic curve cryptography for establishing key agreement schemes, augmented encryption approaches, and digital signature algorithms utilize this operation in their structures. Although much study is performed to realize the EEA in hardware efficiently, research on its reliable implementations needs to be done to achieve fault-immune reliable structures. In this regard, this paper presents a new concurrent error detection (CED) scheme to provide reliability for the aforementioned sensitive and constrained applications. Our proposed CED architecture is a step forward toward more reliable architectures for the EEA algorithm architectures. Through simulations and based on the number of parity bits used, the error detection capability of our CED architecture is derived to be 100% for single-bit errors and close to 99% for the experimented multiple-bit errors. In addition, we present the performance degradations of the proposed approach, leading to low-cost and reliable EEA architectures. The proposed reliable architectures are also suitable for constrained and fault-sensitive embedded applications utilizing the EEA hardware implementations. Mehran Mozaffari Kermani, Reza Azarderakhsh, Chiou-Yng Lee, Siavash Bayat Sarmadi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | Concurrent Error Detection in Finite-Field Arithmetic Operations Using Pipelined and Systolic ArchitecturesabstractIn this work, we consider detection of errors in polynomial, dual, and normal bases arithmetic operations. Error detection is performed by recomputing with the shifted operand method, while the operation unit is in use. This scheme is efficient for pipelined architectures, particularly systolic arrays. Additionally, one semisystolic multiplier for each of the polynomial, dual, type I, and type II optimal normal bases is presented. The results show that for having better or similar space and time overheads compared to a number of related previous work, the multipliers have generally a higher error-detection capability, e.g., the error-detection capability of the RESO-based scheme for single and multiple stuck-at faults in a polynomial basis multiplier is 100 percent. Finally, we also comment on how RESO can be used for concurrent error correction to deal with transient faults. Siavash Bayat Sarmadi, M. Anwar Hasan |
IEEE Trans. Computers | 1 |
| 2007 | Run-Time Error Detection in Polynomial Basis Multiplication Using Linear CodesabstractIn this article we consider detection of errors in polynomial basis multipliers, which have applications in channel coding, VLSI testing, and cryptography. Error detection is performed by applying a class of linear codes while the multiplier is in use. In this article, two error detection schemes are presented. Results show that the probability of error detection of our single-input encoding (SIE) scheme using eight redundant bits is approximately 0.996. Additionally, the time and area overheads of the schemes for our bit-serial implementations are in a reasonable range, e.g., for the SIE scheme with eight redundant bits, the area overhead is 39.71% and the time overhead has been observed to be negligible. Siavash Bayat Sarmadi, M. Anwar Hasan |
ASAP | 1 |
| 2007 | Detecting errors in a polynomial basis multiplier using multiple parity bits for both inputsabstractThis paper investigates the concurrent detection of multiple-bit errors in polynomial basis (PB) multipliers over binary extension fields. To this end, multiple parity bits are considered for both inputs of the multiplier. For the multiplier architecture considered here, the two inputs go through considerably different sets of circuits and this allows us to use different number of parity bits with the inputs. In a bit-parallel implementation of a GF(2163) PB multiplier with eight parity bits for the first input and three parity bits for the second input, the area overhead and the probability of error detection are approximately 55.59% and 0.997, respectively. Additionally, the average time overhead of the scheme implemented in a bit-parallel fashion is approximately 25%. Siavash Bayat Sarmadi, M. Anwar Hasan |
ICCD | 1 |
| 2007 | On Concurrent Detection of Errors in Polynomial Basis MultiplicationabstractThe detection of errors in arithmetic operations is an important issue. This paper discusses the detection of multiple-bit errors due to faults in bit-serial and bit-parallel polynomial basis (PB) multipliers over binary extension fields. Our approach is based on multiple parity bits. Experimental results presented here show that due to an increase in the number of parity bits, the area overhead tends to increase linearly, but the probability of error detection approaches unity fairly quickly, e.g., for eight parity bits. In bit-serial implementation of a GF(2163) PB multiplier using eight parity bits, the area overhead and the probability of error detection are 10.29% and 0.996, respectively. This is achieved without any increase in the computation time of the GF(2163) PB multiplier Siavash Bayat Sarmadi, M. Anwar Hasan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | A Hybrid Fault Injection Approach Based on Simulation and Emulation Co-operationabstractThis paper presents a new fault injection approach, which is based on a co-operation between a simulator and an emulator. This hybrid approach utilizes the advantages of both simulation-based fault injection as well as physical fault injection to provide a good controllability, observability and also a high speed in the fault injection experiments. To do this, parts of a circuit are simulated while the rest parts of the circuit are emulated. A fault injection tool called FITSEC (Fault Injection Tool based on Simulation and Emulation Cooperation) is developed, which supports the entire process of a system design. This is based on both Verilog and VHDL languages and can be used to inject faults at different levels of abstraction. The experimental results show that this approach can significantly reduce the time needed for executing fault injection campaigns. Alireza Ejlali, Seyed Ghassem Miremadi, Hamid R. Zarandi, Hossein Asadi 0001, Siavash Bayat Sarmadi |
DSN | 5 |
| 2002 | Fast Prototyping with Co-operation of Simulation and Emulation
Siavash Bayat Sarmadi, Seyed Ghassem Miremadi, Hossein Asadi 0001, Alireza Ejlali |
FPL | 1 |
| 2002 | Speedup analysis in simulation-emulation co-operationabstractThis paper presents an analytical approach to estimate the speedup in a simulation-emulation cooperation environment. The speedup of this approach as compared with the speedup of a pure simulation is analyzed. Also, an analysis of the speedup is given when different types of application instructions are utilized. The analysis is based on using both Verilog and VHDL. The results show that when only the simulation part of the simulation-emulation co-operation is used, the speedup is higher, than when the pure simulation is used. The total speedup is also depended on the type of application instructions and the communication cycle time between the simulator and the emulator. Seyed Ghassem Miremadi, Siavash Bayat Sarmadi, Hossein Asadi 0001 |
FPT | 2 |