EDBT 2026 Demo / reviewers in the wild / expert
Zhen Gao 0005
dblp:71/1107-5
· DBLP profile ↗
28ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0001-9887-1418ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 9 first-author · 17 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An FPGA-Based Design of Reliable Quantized AdderNet Under Fault Injection
Rong Tian, Huiyang Jin, Yanmao Qi, Zhen Gao 0005 |
ISCAS | 4 |
| 2026 | Fault Modeling and Countermeasures for DRAM-Targeted Electromagnetic Fault Injection
Qiang Liu 0011, Longtao Guo, Xianzhao Xia, Zhen Gao 0005 |
J. Electron. Test. | 4 |
| 2025 | Cloud-edge-end integrated Artificial intelligence based on ensemble learning
Zhen Gao 0005, Daning Su, Chenyang Wang 0001, Cheng Zhang 0019, Xiaofei Wang 0001, Tarik Taleb |
Comput. Commun. | 1 |
| 2025 | Perturbation-based error detection and correction (PBEDC) in dependable large-scale machine learning systems
Ziheng Wang 0005, Pedro Reviriego, Shanshan Liu 0001, Farzad Niknia, Xiaochen Tang, Zhen Gao 0005, Fabrizio Lombardi |
Future Gener. Comput. Syst. | 6 |
| 2025 | Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)abstractStochastic computing (SC) has emerged as an efficient low-power alternative for deploying neural networks (NNs) in resource-limited scenarios, such as the Internet of Things (IoT). By encoding values as serial bitstreams, SC significantly reduces energy dissipation compared to conventional floating-point (FP) designs; however, further improvement of layer-wise mixed-precision implementation for SC remains unexplored. This paper introduces Adjustable Sequence Length (ASL), a novel scheme that applies mixedprecision concepts specifically to SC NNs. By introducing an operator-norm – based theoretical model, this paper shows that truncation noise can cumulatively propagate through the layers by the estimated amplification factors. An extended sensitivity analysis is presented, using Random Forest (RF) regression to evaluate multi-layer truncation effects and validate the alignment of theoretical predictions with practical network behaviors. To accommodate different application scenarios, this paper proposes two truncation strategies (coarse-grained and fine-grained), which apply diverse sequence length configurations at each layer. Evaluations on a pipelined SC MLP synthesized at 32 nm demonstrate that ASL can reduce energy and latency overheads by up to over 60% with negligible accuracy loss. It confirms the feasibility of the ASL scheme for IoT applications and highlights the distinct advantages of mixed-precision truncation in SC designs. Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Zhen Gao 0005, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Internet Things J. | 4 |
| 2025 | Concurrent Linguistic Error Detection (CLED): A New Methodology for Error Detection in Large Language ModelsabstractThe utilization of Large Language Models (LLMs) requires dependable operation in the presence of errors in the hardware (caused by for example radiation) as this has become a pressing concern. At the same time, the scale and complexity of LLMs limit the overhead that can be added to detect errors. Therefore, there is a need for low-cost error detection schemes. Concurrent Error Detection (CED) uses the properties of a system to detect errors, so it is an appealing approach. In this paper, we present a new methodology and scheme for error detection in LLMs: Concurrent Linguistic Error Detection (CLED). Its main principle is that the output of LLMs should be valid and generate coherent text; therefore, when the text is not valid or differs significantly from the normal text, it is likely that there is an error. Hence, errors can potentially be detected by checking the linguistic features of the text generated by LLMs. This has the following main advantages: 1) low overhead as the checks are simple and 2) general applicability, so regardless of the LLM implementation details because the text correctness is not related to the LLM algorithms or implementations. The proposed CLED has been evaluated on two LLMs: T5 and OPUS-MT. The results show that with a 1% overhead, CLED can detect more than 87% of the errors, making it suitable to improve LLM dependability at low cost. Javier Conde, Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2025 | Dependability of the K Minimum Values Sketch: Protection and Comparative AnalysisabstractA basic operation in big data analysis is to find the cardinality estimate; to estimate the cardinality at high speed and with a low memory requirement, data sketches that provide approximate estimates, are usually used. The K Minimum Value (KMV) sketch is one of the most popular options; however, soft errors on memories in KMV may substantially degrade performance. This paper is the first to consider the impact of soft errors on the KMV sketch and to compare it with HyperLogLog (HLL), another widely used sketch for cardinality estimate. Initially, the operation of KMV in the presence of soft errors (so its dependability) in the memory is studied by a theoretical analysis and simulation by error injection. The evaluation results show that errors during the construction phase of KMV may cause large deviations in the estimate results. Subsequently, based on the algorithmic features of the KMV sketch, two protection schemes are proposed. The first scheme is based on using a single parity check (SPC) to detect errors and reduce their impact on the cardinality estimate; the second scheme is based on the incremental property of the memory list in KMV. The presented evaluation shows that both schemes can dramatically improve the performance of KMV, and the SPC scheme performs better even though it requires more memory footprint and overheads in the checking operation. Finally, it is shown that soft errors on the unprotected KMV produce larger worst-case errors than in HLL, but the average impact of errors is lower; also, the protected KMV using the proposed schemes are more dependable than HLL with existing protection techniques. Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2025 | Detect and Replace: Efficient Soft Error Protection of FPGA-Based CNN AcceleratorsabstractConvolutional neural networks (CNNs) are widely used in computer vision and natural language processing. Field-programmable gate arrays (FPGAs) are a popular accelerator for CNNs. However, FPGAs are prone to suffer soft errors, so the reliability of FPGA-based CNNs becomes a key problem when used in safety-critical applications. The convolution module based on a processing element (PE) array is the most complex part of the accelerator, so it is the key to efficient protection. Coding-based schemes have been proposed for efficient protection of the convolution module, where the processing of the PE array is modeled as parallel matrix-vector multiplications (MVMs), and every wrong output would be concurrently detected and corrected. However, these schemes cannot deal with errors in the configuration memory that affects many intermediate results. In this article, a protection scheme is proposed based on faulty PE detection and replace (DR) to deal with such configuration memory errors. The DR scheme is implemented on a CNN accelerator based on Xilinx Zynq 7000 SoC, and fault injection (FI) experiments are performed to evaluate the performance of the proposed DR scheme. The results show that it can effectively mitigate the effect of soft errors in the configuration memory with an overhead of about 1.3 times complexity and 1.4 times power consumption relative to those of the unprotected PE array. Compared with the advanced checksum-of-checksum (CoC) scheme, the DR scheme decreases power consumption by up to 30%. Zhen Gao 0005, Yanmao Qi, Jinchang Shi, Qiang Liu 0011, Guangjun Ge, Yu Wang 0002, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Reducing the Energy Dissipation of Large Language Models (LLMs) with Approximate MemoriesabstractLarge language models (LLMs) have shown impressive performance in a wide range of tasks such as answering questions or summarizing text. However, running LLMs on edge devices is challenging as they require large amounts of energy due to their memory and computation needs. In LLMs most of the memory is needed to store the model parameters which number keeps increasing from one LLM generation to the next. In the last several years, significant efforts have been made to compress and prune parameters, but this is not enough to reduce their memory needs as the number of parameters grows exponentially. In this work, to reduce energy dissipation, rather than trying to reduce the amount of memory used by LLMs, we study the use of approximate memories to store the LLM parameters. Approximate memories can significantly reduce the energy dissipation at the cost of introducing errors in some of the memory bits. Therefore, the impact of errors on LLMs must be understood. To that end, we have performed error injection on different compressed versions of a classic LLM: Bidirectional Encoder Representations from Transformers (BERT). The results show that in some cases compressed BERTs operate reliably at high bit error rates. This makes possible the use of approximate memories with a negligible impact on the LLM performance and a significant reduction in energy dissipation. Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
ISCAS | 1 |
| 2024 | Cognitive Data Fusing for Internet of Things Based on Ensemble Learning and Federated LearningabstractBig data produced by Internet of Things (IoT) devices is the key drive for Prognostic and Health Management (PHM) for industrial equipment or systems. However, data are usually distributed stored in many scenarios due to security and privacy problems. Federated learning (FL) is an effective solution to fuse the data for intelligent decision. But FL faces risk of Denial of Service (DoS) attack or Single Point of Failure (SPOF) problem during training and service phases, and exchange of model parameters poses heavy network traffic between clients. Ensemble learning (EL) is widely used to boost task performance by combining diverse base learners, and it has shown promise in improving distributed intelligent services. Since a decision is collaboratively made by multiple clients in EL in a distributed fashion, DoS and SPOF problem can be inherently avoided, and the deployment cost is much lower than FL. Based on these good properties, we proposed to combine FL and EL for distributed IoT data fusion with a cognitive approach. First, we propose to construct effective EL by generating diverse base models with advanced pruning method, and compare the performance of FL and EL based distributed data fusion. Then a hierarchical combination of FL and EL is proposed based on the cognition of cost and performance at each level for efficient deployment of distributed IoT data fusion. Experiment results show that EL based scheme can achieve close performance to FL based scheme for small number of clients with some data sharing, and the cognitive hierarchical combination of FL and EL can achieve a good tradeoff between task performance and network traffic for large scale distributed IoT data fusion. Zhen Gao 0005 |
IEEE Internet Things J. | 1 |
| 2024 | Verkle-Accumulator-Based Stateless Transaction Validation (VA-STV) Scheme for the Blockchain-Based IoT NetworkabstractThe blockchain-based Internet of Things (IoT) has served widely across various industries for authentication, cooperation, and data sharing but faced the severe challenge of storage scalability. The storage burden gets worse for IoT devices with limited resources. The state data is essential for efficient transaction issuance and validation. This article proposes the Verkle accumulator-based stateless transaction validation (VA-STV) scheme for permissionless blockchains to decrease the storage burden of the state data on each node with the acceptable overhead of computation and communication. In the scheme, the current state is summarized as the commitment maintained in the latest block header, and one witness is generated for each token to guarantee its validity. State transitions are realized by updating the commitment and witnesses so that no state is stored on nodes acting as validators and miners. Only the nodes acting as traders should maintain the tokens controlled by themselves and the witnesses locally. The VA-STV is based on the Verkle accumulator (VA), which is a combination of the Verkle tree (VT) and the KZG polynomial commitment scheme. Simulation results show that the VA-STV provides a smaller witness size ($0.6\times $–$0.74\times $) and faster commitment generation ($6\times $–$14\times$) than the existing stateless schemes in the same settings, which indicates the advantages of VA-STV in succinctness and efficiency. Besides, a tradeoff between the communication and computation requirements can be achieved by adjusting the branching factor, which improves the adaptability of the proposed scheme for different IoT scenarios. Zhaohui Guo, Zhen Gao 0005, Qiang Liu 0011, Lei Liu 0031, Mianxiong Dong, Ning Zhang 0007, Mohammed Atiquzzaman |
IEEE Internet Things J. | 2 |
| 2024 | Game-Based Low Complexity and Near Optimal Task Offloading for Mobile Blockchain SystemsabstractThe Internet of Things (IoT) finds applications across diverse fields but grapples with privacy and security concerns. Blockchain offers a remedy by instilling trust among IoT devices. The development of blockchain in IoT encounters hurdles due to its resource-intensive computation processing, notably in PoW-based systems. Cloud and edge computing can facilitate the application of blockchain in this environment, and the IoT users who want to mine in blockchain need to pay the computation resource rent to the Cloud Computing Service Provider (CCSP) for offloading the mining workload. In this scenario, these IoT miners can form groups to trade with CCSP to maximize their utility. In this paper, a mixed model of the Stackelberg game and coalition formation game is embraced to address the grouping and pricing issues between IoT miners and CCSP. In particular, the Stackelberg game is utilized to handle the pricing problem, and the coalition formation game is employed to tackle the best group partition problem. Moreover, a coalition formation algorithm is proposed to obtain a nearoptimal solution with very low complexity. Simulation results show that our proposed algorithm can obtain a performance that is very near to the exhaustive search method, outperforms other existing schemes, and requires only a small computation overhead. Jing Li 0006, Zhen Gao 0005, Zhu Han 0001, Chao Qiu, Xiaofei Wang 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2024 | Concurrent Classifier Error Detection (CCED) in Large Scale Machine Learning SystemsabstractThe complexity of machine learning (ML) systems increases each year. As these systems are widely utilized, ensuring their reliable operation is becoming a design requirement. Traditional error detection mechanisms introduce circuit or time redundancy that significantly impacts system performance. An alternative is the use of concurrent error detection (CED) schemes that operate in parallel with the system and exploit their properties to detect errors. CED is attractive for large ML systems because it can potentially reduce the cost of error detection. In this article, we introduce concurrent classifier error detection (CCED), a scheme to implement CED in ML systems using a concurrent ML classifier to detect errors. CCED identifies a set of check signals in the main ML system and feed them to the concurrent ML classifier that is trained to detect errors. The proposed CCED scheme has been implemented and evaluated on two widely used large-scale ML models: Contrastive language-image pretraining (CLIP) used for image classification and bidirectional encoder representations from transformers (BERT) used for natural language applications. The results show that more than 95% of the errors are detected when using a simple Random Forest classifier that is orders of magnitude simpler than CLIP or BERT. Pedro Reviriego, Ziheng Wang 0005, Zhen Gao 0005, Farzad Niknia, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Reliab. | 4 |
| 2023 | Efficient Protection of FPGA Implemented LDPC Decoders Against Single Event Upsets (SEUs) on Configuration MemoriesabstractLow Density Parity Check (LDPC) codes are used in 5G systems for traffic channels due to their excellent error correction capability for long sequences, and the Min-Sum algorithm is widely applied in practical implementations of LDPC decoders due to its low complexity. If the decoder is implemented on a SRAM-based field-programmable gate array (SRAM-FPGA), the radiation-induced single-event upsets (SEUs) can affect the operation of the LDPC decoder by corrupting the configuration memory, which can change the circuit functionality and will not be corrected unless the FPGA is reconfigured. Therefore, protection of LDPC decoders with low overhead is an important problem, especially for resource-limited on-board space systems. In this paper, an efficient Duplicate With Comparison (DWC) protection scheme is proposed based on the different distribution of the parity check sum of the LDPC decoder in the error-free case and the faulty case. In particular, the check sum accumulation number and threshold are optimized to achieve high detection probability with short delay. FPGA based implementation and hardware fault injection experiments are conducted to evaluate the performance of the proposed schemes. Experimental results show that, the effect of SEUs on the LDPC decoder can be completely eliminated by the proposed scheme with 2 times computational overhead and 1.69 times power consumption overhead compared to the unprotected decoder. Zhen Gao 0005, Yinghao Cheng, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A Methodology for the Design of Fault Tolerant Parallel Digital Channelizers on SRAM-FPGAsabstractDigital channelizers (DCs) based on the Discrete Fourier Transform (DFT) and polyphase filter banks are widely used in on-board processing (OBP) platforms to extract narrowband sub-channels from a wideband signal efficiently. In high-capacity communication satellite platforms there are always multiple DCs extracting narrowband signals from multiple wideband signals in parallel. Field-programmable gate arrays (FPGAs) are a popular option for the implementation of DCs due to their parallel computing capabilities and good re-configurability, but FPGAs suffer single-event upsets (SEUs) on the space platform. This paper focuses on the efficient protection of parallel DCs with enhanced coding techniques. We first prove that a linear relationship between parallel DCs can be introduced and maintained among the multiple outputs. However, traditional coding schemes cannot be directly applied for the detection of faulty DCs due to the quantization noise introduced by fixed point implementations. To address this issue, we propose an enhanced coding scheme by averaging in the space and time domains to minimize the effect of quantization noise, introducing thresholds and a majority voter to further improve the detection probability. Both theoretical analysis and fault injection experiments prove the effectiveness of the proposed protection scheme. Experimental results show that all the SEUs that cause an SNR lower than 20dB can be detected and recovered, and the resource overheads are about 1.6 times and 1.3 times of that of the unprotected DCs for systems with 8 DCs and 16 DCs, respectively. Zhen Gao 0005, Jiajun Xiao, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Resource Management and Pricing for Cloud Computing Based Mobile Blockchain With PoolingabstractIn a public blockchain system applying Proof of Work (PoW), the participants need to compete with their computing resources for reward, which is challenging for resource-limited devices. Mobile blockchain is proposed to facilitate the application of blockchain for mobile service, in which the lightweight devices can participate mining by renting resources from the Cloud Computing Service Provider (CCSP), but CCSP usually does not have the information about the demand preference of users. In this article, a contract model is adopted to address the cloud computing resource allocation and pricing problem in the mobile blockchain. In particular, an adverse selection contract solution is proposed to overcome the information asymmetry problem, and resource pooling is introduced to improve the stability of users’ rewards. Simulation results show that the information asymmetry problem is well overcome by adverse selection contract so that CCSP can obtain more utility than linear pricing contracts. Furthermore, the resource pooling could effectively improve the users’ and CCSP's utilities. When the size of the mining pool is large enough, it can achieve an improvement effect of more than 10 times. The effect of pool size and user type distribution on CCSP's utility is also studied. Jing Li 0006, Zhen Gao 0005, Zhu Han 0001, Chao Qiu, Xiaofei Wang 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | World State Attack to Blockchain Based IoV and Efficient Protection With Hybrid RSUs ArchitectureabstractBlockchain technology is developing rapidly and has been widely applied in the field of Internet of Vehicles (IoV) to solve trust and security problems. However, due to the high security requirements in IoV scenarios, the security threats of blockchain itself become a big challenge for its applications in IoV. As the largest distributed platform supporting smart contract, Ethereum becomes one of the popular blockchain platforms that has been applied in IoV applications. In Ethereum, the local world state (stored on Road Side Units (RSUs) in IoV) is applied to facilitate account query and transaction verification. However, previous works showed that the local database can be easily tampered, so attackers may issue invalid transactions based on the modified world state, which is not acceptable for IoV applications. In this paper, the success probability and expected time for such an attack are first analyzed theoretically, including the effect of portion of tampered RSUs and the number of required confirmation blocks. Then experiment evaluation verifies the correctness of the theoretical analysis and shows that the attack would succeed with a higher probability within a shorter time when the local database on more RSUs are attacked. On the contrary, increasing of confirmation blocks can effectively reduce the success probability of a single attack and extend the confirmation time of the invalid transaction. Finally, efficient attack detection and recovery methods are proposed based on a novel hierarchical architecture with hybrid RSUs, and the effectiveness and complexity are verified by theoretical analysis and experiments. Zhen Gao 0005, Dongbin Zhang, Jiuzhi Zhang, Lei Liu 0031, Dusit Niyato, Victor C. M. Leung |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Reliability Evaluation and Fault Tolerance Design for FPGA Implemented Reed Solomon (RS) Erasure DecodersabstractReed–Solomon erasure codes (RS-ECs) are widely applied in storage and packet communication systems to recover erasures. When implemented on a field-programmable gate array (FPGA) in a space platform, the RS-EC decoder will suffer single event upsets (SEUs) that can cause failures. In this brief, the reliability of an RS-EC decoder implemented on an FPGA to errors on the configuration memory is first studied based on hardware SEU injection experiments. We found that the reliability is lower for larger number of erased symbols, but there are still about 85% SEUs can be tolerated by the decoder itself even for the maximum number of erased symbols within the recovery capability. In addition, around 10%–25% SEUs on critical bits can cause system exceptions. Based on these results, a duplication with comparison (DWC) scheme is proposed for the protection of the RS-EC decoder. In particular, a checksum parity-based approach is proposed to detect the faulty decoder to reduce the computation overhead. Experimental results show that the reliability of the DWC protected RS-EC decoder to SEUs on the configuration memory is almost the same of a traditional triple modular redundancy (TMR) protection, and the resource usage is only about$2.15\times $that of the unprotected decoder. Zhen Gao 0005, Jinchang Shi, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Evaluation of Blockchain-enabled Mobile Core Network Control Plane for Satellite-terrestrial Integrated NetworksabstractIn order to reduce the signaling transmission latency of 5G satellite-terrestrial integrated networks, this paper proposed a blockchain-enabled mobile core network control plane, namely BC-CN-CP Slice. It can be deployed at the edge of the mobile network adjacent to RAN with blockchain nodes as storage unit. BC-CN-CP Slices in various regions synchronize data automatically by satellite communications, and each slice has complete and consistent data to form the global view of entire network, so that the UE attachment can be processed locally without the assistance of satellite. In this way, signaling transmission latency can be greatly reduced. Furthermore, function calls instead of NFs in traditional 5GC SBA are used to speed up signaling processing in CN CP. The experimental evaluation results showed that, compared with 5GC SBA, BC-CN-CP Slice can reduce the latency of UE attachment to 6.56%, 16.8% and 40.63% for one-way satellite link delay of 250ms, 75ms and 9ms, respectively. Ming Zhao 0001, Zhen Gao 0005 |
ICC | 4 |
| 2022 | Special Session: Fault-Tolerant Deep Learning: A Hierarchical PerspectiveabstractWith the rapid advancements of deep learning in the past decade, it can be foreseen that deep learning will be continuously deployed in more and more safety-critical applications such as autonomous driving and robotics. In this context, reliability turns out to be critical to the deployment of deep learning in these applications and gradually becomes a first-class citizen among the major design metrics like performance and energy efficiency. Nevertheless, the back-box deep learning models combined with the diverse underlying hardware faults make resilient deep learning extremely challenging. In this special session, we conduct a comprehensive survey of fault-tolerant deep learning design approaches with a hierarchical perspective and investigate these approaches from model layer, architecture layer, circuit layer, and cross layer respectively. Cheng Liu 0008, Zhen Gao 0005, Siting Liu 0001, Xuefei Ning, Huawei Li 0001, Xiaowei Li 0001 |
VTS | 2 |
| 2022 | RNS-Based Adaptive Compression Scheme for the Block Data in the Blockchain for IIoTabstractThe Industrial Internet of Things (IIoT) is the essential component of Industry 4.0. Blockchain is a promising technology for secure data sharing and trustable cooperation between IIoT devices. However, the ever-growing transaction records make it difficult for the storage-limited IIoT devices to join the blockchain network. In this article, an adaptive compression scheme is proposed to decrease the storage volume on each node. In the scheme, the block body is compressed by representing the included transactions as their remainders stored in the distributed nodes. The original transaction could be recovered based on the Chinese remainder theorem. In particular, each node adapts its compression ratio according to its storage resource. The nodes storing more data have advantages in transaction recovery, introducing an incentive mechanism for efficient storage utilization. The theoretical analysis and simulation results show that the proposed scheme can achieve a high compression ratio with good service availability. The proposed scheme dramatically lowers the threshold for IIoT devices to join the blockchain network, which is important for the large-scale application of blockchain in Industry 4.0. Zhaohui Guo, Zhen Gao 0005, Qiang Liu 0011, Chinmay Chakraborty, Qiaozhi Hua, Keping Yu, Shaohua Wan 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Soft Error Tolerant Convolutional Neural Networks on FPGAs With Ensemble LearningabstractConvolutional neural networks (CNNs) are widely used in computer vision and natural language processing. Field-programmable gate arrays (FPGAs) are popular accelerators for CNNs. However, if used in critical applications, the reliability of FPGA-based CNNs becomes a priority because FPGAs are prone to suffer soft errors. Traditional protection schemes, such as triple modular redundancy (TMR), introduce a large overhead, which is not acceptable in resource-limited platforms. This article proposes to use an ensemble of weak CNNs to build a robust classifier with low cost. To have a group of base CNNs with low complexity and balanced similarity and diversity, residual neural networks (ResNets) with different layers (20/32/44/56) are combined in the ensemble system to replace a single strong ResNet 110. In addition, a robust combiner is designed based on the reliability evaluation of a single ResNet. Single ResNets with different layers and different ensemble schemes are implemented on the FPGA accelerator based on Xilinx Zynq 7000 SoC. The reliability of the ensemble systems is evaluated based on a large-scale fault injection platform and compared with that of the TMR-protected ResNet 110 and ResNet 20. Experiment results show that the proposed ensembles could effectively improve the system reliability when suffering soft errors with an overhead much lower than TMR. Zhen Gao 0005, Jiajun Xiao, Shulin Zeng, Guangjun Ge, Yu Wang 0002, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Reliability Evaluation of the Count Min Sketch (CMS) against Single Event Transients (SETs)abstractEstimating the frequency of the elements in a data set is commonly needed in data analysis. With the increase of the size of the data sets, accurately computing the number of times that each element appears with a counter becomes impractical. Instead, the Count Min Sketch (CMS) is widely used in big data processing to estimate frequency due to its simplicity and small storage needs. However, soft errors caused by Single Event Transients (SETs) will affect the hardware implementation of the CMS, mainly the hash functions. In this paper, the effect of SETs on the hash functions of the CMS frequency estimate is analyzed theoretically in terms of overestimation probability, underestimation probability, and the equal probability, and further discussed for data with different frequency. Simulation results verify the correctness of the theoretical analysis and reveal several valuable conclusions. First, a large portion of SETs can be tolerated by the CMS itself, and the reliability of the CMS improves when larger number of arrays are used. Second, the average probability for overestimation and underestimation are almost the same, and decrease for larger numbers of arrays. Third, SETs are more likely to cause underestimation for the most frequent data elements. Finally, the overall effect of SETs on the CMS is slightly affected by the number of counters in each array, and seems to be independent of the distribution of the input sequence. The results and analysis presented in this paper provide a starting point for the design of efficient SET fault-tolerant schemes for the CMS. Zhen Gao 0005, Pedro Reviriego |
VTS | 2 |
| 2021 | FTT-NAS: Discovering Fault-tolerant Convolutional Neural ArchitectureabstractWith the fast evolvement of embedded deep-learning computing systems, applications powered by deep learning are moving from the cloud to the edge. When deploying neural networks (NNs) onto the devices under complex environments, there are various types of possible faults: soft errors caused by cosmic radiation and radioactive impurities, voltage instability, aging, temperature variations, malicious attackers, and so on. Thus, the safety risk of deploying NNs is now drawing much attention. In this article, after the analysis of the possible faults in various types of NN accelerators, we formalize and implement various fault models from the algorithmic perspective. We propose Fault-Tolerant Neural Architecture Search (FT-NAS) to automatically discover convolutional neural network (CNN) architectures that are reliable to various faults in nowadays devices. Then, we incorporate fault-tolerant training (FTT) in the search process to achieve better results, which is referred to as FTT-NAS. Experiments on CIFAR-10 show that the discovered architectures outperform other manually designed baseline architectures significantly, with comparable or fewer floating-point operations (FLOPs) and parameters. Specifically, with the same fault settings, F-FTT-Net discovered under the feature fault model achieves an accuracy of 86.2% (VS. 68.1% achieved by MobileNet-V2), and W-FTT-Net discovered under the weight fault model achieves an accuracy of 69.6% (VS. 60.8% achieved by ResNet-18). By inspecting the discovered architectures, we find that the operation primitives, the weight quantization range, the capacity of the model, and the connection pattern have influences on the fault resilience capability of NN models. Xuefei Ning, Guangjun Ge, Zhenhua Zhu 0002, Xiaoming Chen 0003, Zhen Gao 0005, Yu Wang 0002, Huazhong Yang |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2021 | Design of FPGA-Implemented Reed-Solomon Erasure Code (RS-EC) Decoders With Fault Detection and Location on User MemoryabstractReed-Solomon erasure codes (RS-ECs) are widely used in packet communication and storage systems to recover erasures. When the RS-EC decoder is implemented on a field-programmable gate array (FPGA) in a space platform, it will suffer single-event upsets (SEUs) that can cause failures. In this article, the reliability of an RS-EC decoder implemented on an FPGA when there are errors in the user memory is first studied. Then, a fault detection and location scheme is proposed based on partial reencoding for the faults in the user memory of the RS-EC decoder. Furthermore, check bits are added in the generator matrix to improve the fault location performance. The theoretical analysis shows that the scheme could detect most faults with small missing and false detection probability. Experimental results on a case study show that more than 90% of the faults on user memory could be tolerated by the decoder, and all the other faults can be detected by the fault detection scheme when the number of erasures is smaller than the correction capability of the code. Although false alarms exist (with probability smaller than 4%), they can be used to avoid fault accumulation. Finally, the fault location scheme could accurately locate all the faults. The theoretical estimates are very close to the experiment results, which verifies the correctness of the analysis done. Zhen Gao 0005, Yinghao Cheng, Kangkang Guo, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Design of SEU-Tolerant Turbo Decoders Implemented on SRAM-FPGAsabstractTurbo codes are widely used in satellite communications. When a turbo decoder is implemented on a field-programmable gate array (FPGA) in a space platform, it will suffer single-event upsets (SEUs) that can cause failures and disrupt communications. Therefore, the protection of turbo decoders implemented on FPGAs is important. In this article, first the reliability of an SRAM-FPGA-implemented turbo decoder to SEUs on user memory and configuration memory is evaluated based on fault injection experiments. Then, based on the features of the turbo decoder and the characteristics of the failures revealed by the reliability study, a duplication with comparison (DWC) scheme is proposed for the protection of the turbo decoder. Experimental results show that the reliability of the protected turbo decoder to SEUs on user memory and configuration memory is improved by 99.4% and 95.6%, respectively. The resource usage is about 2.2× that of an unprotected turbo decoder, which is significantly lower than the more than 3× required by the traditional triple modular redundancy (TMR) protection. Finally, the proposed scheme is compared with another two protection schemes. Zhen Gao 0005, Tong Yan, Kangkang Guo, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | A Scheme to Design Concurrent Error Detection Techniques for the Fast Fourier Transform Implemented in SRAM-Based FPGAsabstractSoft errors are an important issue for SRAM-based Field Programmable Gate Arrays (FPGAs), since they result in permanent alterations of the mapped circuit when they affect their configuration memory. Concurrent Error Detection (CED) techniques, such as Dual Modular Redundancy (DMR), are usually employed to detect errors that affect the performance of the circuit. When trying to detect errors produced on the complex Fast Fourier Transform (FFT), the Parseval Sum of Squares (SoS) is a widely used technique. In this paper, we present a scheme to implement CED techniques for the complex FFT implemented in SRAM-based FPGAs. These techniques perform checks based on the relationships existing between one or more of the inputs and the outputs of the algorithm. Three examples of these techniques are provided to further clarify how to construct them. These techniques, along with DMR and SoS, have been tested through fault injection. An analysis on their error detection capabilities shows that they achieve high detection rates with much less resource usage than DMR and SoS. In addition, the number of false error detections for these techniques is lower than that of SoS, which leads to less unnecessary reconfigurations of the device. Ricardo Gonzalez-Toral, Pedro Reviriego, Juan Antonio Maestro, Zhen Gao 0005 |
IEEE Trans. Computers | 4 |
| 2018 | An Efficient Fault-Tolerance Design for Integer Parallel Matrix-Vector MultiplicationsabstractParallel matrix processing is a typical operation in many systems, and in particular matrix-vector multiplication (MVM) is one of the most common operations in the modern digital signal processing and digital communication systems. This paper proposes a fault-tolerant design for integer parallel MVMs. The scheme combines ideas from error correction codes with the self-checking capability of MVM. Field-programmable gate array evaluation shows that the proposed scheme can significantly reduce the overheads compared to the protection of each MVM on its own. Therefore, the proposed technique can be used to reduce the cost of providing fault tolerance in practical implementations. Zhen Gao 0005, Qingqing Jing, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |