VLDB 2026 Research / reviewers in the wild / expert
Pedro Reviriego
dblp:60/2579 · also Pedro Reviriego Vasallo
· DBLP profile ↗
121ranked-venue papers
41as first author
52since 2021 · last 2025
0000-0003-2540-5234ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 20 first-author · 26 since 2021Computer networks · 16 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 1 since 2021Security and privacy · 8 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Theory of computation · 4 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin AmericaabstractMaría Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez, Gonzalo Santamaria Gomez, Rodrigo Agerri, Nuria Aldama García, Luis Chiruzzo, Javier Conde, Helena Gomez Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia López Fuertes, Flor Miriam Plaza-del-Arco, María-Teresa Martín-Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, Estrella Vallecillo-Rodríguez, Jorge Vallego, Irune Zubiaga. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez 0001, Gonzalo Santamaría Gómez, Rodrigo Agerri, Nuria Aldama-García, Luis Chiruzzo, Javier Conde, Helena Gómez-Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia Fuertes, Flor Miriam Plaza del Arco, María Teresa Martín Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, María Estrella Vallecillo Rodríguez, Jorge Vallego, Irune Zubiaga |
ACL (1) | 20 |
| 2025 | AIQUIZ: An Open-Source Artificial Intelligence Multiple-Choice Question Generation PlatformabstractThis paper presents AIQUIZ, an open-source webbased platform designed for the automatic generation of multiplechoice questions (MCQs) using large language models (LLMs). AIQUIZ aims to streamline the labor-intensive process of manually creating MCQs to enhance the learning experience by providing adaptive questions tailored to students' past performance. The platform supports A/B testing, allowing educators and researchers to experiment and evaluate different prompts and LLMs to optimize the quality and relevance of the generated questions. AIQUIZ was applied to four engineering courses at Polytechnic University of Madrid (UPM), where it demonstrated promising results, including a high accuracy rate in student responses and a low rate of reported errors in generated questions. The paper describes the platform's key functionalities and results. Enrique Barra, Javier Conde, Anabel Pilicita-Garrido, Alejandro Pozo, Sonsoles López-Pernas, Pedro Reviriego |
ICALT | 6 |
| 2025 | Perturbation-based error detection and correction (PBEDC) in dependable large-scale machine learning systems
Ziheng Wang 0005, Pedro Reviriego, Shanshan Liu 0001, Farzad Niknia, Xiaochen Tang, Zhen Gao 0005, Fabrizio Lombardi |
Future Gener. Comput. Syst. | 2 |
| 2025 | Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)abstractStochastic computing (SC) has emerged as an efficient low-power alternative for deploying neural networks (NNs) in resource-limited scenarios, such as the Internet of Things (IoT). By encoding values as serial bitstreams, SC significantly reduces energy dissipation compared to conventional floating-point (FP) designs; however, further improvement of layer-wise mixed-precision implementation for SC remains unexplored. This paper introduces Adjustable Sequence Length (ASL), a novel scheme that applies mixedprecision concepts specifically to SC NNs. By introducing an operator-norm – based theoretical model, this paper shows that truncation noise can cumulatively propagate through the layers by the estimated amplification factors. An extended sensitivity analysis is presented, using Random Forest (RF) regression to evaluate multi-layer truncation effects and validate the alignment of theoretical predictions with practical network behaviors. To accommodate different application scenarios, this paper proposes two truncation strategies (coarse-grained and fine-grained), which apply diverse sequence length configurations at each layer. Evaluations on a pipelined SC MLP synthesized at 32 nm demonstrate that ASL can reduce energy and latency overheads by up to over 60% with negligible accuracy loss. It confirms the feasibility of the ASL scheme for IoT applications and highlights the distinct advantages of mixed-precision truncation in SC designs. Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Zhen Gao 0005, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Internet Things J. | 2 |
| 2025 | Concurrent Linguistic Error Detection (CLED): A New Methodology for Error Detection in Large Language ModelsabstractThe utilization of Large Language Models (LLMs) requires dependable operation in the presence of errors in the hardware (caused by for example radiation) as this has become a pressing concern. At the same time, the scale and complexity of LLMs limit the overhead that can be added to detect errors. Therefore, there is a need for low-cost error detection schemes. Concurrent Error Detection (CED) uses the properties of a system to detect errors, so it is an appealing approach. In this paper, we present a new methodology and scheme for error detection in LLMs: Concurrent Linguistic Error Detection (CLED). Its main principle is that the output of LLMs should be valid and generate coherent text; therefore, when the text is not valid or differs significantly from the normal text, it is likely that there is an error. Hence, errors can potentially be detected by checking the linguistic features of the text generated by LLMs. This has the following main advantages: 1) low overhead as the checks are simple and 2) general applicability, so regardless of the LLM implementation details because the text correctness is not related to the LLM algorithms or implementations. The proposed CLED has been evaluated on two LLMs: T5 and OPUS-MT. The results show that with a 1% overhead, CLED can detect more than 87% of the errors, making it suitable to improve LLM dependability at low cost. Javier Conde, Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 4 |
| 2025 | Dependability of the K Minimum Values Sketch: Protection and Comparative AnalysisabstractA basic operation in big data analysis is to find the cardinality estimate; to estimate the cardinality at high speed and with a low memory requirement, data sketches that provide approximate estimates, are usually used. The K Minimum Value (KMV) sketch is one of the most popular options; however, soft errors on memories in KMV may substantially degrade performance. This paper is the first to consider the impact of soft errors on the KMV sketch and to compare it with HyperLogLog (HLL), another widely used sketch for cardinality estimate. Initially, the operation of KMV in the presence of soft errors (so its dependability) in the memory is studied by a theoretical analysis and simulation by error injection. The evaluation results show that errors during the construction phase of KMV may cause large deviations in the estimate results. Subsequently, based on the algorithmic features of the KMV sketch, two protection schemes are proposed. The first scheme is based on using a single parity check (SPC) to detect errors and reduce their impact on the cardinality estimate; the second scheme is based on the incremental property of the memory list in KMV. The presented evaluation shows that both schemes can dramatically improve the performance of KMV, and the SPC scheme performs better even though it requires more memory footprint and overheads in the checking operation. Finally, it is shown that soft errors on the unprotected KMV produce larger worst-case errors than in HLL, but the average impact of errors is lower; also, the protected KMV using the proposed schemes are more dependable than HLL with existing protection techniques. Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2025 | Beware of Words: Evaluating the Lexical Diversity of Conversational LLMs using ChatGPT as Case StudyabstractThe performance of conversational Large Language Models (LLMs) in general, and of ChatGPT in particular, is currently being evaluated on many different tasks, from logical reasoning or math to answering questions on a myriad of topics. Instead, much less attention is being devoted to the study of the linguistic features of the texts generated by these LLMs. This is surprising since LLMs are models for language, and understanding how they use the language is important. Indeed, conversational LLMs are poised to have a significant impact on the evolution of languages as they may eventually dominate the creation of new text. This means that for example, if conversational LLMs do not use a word it may become less and less frequent and eventually stop being used altogether. Therefore, evaluating the linguistic features of the text they produce and how those depend on the model parameters is the first step toward understanding the potential impact of conversational LLMs on the evolution of languages. In this article, we consider the evaluation of the lexical diversity of the text generated by LLMs in English and how it depends on the model parameters. A methodology is presented and used to conduct a comprehensive evaluation of lexical diversity using ChatGPT as a case study. The results show how lexical diversity depends on the version of ChatGPT and some of its parameters, such as the presence penalty, or the role assigned to the model. The dataset and tools used in our analysis are released under open licenses with the goal of drawing much-needed attention to the evaluation of the linguistic features of LLM-generated text. Gonzalo Martínez 0001, José Alberto Hernández 0001, Javier Conde, Pedro Reviriego, Elena Merino Gómez |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | Detect and Replace: Efficient Soft Error Protection of FPGA-Based CNN AcceleratorsabstractConvolutional neural networks (CNNs) are widely used in computer vision and natural language processing. Field-programmable gate arrays (FPGAs) are a popular accelerator for CNNs. However, FPGAs are prone to suffer soft errors, so the reliability of FPGA-based CNNs becomes a key problem when used in safety-critical applications. The convolution module based on a processing element (PE) array is the most complex part of the accelerator, so it is the key to efficient protection. Coding-based schemes have been proposed for efficient protection of the convolution module, where the processing of the PE array is modeled as parallel matrix-vector multiplications (MVMs), and every wrong output would be concurrently detected and corrected. However, these schemes cannot deal with errors in the configuration memory that affects many intermediate results. In this article, a protection scheme is proposed based on faulty PE detection and replace (DR) to deal with such configuration memory errors. The DR scheme is implemented on a CNN accelerator based on Xilinx Zynq 7000 SoC, and fault injection (FI) experiments are performed to evaluate the performance of the proposed DR scheme. The results show that it can effectively mitigate the effect of soft errors in the configuration memory with an overhead of about 1.3 times complexity and 1.4 times power consumption relative to those of the unprotected PE array. Compared with the advanced checksum-of-checksum (CoC) scheme, the DR scheme decreases power consumption by up to 30%. Zhen Gao 0005, Yanmao Qi, Jinchang Shi, Qiang Liu 0011, Guangjun Ge, Yu Wang 0002, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Combined Filtering and Frequency Estimation with the Integrated Xor Filter and Count-Min SketchabstractMonitoring and analysis of network traffic with proper accuracy and efficiency are paramount in computer networks management and security. An example of function required for this type of applications is per-flow packet accounting. Having an efficient data structure that provides such functionality while maintaining good performance can become challenging. In this paper, we propose an integrated data structure that combines the ability of the xor filter to detect packets from selected flows and the Count-Min sketch (CMS) to efficiently estimate the frequencies. The results show that the integrated filter improves performance compared to keeping both structures separate, achieving a reduction of up to 23% in the probability of false positives, fewer memory access per operation, and a reduction of the Average Relative Error in the CMS as more packets are analyzed. Roberto Martínez Aguilar, Pedro Reviriego, David Larrabeiti |
HPSR | 2 |
| 2024 | Detection of NFT Duplications with Image Hash FunctionsabstractNon-fungible tokens (NFTs) are digital assets representing ownership or proof of authenticity of a unique item. NFTs are blockchain-based and rely on smart contracts. The increase in duplicate NFTs in recent years brings the need for discovery tools for forged NFTs, some of which include using image hash functions. Though the problem of image duplication is widely discussed, detecting NFT duplications requires using fast detection methods as a new NFT image needs to be compared with the entire NFT history on the blockchain. In this paper, we analyze the performance of several image-hash functions, examine the cases where each function performs well, and evaluate multiple image-hash-functions-based NFT duplication detectors. Our approach achieves high accuracy in detecting NFT duplications and demonstrates that using several hash functions rather than one increases the ability to detect duplications. Arad Kotzer, Mostafa Naamneh, Ori Rottenstreich, Pedro Reviriego |
ICBC | 4 |
| 2024 | Reducing the Energy Dissipation of Large Language Models (LLMs) with Approximate MemoriesabstractLarge language models (LLMs) have shown impressive performance in a wide range of tasks such as answering questions or summarizing text. However, running LLMs on edge devices is challenging as they require large amounts of energy due to their memory and computation needs. In LLMs most of the memory is needed to store the model parameters which number keeps increasing from one LLM generation to the next. In the last several years, significant efforts have been made to compress and prune parameters, but this is not enough to reduce their memory needs as the number of parameters grows exponentially. In this work, to reduce energy dissipation, rather than trying to reduce the amount of memory used by LLMs, we study the use of approximate memories to store the LLM parameters. Approximate memories can significantly reduce the energy dissipation at the cost of introducing errors in some of the memory bits. Therefore, the impact of errors on LLMs must be understood. To that end, we have performed error injection on different compressed versions of a classic LLM: Bidirectional Encoder Representations from Transformers (BERT). The results show that in some cases compressed BERTs operate reliably at high bit error rates. This makes possible the use of approximate memories with a negligible impact on the LLM performance and a significant reduction in energy dissipation. Zhen Gao 0005, Pedro Reviriego, Shanshan Liu 0001, Fabrizio Lombardi |
ISCAS | 3 |
| 2024 | Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of ThingsabstractThe implementation of Machine Learning (ML) in Internet of Things (IoT) devices poses significant operational challenges due to limited energy and computation resources. In recent years, significant efforts have been made to implement simplified ML models that can achieve reasonable performance while reducing computation and energy, for example by pruning weights in neural networks, or using reduced precision for the parameters and arithmetic operations. However, this type of approach is limited by the performance of the ML implementation, i.e., by the loss for example in accuracy due to the model simplification. In this paper, we present Adaptive Resolution Inference (ARI), a novel approach that enables to evaluate new trade-offs between energy dissipation and model performance in ML implementations. The main principle of the proposed approach is to run inferences with reduced precision (quantization) and use the margin over the decision threshold to determine if either the result is reliable, or the inference must run with the full model. The rationale is that quantization only introduces small deviations in the inference scores, such that if the scores have a sufficient margin over the decision threshold, it is very unlikely that the full model would have a different result. Therefore, we can run the quantized model first, and only when the scores do not have a sufficient margin, the full model is run. This enables most inferences to run with the reduced precision model and only a small fraction requires the full model, so significantly reducing computation and energy while not affecting model performance. The proposed ARI approach is presented, analyzed in detail, and evaluated using different datasets both for floating-point and stochastic computing implementations. The results show that ARI can significantly reduce the energy for inference in different configurations with savings between 40% and 85%. Ziheng Wang 0005, Pedro Reviriego, Farzad Niknia, Javier Conde, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Internet Things J. | 2 |
| 2024 | On the Security of Quotient Filters: Attacks and Potential CountermeasuresabstractThe security of probabilistic data structures is increasingly important due to their wide adoption in many computing systems and applications. In particular, the security of approximate membership check filters such as Bloom or cuckoo filters has been recently studied showing how an attacker can degrade the filter performance in some settings. In this paper, we consider for the first time the security of another popular approximate membership check filter, the Quotient Filter (QF). Our analysis and simulations show that quotient filters are vulnerable to both white and black box attackers that can cause insertion failures and degrade the filter performance very significantly. An interesting finding is that quotient filters are vulnerable to a new type of attack, not applicable to Bloom or cuckoo filters, that can degrade the speed of queries dramatically. The paper also briefly discusses and evaluates potential countermeasures to detect and protect against those attacks. Pedro Reviriego, Miguel González 0005, Niv Dayan, Gabriel Huecas, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2024 | On the Privacy of Multi-Versioned Approximate Membership Check FiltersabstractApproximate membership filters are increasingly used in many computing and networking applications and new filter designs are being continuously presented to improve one or more performance metrics. Therefore, understanding their security and privacy is an important issue. Previous works have considered attackers that only have access to an individual filter in isolation. For applications that generate many related filters, such as a filter for a deny list that evolves over time, that analysis is insufficient. This paper considers an attacker with access to several versions of a filter that share most of the same input elements. We find that for typical implementations of Bloom, cuckoo, and quotient filters, the attacker gains little or no advantage with access to multiple versions of a filter. However, typical xor filters do reveal more information about their input elements by querying multiple versions of a filter, and we propose techniques to enhance the privacy of xor filters and others. Pedro Reviriego, Alfonso Sánchez-Macián, Peter C. Dillinger, Stefan Walzer |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | On the Privacy of Adaptive Cuckoo Filters: Analysis and ProtectionabstractAs probabilistic data structures are widely adopted in computing systems, their privacy is a major issue. Recent works have shown that even though the values stored in these structures look random, information can be extracted from them in some settings. In this paper, we consider the privacy of adaptive cuckoo filters, a probabilistic data structure that implements approximate membership checking. The main novelty and benefit of these filters are that they can adapt to removing false-positives. Unfortunately, our analysis shows that adaptation can dramatically reduce the privacy of the filters, allowing an attacker to extract the set of elements stored in the filter. Indeed, in some settings, the attacker can identify 100% of the elements stored in the filter. This means that the protection of the privacy of adaptive cuckoo filters should be considered. To that end, we propose preprocessing reduction (PR), a scheme that prevents an attacker from extracting the set of elements stored in the filter at the cost of increasing the false-positive probability of the filter. In many settings, the impact on false-positives will be negligible. For example, in a case study with 32-bit universes, the increase in the false-positive probability was smaller than 8% in all the configurations tested. Interestingly, PR is applicable not only to adaptive filters but also to approximate membership check filters in general and thus can be used to protect, for example, Bloom filters. Pedro Reviriego, Jim Apple, David Larrabeiti, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Supporting Dynamic Insertions in xor and Binary Fuse Filters With the Integrated xor/BIF-Bloom FilterabstractApproximate membership check filters are widely used in networking applications to resolve membership queries at high speed with a low memory cost. Due to their extensive use, many filter types have been proposed. Two recent and interesting alternatives are the xor filter and the binary fuse filter, which in certain configurations have one of the lowest false positive rates, are faster and use less memory than other filters. However, one of the main drawbacks of xor and binary fuse filters is that it is not possible to add keys once the filter has been built. This limits their use in many network related applications where keys have to be added dynamically. This paper presents the Integrated xor-Bloom filter (IXOR) and the Integrated binary fuse-Bloom filter (IBIF), both schemes allow dynamic insertions in xor and binary fuse filters without the need to reconstruct the filters. The schemes have been implemented and evaluated showing that a large number of dynamic insertions can be supported with a limited memory overhead and a small impact on the false positive probability and lookup speed. Therefore, the proposed filters can bring the benefits of xor and binary fuse filters to networking applications that need to support dynamic insertions. Roberto Martínez, Pedro Reviriego, David Larrabeiti |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Cardinality Estimation Adaptive Cuckoo Filters (CE-ACF): Approximate Membership Check and Distinct Query Count for High-Speed Network MonitoringabstractIn network monitoring applications, it is often beneficial to employ a fast approximate set-membership filter to check if a given packet belongs to a monitored flow. Recent adaptive filter designs, such as the Adaptive Cuckoo Filter, are especially promising for such use cases as they adapt fingerprints to eliminate recurring false positives. In many traffic monitoring applications, it is also of interest to know the number of distinct flows that traverse a link or the number of nodes that are sending traffic. This is commonly done using cardinality estimation sketches. Therefore, on a given switch or network device, the same packets are typically processed using both a filter and a cardinality estimator. Having to process each packet with two independent data structures adds complexity to the implementation and limits performance. This paper shows that adaptive cuckoo filters can also be used to estimate the number of distinct negative elements queried on the filter. In flow monitoring, those distinct queries correspond to distinct flows. This is interesting as we get the cardinality estimation for free as part of the normal adaptive filter’s operation. We provide (1) a theoretical analysis, (2) simulation results, and (3) an evaluation with real packet traces to show that adaptive cuckoo filters can accurately estimate a wide range of cardinalities in practical scenarios. Pedro Reviriego, Jim Apple, Otmar Ertl, Niv Dayan |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Concurrent Classifier Error Detection (CCED) in Large Scale Machine Learning SystemsabstractThe complexity of machine learning (ML) systems increases each year. As these systems are widely utilized, ensuring their reliable operation is becoming a design requirement. Traditional error detection mechanisms introduce circuit or time redundancy that significantly impacts system performance. An alternative is the use of concurrent error detection (CED) schemes that operate in parallel with the system and exploit their properties to detect errors. CED is attractive for large ML systems because it can potentially reduce the cost of error detection. In this article, we introduce concurrent classifier error detection (CCED), a scheme to implement CED in ML systems using a concurrent ML classifier to detect errors. CCED identifies a set of check signals in the main ML system and feed them to the concurrent ML classifier that is trained to detect errors. The proposed CCED scheme has been implemented and evaluated on two widely used large-scale ML models: Contrastive language-image pretraining (CLIP) used for image classification and bidirectional encoder representations from transformers (BERT) used for natural language applications. The results show that more than 95% of the errors are detected when using a simple Random Forest classifier that is orders of magnitude simpler than CLIP or BERT. Pedro Reviriego, Ziheng Wang 0005, Zhen Gao 0005, Farzad Niknia, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Reliab. | 1 |
| 2024 | Enhancing data protection with a distributed storage system based on the redundant residue number systemabstractAbstract Big data becomes the key for ubiquitous computing and intelligence, and Distributed Storage Systems (DSS) are widely used in large-scale data centers or in the cloud for efficient data management. However, the data on stored are likely to be unavailable due to hardware failures and cyberattacks, e.g. DDoS. Maximum Distance Separable (MDS) codes are commonly used for the recovery of faulty storage nodes or unavailable data. However, the recovery of data nodes usually involves access to multiple nodes, which introduces significant communication overheads to the DSS. In this paper, a new DSS based on the Redundant Residue Number System (RRNS) is proposed, where efficient recovery is enabled by applying the second version of Chinese Remainder Theorem (CRT-II). The complexity and network traffic of the proposed data protection scheme is analyzed theoretically and compared with that of traditional MDS based DSSs. Experimental results show that the proposed DSS achieves lower encoding complexity, lower recovery complexity and lower network traffic than the MDS based schemes. Although the proposed data protection scheme introduces computation overheads for the case on which there are no failing nodes, its complexity is still lower for scenarios with frequent data updates. In addition, the proposed scheme introduces additional advantages in terms of security and storage flexibility. Pedro Reviriego |
Wirel. Networks | 3 |
| 2023 | Feature-Embedding Triplet Networks with a Separately Constrained Loss FunctionabstractFeature-embedding triplet networks (TNs) with three symmetric subchannels are very promising for similarity-measuring applications. This paper proposes a novel separately constrained triple loss (SCTL) function that applies to TNs for classification. Through minimizing the intra-class distance and maximizing the inter-class distance, SCTL eliminates possible false solutions and provides insight into the dependency of training based on these two terms. Based on this dependency, the strategy of selecting hyperparameters in SCTL is also analyzed to further improve performance. The effectiveness of the proposed SCTL is evaluated based on TNs with multi-layer perceptrons; the results show that compared to all existing loss functions, the use of SCTL offers the best classification accuracy for the TNs, while incurring in negligible hardware overhead (e.g., only a 0.0002% area overhead of the subnetworks). Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Honglan Jiang, Siting Liu 0001, Pedro Reviriego, Fabrizio Lombardi |
ISCAS | 6 |
| 2023 | InfiniFilter: Expanding Filters to Infinity and BeyondabstractFilter data structures have been used ubiquitously since the 1970s to answer approximate set-membership queries in various areas of computer science including architecture, networks, operating systems, and databases. Such filters need to be allocated with a given capacity in advance to provide a guarantee over the false positive rate. In many applications, however, the data size is not known in advance, requiring filters to dynamically expand. This paper shows that existing methods for expanding filters exhibit at least one of the following flaws: (1) they entail an expensive scan over the whole data set, (2) they require a lavish memory footprint, (3) their query, delete and/or insertion performance plummets, (4) their false positive rate skyrockets, and/or (5)~they cannot expand indefinitely. We introduce InfiniFilter, a new method for expanding filters that addresses these shortcomings. InfiniFilter is a hash table that stores a fingerprint for each entry. It doubles in size when it reaches capacity, and it sacrifices one bit from each fingerprint to map it to the expanded hash table. The core novelty is a new and flexible hash slot format that sets longer fingerprints to newer entries. This keeps the average fingerprint length long and thus the false positive rate stable. At the same time, InfiniFilter provides stable insertion/query/delete performance as it is comprised of a unified hash table. We implement InfiniFilter on top of Quotient Filter, and we demonstrate theoretically and empirically that it offers superior cost properties compared to existing methods: it better scales performance, the false positive rate, and the memory footprint, all at the same time. Niv Dayan, Ioana O. Bercea, Pedro Reviriego, Rasmus Pagh |
Proc. ACM Manag. Data | 3 |
| 2023 | Attacking the Privacy of Approximate Membership Check Filters by Positive ConcentrationabstractApproximate membership check filters are increasingly used to speed up data processing in many applications. Also, privacy is becoming a key design objective for many systems and thus, the privacy of filters needs to be carefully considered. Previous works have shown that an attacker that knows the implementation details of the filter and has access to its content, may be able to extract some information about the elements stored in the filter. This attack is, however, specific to Bloom filters and requires that the universe of elements must be small. In this article, we show that in many practical settings, an attacker that has only a black-box access to the filter, can extract information about the elements stored in the filter regardless of the specific filter type and the universe size. This is possible based on the key observation that in many applications, the elements stored in the filter are not randomly chosen, but they are concentrated in one or more parts of the universe of elements. To identify these parts, the positive probability can be measured on different parts of the universe; the parts having significantly larger values than the average positive probability for the filter are the ones on which the filter elements are concentrated. This approach is formalized and applied to several case studies showing the process by which the attacker can get additional information about the elements stored for the filters in a wide range of scenarios. Pedro Reviriego, Alfonso Sánchez-Macián, Elena Merino Gómez, Ori Rottenstreich, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2023 | Tolerance of Siamese Networks (SNs) to Memory Errors: Analysis and DesignabstractThis article considers memory errors in a Siamese Network (SN) through an extensive analysis and proposes two schemes (using a weight filter and a code) to provide efficient hardware solutions for error tolerance. Initially the impact of memory errors on the weights of the SN (stored as floating-point (FP) numbers) is analyzed; this shows that the degradation is mostly caused by outliers in weights. Two schemes are subsequently proposed. An analysis is pursued to establish the filter's bounds selection by the maximum/minimum values of the weight distributions, by which outliers can be removed from the operation of the SN. A code scheme for protecting the sign and exponent bits of each weight in an FP number, is also proposed; this code incurs in no memory overhead by utilizing the 4 least significant bits (LSB) to store parity bits. Simulation shows that the filter has a better performance for multi-bit errors correction (a reduction of 95.288% in changed predictions), while the code achieves superior results in single-bit errors correction (a reduction of 99.775% in changed predictions). The combined method that uses the two proposed schemes, retains their advantages, so adaptive to all scenarios; The ASIC-based FP designs of the SN using serial and hybrid implementations are also presented; these pipelined designs utilize a novel multi-layer perceptron (MLP) (as branch networks of the SN) that operates at a frequency of 681.2 MHz (at a 32nm technology node), so significantly higher than existing designs found in the technical literature. The proposed error-tolerant approaches also show advantages in overheads comparing with for example traditional error correction code (ECC). These error-tolerant MLP-based designs are well suited to hardware/power-constrained platforms. Ziheng Wang 0005, Farzad Niknia, Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 4 |
| 2023 | Efficient Protection of FPGA Implemented LDPC Decoders Against Single Event Upsets (SEUs) on Configuration MemoriesabstractLow Density Parity Check (LDPC) codes are used in 5G systems for traffic channels due to their excellent error correction capability for long sequences, and the Min-Sum algorithm is widely applied in practical implementations of LDPC decoders due to its low complexity. If the decoder is implemented on a SRAM-based field-programmable gate array (SRAM-FPGA), the radiation-induced single-event upsets (SEUs) can affect the operation of the LDPC decoder by corrupting the configuration memory, which can change the circuit functionality and will not be corrected unless the FPGA is reconfigured. Therefore, protection of LDPC decoders with low overhead is an important problem, especially for resource-limited on-board space systems. In this paper, an efficient Duplicate With Comparison (DWC) protection scheme is proposed based on the different distribution of the parity check sum of the LDPC decoder in the error-free case and the faulty case. In particular, the check sum accumulation number and threshold are optimized to achieve high detection probability with short delay. FPGA based implementation and hardware fault injection experiments are conducted to evaluate the performance of the proposed schemes. Experimental results show that, the effect of SEUs on the LDPC decoder can be completely eliminated by the proposed scheme with 2 times computational overhead and 1.69 times power consumption overhead compared to the unprotected decoder. Zhen Gao 0005, Yinghao Cheng, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | A Methodology for the Design of Fault Tolerant Parallel Digital Channelizers on SRAM-FPGAsabstractDigital channelizers (DCs) based on the Discrete Fourier Transform (DFT) and polyphase filter banks are widely used in on-board processing (OBP) platforms to extract narrowband sub-channels from a wideband signal efficiently. In high-capacity communication satellite platforms there are always multiple DCs extracting narrowband signals from multiple wideband signals in parallel. Field-programmable gate arrays (FPGAs) are a popular option for the implementation of DCs due to their parallel computing capabilities and good re-configurability, but FPGAs suffer single-event upsets (SEUs) on the space platform. This paper focuses on the efficient protection of parallel DCs with enhanced coding techniques. We first prove that a linear relationship between parallel DCs can be introduced and maintained among the multiple outputs. However, traditional coding schemes cannot be directly applied for the detection of faulty DCs due to the quantization noise introduced by fixed point implementations. To address this issue, we propose an enhanced coding scheme by averaging in the space and time domains to minimize the effect of quantization noise, introducing thresholds and a majority voter to further improve the detection probability. Both theoretical analysis and fault injection experiments prove the effectiveness of the proposed protection scheme. Experimental results show that all the SEUs that cause an SNR lower than 20dB can be detected and recovered, and the resource overheads are about 1.6 times and 1.3 times of that of the unprotected DCs for systems with 8 DCs and 16 DCs, respectively. Zhen Gao 0005, Jiajun Xiao, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Error-Resilient Data Compression With Tunstall CodesabstractData compression has been commonly employed to reduce the required memory size for emerging applications with large storage needs like Big Data and Machine Learning (ML). When considering the flexibility of decompression and its hardware implementation, variable-to-fixed length codes (e.g., Tunstall codes) are usually selected. However, memories are prone to suffer different types of errors, causing the stored data to be corrupted; if an error affects the compressed data, it can propagate and cause corruption in a sequence of bits of the decompressed data. Therefore, error resilience should be built-in as part of the memory design to provide reliable data, especially for safety-critical applications. However, Error Correction Codes (ECCs) that are widely used for memory protection, are not very efficient to protect compressed data, because ECCs further increase the memory size and the additional decoding process can impact the latency to decompress the stored data. In this paper, an efficient error-resilient data compression technique with Tunstall codes is proposed; it requires almost no memory overhead and can correct most errors during the decompression process by introducing a conversion table. An enhanced design is also presented to reduce the impact of errors when they cannot be corrected. The proposed scheme has been implemented and evaluated on three ML datasets; results show that it can deal with up to 99.98% errors with almost no memory overhead when Tunstall codes with smaller than 16-bit symbols are employed. The scheme has also been evaluated for two ML applications; results show that even though a small number of errors cannot be corrected in the proposed scheme, they have an extremely low impact on the classification results and the protection overhead is significantly lower than existing ECC techniques. Shanshan Liu 0001, Pedro Reviriego, Anees Ullah, Ahmed Louri, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | On the Privacy of Counting Bloom Filters Under a Black-Box AttackerabstractCounting Bloom Filters (CBFs) areapproximatemembership checking data structures, and it is normally believed that at most anapproximatereconstruction of the underlying set can be derived when interacting with a CBF. This paper decisively refutes this assumption. In a recent paper, we considered the privacy of CBFs when the attacker has access to the implementation details and thus, it sees the filter as a white-box. In that setting, we showed that the attacker may be able to extract the elements stored in the filter when the number of false positives over the entire universe is not significantly larger than the number of elements stored in the filter. In this work, we consider a black-box attacker that can only perform user interactions on the CBF to insert, remove and query elements with no knowledge of the filter implementation details. We show that even in this case, an attacker may be able to extract information from the filter at the cost of using more complex and time-consuming attack algorithms. The proposed algorithms have been implemented and compared with the white-box attack, showing that in most cases, almost the same information can be extracted from the filter. Sergio Galán, Pedro Reviriego, Stefan Walzer, Alfonso Sánchez-Macián, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | On the Privacy of Counting Bloom FiltersabstractBloom filters are widely used in networking and computing to accelerate membership checking. In many applications filters store sensitive data, so their privacy is of primary concern. At first glance, it seems that extracting the set of elements inserted from the filter would not be possible, because in Bloom filters elements are mapped to positions using hash functions. However, previous works have shown that for the Bloom filter, it may be possible to identify few of the elements inserted in the filter. In this work, we consider the case of counting Bloom filters (CBFs) and show that in some cases, the entire set of elements used to create the filter can be extracted from the filter. This poses serious privacy and security concerns when an attacker can get access to the filter contents. In this article, an algorithm to extract the elements inserted from the filter is presented and analyzed theoretically; then, the feasibility of the CBF inversion is shown by simulation. A case study is presented in detail to illustrate that in practical applications, these conditions can be met by using additional restrictions that are implicit in the nature of the application itself. Pedro Reviriego, Alfonso Sánchez-Macián, Stefan Walzer, Elena Merino Gómez, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Toward Optimal Softcore Carry-aware Approximate Multipliers on Xilinx FPGAsabstractDomain-specific accelerators for signal processing, image processing, and machine learning are increasingly being implemented on SRAM-based field-programmable gate arrays (FPGAs). Owing to the inherent error tolerance of such applications, approximate arithmetic operations, in particular, the design of approximate multipliers, have become an important research problem. Truncation of lower bits is a widely used approximation approach; however, analyzing and limiting the effects of carry-propagation due to this approximation has not been explored in detail yet. In this article, an optimized carry-aware approximate radix-4 Booth multiplier design is presented that leverages the built-in slice look-up tables (LUTs) and carry-chain resources in a novel configuration. The proposed multiplier simplifies the computation of the upper and lower bits and provides significant benefits in terms of FPGA resource usage (LUTs saving 38.5%–42.9%), Power Delay Product (PDP saving 49.4%–53%), performance metric (LUTs × critical path delay (CPD) × PDP saving 68.9%–73.1%) and errors (70% improvement in mean relative error distance) compared to the latest state-of-the-art designs. Therefore, the proposed designs are an attractive choice to implement multiplication on FPGA-based accelerators. Muhammad Awais Khan 0002, Ali Zahir, Syed Ayaz Ali Shah, Pedro Reviriego, Anees Ullah, Nasim Ullah, Adam Khan, Hazrat Ali |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Reliability Evaluation and Fault Tolerance Design for FPGA Implemented Reed Solomon (RS) Erasure DecodersabstractReed–Solomon erasure codes (RS-ECs) are widely applied in storage and packet communication systems to recover erasures. When implemented on a field-programmable gate array (FPGA) in a space platform, the RS-EC decoder will suffer single event upsets (SEUs) that can cause failures. In this brief, the reliability of an RS-EC decoder implemented on an FPGA to errors on the configuration memory is first studied based on hardware SEU injection experiments. We found that the reliability is lower for larger number of erased symbols, but there are still about 85% SEUs can be tolerated by the decoder itself even for the maximum number of erased symbols within the recovery capability. In addition, around 10%–25% SEUs on critical bits can cause system exceptions. Based on these results, a duplication with comparison (DWC) scheme is proposed for the protection of the RS-EC decoder. In particular, a checksum parity-based approach is proposed to detect the faulty decoder to reduce the computation overhead. Experimental results show that the reliability of the DWC protected RS-EC decoder to SEUs on the configuration memory is almost the same of a traditional triple modular redundancy (TMR) protection, and the resource usage is only about$2.15\times $that of the unprotected decoder. Zhen Gao 0005, Jinchang Shi, Qiang Liu 0011, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Is your FPGA bitstream Hardware Trojan-free? Machine learning can provide an answer
Alessandro Palumbo, Luca Cassano, Bruno Luzzi, José Alberto Hernández 0001, Pedro Reviriego, Giuseppe Bianchi 0001, Marco Ottavi |
J. Syst. Archit. | 5 |
| 2022 | Selective Neuron Re-Computation (SNRC) for Error-Tolerant Neural NetworksabstractArtificial Neural networks (ANNs) are widely used to solve classification problems for many machine learning applications. When errors occur in the computational units of an ANN implementation due to for example radiation effects, the result of an arithmetic operation can be changed, and therefore, the predicted classification class may be erroneously affected. This is not acceptable when ANNs are used in many safety-critical applications, because the incorrect classification may result in a system failure. Existing error-tolerant techniques usually rely on physically replicating parts of the ANN implementation or incurring in a significant computation overhead. Therefore, efficient protection schemes are needed for ANNs that are run on a processor and used in resource-limited platforms. A technique referred to as Selective Neuron Re-Computation (SNRC), is proposed in this paper. As per the ANN structure and algorithmic properties, SNRC can identify the cases in which the errors have no impact on the outcome; therefore, errors only need to be handled by re-computation when the classification result is detected as unreliable. Compared with existing temporal redundancy-based protection schemes, SNRC saves more than 60 percent of the re-computation (more than 90 percent in many cases) overhead to achieve complete error protection as assessed over a wide range of datasets. Different activation functions are also evaluated. Shanshan Liu 0001, Pedro Reviriego, Fabrizio Lombardi |
IEEE Trans. Computers | 2 |
| 2022 | A Delta Sigma Modulator-Based Stochastic DividerabstractThe divider is one of the most complex hardware units in Stochastic Computing (SC); even though several new designs have been presented to reduce the computation latency of the conventional divider, all of them still require a considerable number of clock cycles. Moreover, they incur in low performance due to the employed arithmetic computational scheme. In this paper, a Delta Sigma Modulator (DSM) based stochastic divider is proposed. As an entirely digital circuit, the proposed divider offers the best computation latency and accuracy over all existing stochastic dividers found in the technical literature (with a typical reduction between 66.8% and 96.9% in the number of clock cycles and a reduction from$10^{\mathrm {-3.4}}$to$10^{\mathrm {-3.9}}$in the average mean square error for a 10-bit resolution). An SC-based Neural Network (NN) is considered as an initial case study to evaluate the advantages of the proposed design in an emerging application; results show that the proposed divider enables an SC-based NN to achieve a higher classification accuracy and hardware efficiency than existing designs. To show the flexibility of the proposed divider design, its application to Sobol-based sequences is also presented; also in this case, its superiority over other designs is confirmed. These features make the proposed design very attractive for hardware-constrained platforms; moreover, such a novel design approach that incorporates ideas from analog/mixed signal circuit design into a digital circuit design, can motivate other researchers to design efficient SC designs using similar schemes. Xiaochen Tang, Shanshan Liu 0001, Farzad Niknia, Pedro Reviriego, Ziheng Wang 0005, Wei Tang 0002, Ahmed Louri, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Remove Minimum (RM): An Error-Tolerant Scheme for Cardinality Estimate by HyperLogLogabstractEstimating the number of distinct elements is required in many computing applications. One of the state-of-the-art algorithms for cardinality estimate is the HyperLogLog; it provides a good estimate over a large range of cardinality values using a small array of counters. As HLL is implemented in computing systems, it is exposed to soft errors that can corrupt bits stored in memories or registers. To avoid data corruption, memories are commonly protected with Error Correction Codes (ECCs). ECCs however incur in significant overhead because protection needs additional memory cells per word to store the parity check bits as well as additional computation for checking them. In this paper, we first study the impact of soft errors on the HLL algorithm by performing simulation by error injection. The results show that the algorithm is quite robust and can filter out most errors. However, for large cardinalities, there are some errors that can cause a large discrepancy in the HLL estimate. Based on the analysis of the experimental results and the HLL algorithm, a protection technique is proposed that effectively mitigates the impact of soft errors at a small overhead. The proposed Remove Minimum (RM) scheme has been validated by error injection experiments. Pedro Reviriego, Jorge Martínez 0001, Ori Rottenstreich, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | On the Security of the K Minimum Values (KMV) SketchabstractData sketches are widely used to accelerate operations in big data analytics. For example, algorithms use sketches to compute the cardinality of a set, or the similarity between two sets. Sketches achieve significant reductions in computing time and storage requirements by providing probabilistic estimates rather than exact values. In many applications, an estimate is sufficient and thus, it is possible to trade accuracy for computational complexity; this enables the use of probabilistic sketches. However, the use of probabilistic data structures may create security issues because an attacker may manipulate the data in such a way that the sketches produce an incorrect estimate. For example, an attacker could potentially inflate the estimate of the number of distinct users to increase its revenues or popularity. Recent works have shown that an attacker can manipulate Hyperloglog, a sketch widely used for cardinality estimate, with no knowledge of its implementation details. This paper considers the security of K Minimum Values (KMV), a sketch that is also widely used to implement both cardinality and similarity estimates. Next sections characterize vulnerabilities at an implementation-independent level, with attacks formulated as part of a novel adversary model that manipulates the similarity estimate. Therefore, the paper pursues an analysis and simulation; the results suggest that as vulnerable to attacks, an increase or reduction of the estimate may occur. The execution of the attacks against the KMV implementation in the Apache DataSketches library validates these scenarios. Experiments show an excellent agreement between theory and experimental results. Pedro Reviriego, Alfonso Sánchez-Macián, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Breaking Cuckoo Hash: Black Box AttacksabstractIntroduced less than twenty years ago, cuckoo hashing has a number of attractive features like a constant worst case number of memory accesses for queries and close to full memory utilization. Cuckoo hashing has been widely adopted to perform exact matching of an incoming key with a set of stored (key, value) pairs in both software and hardware implementations. This widespread adoption makes it important to consider the security of cuckoo hashing. Most hash based data structures can be attacked by generating collisions that reduce their performance. In fact, for cuckoo hashing collisions could lead to insertion failures which in some systems would lead to a system failure. For example, if cuckoo hashing is used to perform Ethernet lookup and a given MAC address cannot be added to the cuckoo hash, the switch would not be able to correctly forward frames to that address. Previous works have shown that this can be done when the attacker knows the hash functions used in the implementation. However, in many cases the attacker would not have that information and would only have access to the cuckoo hash operations to perform insertions, removals or queries. This article considers the security of a cuckoo hash to an attacker that has only a black box access to it. The analysis shows that by carefully performing user operations on the cuckoo hash, the attacker can force insertion failures with a small set of elements. The proposed attack has been implemented and tested for different configurations to demonstrate its feasibility. The fact that cuckoo hash can be broken with only access to its user functions should be taken into account when implementing it in critical systems. The article also discusses potential approaches to mitigate this vulnerability. Pedro Reviriego, Daniel Ting |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Attacking Adaptive Cuckoo Filters: Too Much Adaptation Can Kill YouabstractAdaptation has recently been proposed to reduce the false positive rate of approximate membership check filters for applications in which the same elements are checked multiple times. Its operational principle is to adapt the filter when a false positive occurs for a given element, such that subsequent checks of that element do not cause a positive result (as beneficial for example in networking). Security is an important consideration for approximate membership check filters and several attacks have been described in the literature; therefore, it is of interest to study the security of adaptive filters. In this paper, we consider adaptive cuckoo filters and show that an attacker can generate sequences of lookups that cause the filter to continuously adapt and not being able to remove the false positives. This degrades the filter performance due to the adaptation overhead; it also makes it harder for other false positives to be removed, because adaptation can be monopolized by the attacker. This can be done when the attacker has only a black-box access to the filter being able to perform lookups but with no knowledge of the implementation of the filter. The proposed attacks have been implemented and tested to validate their effectiveness in terms of the construction of the attack set and the impact of the attack itself. The evaluation results confirm that adaptation unfortunately increases the attack surface of filters and new mechanisms to protect them should be developed. Pedro Reviriego, Alfonso Sánchez-Macián, Salvatore Pontarelli, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | Adaptive One Memory Access Bloom FiltersabstractBloom filters are widely used to perform fast approximate membership checking in networking applications. The main limitation of Bloom filters is that they suffer from false positives that can only be reduced by using more memory. We suggest to take advantage of a common repetition in the identity of queried elements to adapt Bloom filters for avoiding false positives for elements that repeat upon queries. In this paper, one memory access Bloom filters are used to design an adaptation scheme that can effectively remove false positives while completing all queries in a single memory access. The proposed filters are well suited for scenarios on which the number of memory bits per element is low and thus complement existing adaptive cuckoo filters that are not efficient in that case. The evaluation results using packet traces show that the proposed adaptive Bloom filters can significantly reduce the false positive rate in networking applications with the single memory access. In particular, when using as few as four bits per element, false positive rates below 5% are achieved. Pedro Reviriego, Alfonso Sánchez-Macián, Ori Rottenstreich, David Larrabeiti |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | Processor Security: Detecting Microarchitectural Attacks via Count-Min SketchesabstractThe continuous quest for performance pushed processors to incorporate elements such as multiple cores, caches, acceleration units, or speculative execution that make systems very complex. On the other hand, these features often expose unexpected vulnerabilities that pose new challenges. For example, the timing differences introduced by caches or speculative execution can be exploited to leak information or detect activity patterns. Protecting embedded systems from existing attacks is extremely challenging, and it is made even harder by the continuous rise of new microarchitectural attacks (e.g., the Spectre and Orchestration attacks). In this article, we present a new approach based on count-min sketches for detecting microarchitectural attacks in the microprocessors featured by embedded systems. The idea is to add to the system a security checking module (without modifying the microprocessor under protection) in charge of observing the fetched instructions and identifying and signaling possible suspicious activities without interfering with the nominal activity of the system. The proposed approach can be programmed at design time (and reprogrammed after deployment) in order to always keep updated the list of the attacks that the checker is able to identify. We integrated the proposed approach in a large RISC-V core, and we proved its effectiveness in detecting several versions of the Spectre, Orchestration, Rowhammer, and Flush + Reload attacks. In its best configuration, the proposed approach has been able to detect 100% of the attacks, with no false alarms and introducing about 10% area overhead, about 4% power increase, and without working frequency reduction. Kerem Arikan, Alessandro Palumbo, Luca Cassano, Pedro Reviriego, Salvatore Pontarelli, Giuseppe Bianchi 0001, Oguz Ergin, Marco Ottavi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Soft Error Tolerant Convolutional Neural Networks on FPGAs With Ensemble LearningabstractConvolutional neural networks (CNNs) are widely used in computer vision and natural language processing. Field-programmable gate arrays (FPGAs) are popular accelerators for CNNs. However, if used in critical applications, the reliability of FPGA-based CNNs becomes a priority because FPGAs are prone to suffer soft errors. Traditional protection schemes, such as triple modular redundancy (TMR), introduce a large overhead, which is not acceptable in resource-limited platforms. This article proposes to use an ensemble of weak CNNs to build a robust classifier with low cost. To have a group of base CNNs with low complexity and balanced similarity and diversity, residual neural networks (ResNets) with different layers (20/32/44/56) are combined in the ensemble system to replace a single strong ResNet 110. In addition, a robust combiner is designed based on the reliability evaluation of a single ResNet. Single ResNets with different layers and different ensemble schemes are implemented on the FPGA accelerator based on Xilinx Zynq 7000 SoC. The reliability of the ensemble systems is evaluated based on a large-scale fault injection platform and compared with that of the TMR-protected ResNet 110 and ResNet 20. Experiment results show that the proposed ensembles could effectively improve the system reliability when suffering soft errors with an overhead much lower than TMR. Zhen Gao 0005, Jiajun Xiao, Shulin Zeng, Guangjun Ge, Yu Wang 0002, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2021 | Analyzing and Assessing Pollution Attacks on Bloom Filters: Some Filters are More Vulnerable than OthersabstractBloom filters are probabilistic data structures that are popular in networking for set representation; however, they show an inherent inaccuracy due to false positives. One of the potential attacks on Bloom filters is to pollute them with elements that cause the filter to have a larger false positive probability than under normal operation; Pollution is simple when an attacker knows the details of the filter implementation. Recent research has shown that also black-box adversaries can pollute a counting Bloom filter (a common variant of the filter that also supports removals) with no knowledge of its implementation. As over time, many variants and improvements of Bloom filters have been proposed, it is of interest to study whether they can also be polluted and if so also the increase in their false positive probability. This paper first proposes and then evaluates pollution attacks for some of the most common variants including the Block Bloom filters (BBFs), the Variable Increment and Fingerprint Counting Bloom filters (VI-CBFs and FP-CBFs). The results show that with or without knowledge of the implementation, these variants of the Bloom filter are significantly more vulnerable to pollution attacks than the traditional Bloom filter. In particular, BBFs are extremely vulnerable, so providing an insight on their impact and use in practical systems when the number of memory accesses per lookup must be reduced. Pedro Reviriego, Ori Rottenstreich, Shanshan Liu 0001, Fabrizio Lombardi |
CNSM | 1 |
| 2021 | Perfect cuckoo filtersabstractBloom filters and cuckoo filters are used in many applications to reduce the amount of memory needed to check if an element belongs to a set. The main drawback of these filters is that with low probability, a positive is returned for an element that is not in the set. Recently, the concept of Bloom filters with a false positive free zone has been introduced showing that false positives can be avoided when the universe from which elements are taken and the number of elements inserted in the filter are both small. Unfortunately, this limits the use of such false positive free Bloom filters in many practical applications. In this paper, a false positive free, i.e. perfect, cuckoo filter is presented and evaluated. The proposed design supports universe sizes of billions of elements and stores millions of elements, making it practical for a wide range of applications. The perfect cuckoo filter can be also used to perform mapping, further extending the range of scenarios in which can be used. The benefits of the proposed perfect cuckoo filter are illustrated with two case studies: IP address blacklisting and longest prefix match for IP forwarding. Pedro Reviriego, Salvatore Pontarelli |
CoNEXT | 1 |
| 2021 | Learned Bloom Filters in Adversarial Environments: A Malicious URL Detection Use-CaseabstractLearned Bloom Filters (LBFs) have been recently proposed as an alternative to traditional Bloom filters that can reduce the amount of memory needed to achieve a target false positive probability when representing a given set of elements. LBFs rely on Machine Learning models combined with traditional Bloom filters. However, if LBFs are going to be used as an alternative to Bloom filters, their security must be also be considered. In this paper, the security of LBFs is studied for the first time and a vulnerability different from those of traditional Bloom filters is uncovered. In more detail, an attacker can easily create a set of elements that are not in the filter with a much larger false positive probability than the target for which the filter has been designed. The constructed attack set can then be used to for example launch a denial of service attack against the system that uses the LBF. A malicious URL case study is used to illustrate the proposed attacks and show their effectiveness in increasing the false positive probability of LBFs. The dataset under consideration includes nearly 485K URLs where 16.47% of them are malicious URLs. Unfortunately, it seems that mitigating this vulnerability is not straightforward. Pedro Reviriego, José Alberto Hernández 0001, Zhenwei Dai, Anshumali Shrivastava |
HPSR | 1 |
| 2021 | The Logarithmic Dynamic Cuckoo FilterabstractThe emergence of big data applications makes efficient representation for large-scale dynamic data sets a challenge. The state-of-the-art design, i.e., the dynamic cuckoo filter (DCF), provides extensible approximate set representation by employing a novel chain based data structure which allows appending new building cuckoo filter blocks. However, such a design needs linearly increasing computation costs and memory space when a set scales. This makes it inefficient for big data sets. In this paper, we propose a novel data structure for dynamic big data sets, called logarithmic dynamic cuckoo filter (LDCF). LDCF uses a novel multi-level tree structure and reduces the worst insertion and membership testing times from O(N) to O(1), where N is the size of the set. At the same time, LDCF reduces the memory cost of DCF as the cardinality of the set increases. Comprehensive experiment results show that LDCF significantly reduces the membership checking time and the memory space cost for large-scale datasets compared to state-of-the-art designs. Fan Zhang 0024, Hanhua Chen, Hai Jin 0001, Pedro Reviriego |
ICDE | 4 |
| 2021 | Reliability Evaluation of the Count Min Sketch (CMS) against Single Event Transients (SETs)abstractEstimating the frequency of the elements in a data set is commonly needed in data analysis. With the increase of the size of the data sets, accurately computing the number of times that each element appears with a counter becomes impractical. Instead, the Count Min Sketch (CMS) is widely used in big data processing to estimate frequency due to its simplicity and small storage needs. However, soft errors caused by Single Event Transients (SETs) will affect the hardware implementation of the CMS, mainly the hash functions. In this paper, the effect of SETs on the hash functions of the CMS frequency estimate is analyzed theoretically in terms of overestimation probability, underestimation probability, and the equal probability, and further discussed for data with different frequency. Simulation results verify the correctness of the theoretical analysis and reveal several valuable conclusions. First, a large portion of SETs can be tolerated by the CMS itself, and the reliability of the CMS improves when larger number of arrays are used. Second, the average probability for overestimation and underestimation are almost the same, and decrease for larger numbers of arrays. Third, SETs are more likely to cause underestimation for the most frequent data elements. Finally, the overall effect of SETs on the CMS is slightly affected by the number of counters in each array, and seems to be independent of the distribution of the input sequence. The results and analysis presented in this paper provide a starting point for the design of efficient SET fault-tolerant schemes for the CMS. Zhen Gao 0005, Pedro Reviriego |
VTS | 4 |
| 2021 | Less-is-Better Protection (LBP) for memory errors in kNNs classifiers
Shanshan Liu 0001, Pedro Reviriego, Paolo Montuschi, Fabrizio Lombardi |
Future Gener. Comput. Syst. | 2 |
| 2021 | Designs for efficient low power cardinality and similarity sketches by Two-Step Hashing (TSH)
Jie Li 0030, Pedro Reviriego, Shanshan Liu 0001, Liyi Xiao, Fabrizio Lombardi |
Integr. | 2 |
| 2021 | Soft Error Tolerant Count Min SketchesabstractThe estimation of the frequency of the elements on a set is needed in a wide range of computing applications. For example, to estimate the number of hits that a video gets or the number of packets in a network flow. In some cases, the number of elements in the set is very large and it is not practical to maintain a table with the exact count for each of them. Instead, simpler and more efficient data structures, commonly referred to as sketches, that provide an estimate are used. Among those structures the Count Min Sketch (CMS) is one of the most popular sketches. The CMS provides estimates that have one sided errors. In more detail, the CMS returns an estimate that is equal to or larger than the actual value. An update or check requires a small and constant number of memory accesses and the memory footprint is fixed and does not depend on the number of elements. The CMS relies on several arrays of counters that are stored in memory. Memories are prone to suffer soft errors that flip the contents of memory cells due for example to ionizing radiation. Therefore, it is of interest to study the impact that soft errors can have on the CMS estimates and to propose protection techniques that minimize their effect while requiring low overhead in terms of additional memory and circuitry. To the best of our knowledge this has not been done before. In this article, first the effect of soft errors on the CMS is evaluated by injecting errors. Then, in the second part a protection technique that does not require additional memory bits is presented and compared with the protection using a parity bit. In the last part of the article, the technique is extended to protect also against double adjacent bit errors. Pedro Reviriego, Jorge Martínez 0001, Marco Ottavi |
IEEE Trans. Computers | 1 |
| 2021 | Towards Low Latency and Resource-Efficient FPGA Implementations of the MUSIC Algorithm for Direction of Arrival EstimationabstractThe estimation of the Direction of Arrival (DoA) is one of the most critical parameters for target recognition, identification and classification. MUltiple SIgnal Classification (MUSIC) is a powerful technique for DoA estimation. The algorithm requires complex mathematical operations like the computation of the covariance matrix for the input signals, eigenvalue decomposition and signal peak search. All these signal processing operations make real-time and resource-efficient implementation of the MUSIC algorithm on Field Programmable Gate Arrays (FPGAs) a challenge. In this paper, a novel design approach is proposed for the FPGA-implementation of the MUSIC algorithm. This approach enables a significant reduction in both FPGA resources and latency. In more detail, the proposed design enables the estimation of DoA in real-time scenarios in 2μsec with 30% to 50% fewer resources as compared to existing techniques. Uzma M. Butt, Shoab A. Khan, Anees Ullah, Abdul Khaliq, Pedro Reviriego, Ali Zahir |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Stochastic Dividers for Low Latency Neural NetworksabstractDue to the low complexity in arithmetic unit design, stochastic computing (SC) has attracted considerable interest to implement Artificial Neural Networks (ANNs) for resources-limited applications, because ANNs must usually perform a large number of arithmetic operations. To attain a high computation accuracy in an SC-based ANN, extended stochastic logic is utilized together with standard SC units and thus, a stochastic divider is required to perform the conversion between these logic representations. However, the conventional divider incurs in a large computation latency, so limits an SC implementation for ANNs used in applications needing high performance. Therefore, there is a need to design fast stochastic dividers for SC-based ANNs. Recent works (e.g., a binary searching and triple modular redundancy (BS-TMR) based stochastic divider) are targeting a reduction in computation latency, while keeping the same accuracy compared with the traditional design. However, this divider still requires$N$iterations to deal with$2^{N}$-bit stochastic sequences, and thus the latency increases in proportion to the sequence length. In this paper, a decimal searching and TMR (DS-TMR) based stochastic divider is initially proposed to further reduce the computation latency; it only requires two iterations to calculate the quotient, so regardless of the sequence length. Moreover, a trade-off design between accuracy and hardware is also presented. An SC-based Multi-Layer Perceptron (MLP) is then considered to show the effectiveness of the proposed dividers over current designs. Results show that when utilizing the proposed dividers, the MLP achieves the lowest computation latency while keeping the same classification accuracy; although incurring in an area increase, the overhead due to the proposed dividers is low over the entire MLP. When using as combined metric for both hardware design and computation complexity the product of the implementation area, latency, power and number of clock cycles, the proposed designs are also shown to be superior to the SC-based MLPs (at the same level of accuracy) employing other dividers found in the technical literature as well as the commonly used 32-bit floating point implementation. Shanshan Liu 0001, Xiaochen Tang, Farzad Niknia, Pedro Reviriego, Weiqiang Liu 0001, Ahmed Louri, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Avoiding Flow Size Overestimation in Count-Min Sketch With Bloom Filter ConstructionsabstractThe Count-Min sketch is the most popular data structure for flow size estimation, a basic measurement task required in many networks. Typically the number of potential flows is large, eliminating the possibility to maintain a counter per flow within memory of high access rate. The Count-Min sketch is probabilistic and relies on mapping each flow to multiple counters through hashing. This implies potential estimation error such that the size of a flow is overestimated when all flow counters are shared with other flows with observed traffic. Although the error in the estimation can be probabilistically bounded, many applications can benefit from accurate flow size estimation and the guarantee to completely avoid overestimation. We describe a design of the Count-Min sketch with accurate estimations whenever the number of flows with observed traffic follows a known bound, regardless of the identity of these particular flows. We make use of a concept of Bloom filters that avoid false positives and indicate the limitations of existing Bloom filter designs towards accurate size estimation. We suggest new Bloom filter constructions that allow scalability with the support for a larger number of flows and explain how these can imply the unique guarantee of accurate flow size estimation in the well known Count-Min sketch. Ori Rottenstreich, Pedro Reviriego, Ely Porat, S. Muthukrishnan 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | Design of FPGA-Implemented Reed-Solomon Erasure Code (RS-EC) Decoders With Fault Detection and Location on User MemoryabstractReed-Solomon erasure codes (RS-ECs) are widely used in packet communication and storage systems to recover erasures. When the RS-EC decoder is implemented on a field-programmable gate array (FPGA) in a space platform, it will suffer single-event upsets (SEUs) that can cause failures. In this article, the reliability of an RS-EC decoder implemented on an FPGA when there are errors in the user memory is first studied. Then, a fault detection and location scheme is proposed based on partial reencoding for the faults in the user memory of the RS-EC decoder. Furthermore, check bits are added in the generator matrix to improve the fault location performance. The theoretical analysis shows that the scheme could detect most faults with small missing and false detection probability. Experimental results on a case study show that more than 90% of the faults on user memory could be tolerated by the decoder, and all the other faults can be detected by the fault detection scheme when the number of erasures is smaller than the correction capability of the code. Although false alarms exist (with probability smaller than 4%), they can be used to avoid fault accumulation. Finally, the fault location scheme could accurately locate all the faults. The theoretical estimates are very close to the experiment results, which verifies the correctness of the analysis done. Zhen Gao 0005, Yinghao Cheng, Kangkang Guo, Anees Ullah, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | Pollution Attacks on Counting Bloom Filters for Black Box AdversariesabstractThe wide adoption of Bloom filters makes their security an important issue to be addressed. For example, an attacker can increase their error rate through polluting and eventually saturating the filter by inserting elements that set to one a large number of positions in the filter. This is known as a pollution attack and requires that the attacker knows the hash functions used to construct the filter. Such information is not available in many practical settings and in addition a simple protection can be achieved through using a random salt in the hash functions. The same pollution attacks can also be done to counting Bloom filters that in addition to insertions and lookups support removals. This paper considers pollution attacks on counting Bloom filters. We describe two novel pollution attacks that do not require any knowledge of the counting Bloom filter implementation details and evaluate them. These methods show that a counting Bloom filter is vulnerable to pollution attacks even when the attacker has only access to the filter as a black box to perform insertions, removals, and lookups. Pedro Reviriego, Ori Rottenstreich |
CNSM | 1 |
| 2020 | When filtering is not possible caching negatives with fingerprints comes to the rescueabstractBloom filters are widely used in networking to accelerate checks and in particular to avoid accessing slow memories when a match will not be found. Unfortunately, filtering requires several on-chip memory bits per element and thus when the tables are large and the on-chip memory small is not applicable. In those cases, caching the most frequently accessed elements on-chip seems the only viable option. However, the key is typically formed by several packet header fields, which means that each cache entry consumes a significant amount of on-chip memory bits. In this paper, an efficient scheme to cache negatives, that is elements that will not find a match on the table is presented. In more detail, the proposed scheme enables the caching of negatives using less than 16 bits per entry regardless of the size of the key. This translates to a reduction of at least 6x in the size of the cache when used for example flow tracking. Pedro Reviriego, Salvatore Pontarelli |
CoNEXT | 1 |
| 2020 | Reliability Evaluation of Turbo Decoders Implemented on SRAM-FPGAsabstractTurbo codes are widely used in satellite communications. When a Turbo decoder is implemented on a Field Programmable Gate Array (FPGA) in a space platform, it will suffer Single Event Upsets (SEUs) that can cause failures and disrupt communications. In this paper, the reliability of Turbo decoders implemented on FPGAs is evaluated. The Turbo decoder with Log-MAP algorithm is implemented on an SRAM-FPGA. Then, fault injection experiments are conducted to simulate the effects of SEU on the user memory and on the configuration memory of the Turbo decoder. Experimental results show that, for user memory, the SEU tolerance rate is over 95%, and the effect of SEU is related to the iteration period, bit position and Signal to Noise Ratio (SNR). In particular, SEUs on the control/address registers and on the interleaving table have a larger impact than on other registers or memories. For the configuration memory, the SEU tolerance rate is higher than 86%, and decreases as SNR increases. In general, the Turbo decoder exhibits a high reliability against SEUs, and the user memory is more reliable than the configuration memory. Zhen Gao 0001, Ruishi Han, Pedro Reviriego |
VTS | 4 |
| 2020 | Cuckoo Filters and Bloom Filters: Comparison and Application to Packet ClassificationabstractBloom filters are used to perform approximate membership checking in a wide range of applications in both computing and networking, but the recently introduced cuckoo filter is also gaining popularity. Therefore, it is of interest to compare both filters and provide insights into their features so that designers can make an informed decision when implementing approximate membership checking in a given application. This article first compares Bloom and cuckoo filters focusing on a packet classification application. The analysis identifies a shortcoming of cuckoo filters in terms of false positive rate when they do not operate close to full occupancy. Based on that observation, this article also proposes the use of a configurable bucket to improve the scaling of the false positive rate of the cuckoo filter with occupancy. Pedro Reviriego, Jorge Martínez 0001, David Larrabeiti, Salvatore Pontarelli |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | Design of SEU-Tolerant Turbo Decoders Implemented on SRAM-FPGAsabstractTurbo codes are widely used in satellite communications. When a turbo decoder is implemented on a field-programmable gate array (FPGA) in a space platform, it will suffer single-event upsets (SEUs) that can cause failures and disrupt communications. Therefore, the protection of turbo decoders implemented on FPGAs is important. In this article, first the reliability of an SRAM-FPGA-implemented turbo decoder to SEUs on user memory and configuration memory is evaluated based on fault injection experiments. Then, based on the features of the turbo decoder and the characteristics of the failures revealed by the reliability study, a duplication with comparison (DWC) scheme is proposed for the protection of the turbo decoder. Experimental results show that the reliability of the protected turbo decoder to SEUs on user memory and configuration memory is improved by 99.4% and 95.6%, respectively. The resource usage is about 2.2× that of an unprotected turbo decoder, which is significantly lower than the more than 3× required by the traditional triple modular redundancy (TMR) protection. Finally, the proposed scheme is compared with another two protection schemes. Zhen Gao 0005, Tong Yan, Kangkang Guo, Pedro Reviriego |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | FracTCAM: Fracturable LUTRAM-Based TCAM Emulation on Xilinx FPGAsabstractIn this brief, we present FracTCAM, an efficient methodology for ternary content addressable memory (TCAM) emulation on Xilinx field-programmable gate arrays (FPGAs) by leveraging primitive architectural resources. The proposed methodology exploits the fracturable nature of lookup table random access memories (LUTRAMs) and built-in slice flip-flops for deeper pipelining. Multiple slices can be combined together to build deeper and wider TCAMs using ANDing operations. This results in TCAM implementations that achieve lower resources utilization, lower delay, and power consumption. A comparison with the existing schemes shows that FracTCAM consistently achieves the best performance per area (PA) and performance per area per watt (PAW). Ali Zahir, Shadan Khan Khattak, Anees Ullah, Pedro Reviriego, Fahad Bin Muslim, Waleed Ahmad |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | A Radiation Tolerant 10/100 Ethernet Transceiver for Space ApplicationsabstractAs space systems evolve to become more complex, they need larger computing and communication capabilities. For example, larger data rates must be supported and more flexible technologies that enable several applications to share the network resources while providing predictable and reliable performance are needed. One of the approaches to address those issues is the adoption of Ethernet standards in space. This has the benefit of reusing existing and field proven technology that also provides evolution to larger data rates. Ethernet is currently used in some space systems and it is being part of the implementation roadmap of many others, such as the next generation of Ariane launchers. Integrated circuits used in space systems must be designed to withstand the effects of radiation that causes errors and failures. These devices, known as rad-hard devices, need to be designed and manufactured using specific techniques and processes. In order to adopt the use of Ethernet in space, corresponding radhard Integrated Circuits (ICs) need to be available. The European industry is working on several such ICs, including an Ethernet switch and a physical layer transceiver. In this paper, SEPHY a 10/100 Mb/s rad-hard Ethernet transceiver designed for space applications is presented. Anselm Breitenreiter, Jesús López, Pedro Reviriego, Milos Krstic, Úrsula Gutierro, Manuel Sanchez-Renedo, Daniel González |
IOLTS | 3 |
| 2019 | Selective Fault Tolerance by Counting Gates with Controlling ValueabstractThe protection of flip-flops against soft errors in digital circuits incurs significant overheads. To reduce the protection costs, it is common to identify the flip-flops in which errors can produce an effect on the system output or a persistent error in its state. Then, only those critical flip-flops are protected. To identify those flip-flops, one option is to perform fault injection on all the flips flops during functional simulations but this does not scale well for large circuits as the time required to perform an evaluation would not be practical. Another option is to perform the identification based only on the structural properties of the circuit. For example, flip-flops that are in loops are more likely to produce persistent errors and those closer to the system outputs to affect them. In this paper, an enhancement to the structural analysis is proposed to improve its accuracy. The idea is to identify gates that can mask error propagation as they have a controlling value and use only those to measure distances on the circuit to estimate the criticality of flip-flops. This enables us to better predict the effect of errors while keeping the analysis simple and based only on structural properties and gate types. The proposed scheme has been implemented and tested on a realistic circuit to show its effectiveness. Anselm Breitenreiter, Stefan Weidling, Oliver Schrape, Steffen Zeidler 0001, Pedro Reviriego, Milos Krstic |
IOLTS | 5 |
| 2019 | Efficient Concurrent Error Detection for SEC-DAEC EncodersabstractIn the last decade, a number of Single Error Correction Double Adjacent Error Correction (SEC-DAEC) codes have been proposed to protect memories against Multiple Cell Upsets (MCUs). These codes are able to correct errors that affect two adjacent bits that is one of the most common MCU patterns. However, soft errors can also affect the encoder and decoder circuitry creating data corruption. An alternative to protect the encoders is to use parity prediction Concurrent Error Detection (CED) to detect errors and avoid writing erroneous words in the memory. This approach has been previously studied for Orthogonal Latin Square (OLS) codes and for matrix codes. In this paper, the implementation of parity prediction Concurrent Error Detection (CED) for SEC-DAEC codes is considered. To that end, first it is shown that CED has a significant cost for the existing SEC-DAEC codes. This is because they are odd weight codes and parity prediction is much simpler for even weight codes. Based on that observation, even weight SEC-DAEC codes are designed and evaluated. The results show that CED can be efficiently implemented in the proposed codes that achieve a significant reduction in encoder circuit complexity compared to previously proposed SEC-DAEC codes. Jiaqiang Li, Pedro Reviriego, Costas Argyrides, Liyi Xiao |
IOLTS | 2 |
| 2019 | Low Delay 3-Bit Burst Error Correction Codes
Jiaqiang Li, Pedro Reviriego, Liyi Xiao |
J. Electron. Test. | 2 |
| 2019 | CuCoTrack: Cuckoo filter based connection tracking
Pedro Reviriego, Salvatore Pontarelli, Gil Levy |
Inf. Process. Lett. | 1 |
| 2019 | Efficient Implementations of Reduced Precision Redundancy (RPR) Multiply and Accumulate (MAC)abstractMultiply and Accumulate (MAC) is one of the most common operations in modern computing systems. It is for example used in matrix multiplication and in new computational environments such as those executed on neural networks for deep machine learning. MAC is also used in critical systems that must operate reliably such as object recognition for vehicles. Therefore, MAC implementations must be able to cope with errors that may be caused for example by radiation. A common scheme to deal with soft errors in arithmetic circuits is the use of Reduced Precision Redundancy (RPR). RPR instead of replicating the entire circuit, uses reduced precision copies which significantly reduce the overhead while still being able to correct the largest errors. This paper considers the implementation of RPR Multiply and Accumulate circuits. First, it is shown that the properties of signed integer multiplication (two´s complement format) can be used to make RPR more efficient. Then its principles are extended to the MAC operation by proposing RPR implementations that improve the error correction capabilities with a limited impact on circuit overhead. The proposed schemes have been implemented and tested. The results show that they can significantly reduce the Mean Square Error (MSE) at the output when the circuit is affected by a soft error and the implementation overhead of the proposed schemes is extremely low. Ke Chen 0018, Linbin Chen, Pedro Reviriego, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2019 | An ALU Protection Methodology for Soft Processors on SRAM-Based FPGAsabstractThe use of microprocessors in space missions implies that they should be protected against the effects of cosmic radiation. Commonly this objective has been achieved by applying modular redundancy techniques which provide good results in terms of reliability but increase significantly the number of used resources. Because of that, new protection techniques have appeared, trying to establish a trade-off between reliability and resource utilization. In this paper, we propose an application-based methodology, to protect a soft processor implemented in an SRAM-based FPGA, against the effect of soft errors. This is done creating a library of adaptive protection configurations, based on the profiling of the application. This hardware configuration library, combined with the reprogramming capabilities of the FPGA, helps to create an adaptive protection for each application. We propose two partial TMR configurations for the Arithmetic Logic Unit (ALU) as an example of this methodology. The proposed scheme has been tested in a RISC-V soft processor. A fault injection campaign has been performed to test its reliability. Alexis Ramos, Ricardo Gonzalez-Toral, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Computers | 3 |
| 2019 | Two Bit Overlap: A Class of Double Error Correction One Step Majority Logic Decodable CodesabstractError Correction Codes (ECCs) are commonly used to protect memories against soft errors with an impact on memory area and delay. For large memories, the area overhead is mostly due to the additional cells needed to store the parity check bits. In terms of delay, the overhead is mostly needed to detect and correct errors when the data is read from the memory. Most ECCs that can correct more than one error have a complex decoding process and so are limited in high speed memory applications. One exception is One Step Majority Logic Decodable (OS-MLD) codes for which decoding can be done in parallel at high speed. Unfortunately, there are only a few OS-MLD codes that provide a limited choice in terms of block sizes, error correction capabilities and code rate. Therefore, there is considerable interest in a novel construction of OS-MLD codes to provide additional choices for protecting memories. In this paper, a new method to construct Double Error Correction (DEC) OS-MLD codes is presented. This method is based on the use of parity check matrices in which two bits have at most two parity check equations in common; the proposed method provides codes that require a smaller number of parity check bits than existing codes like Orthogonal Latin Square (OLS) codes. The drawback of the proposed Two Bit Overlap (TBO) codes is that they require slightly more complex decoding than OLS codes. Therefore, they provide an intermediate solution between OLS and non OS-MLD codes in terms of decoding delay and number of parity check bits. The proposed TBO codes have been implemented for some block sizes and compared to both OLS and BCH codes to illustrate the trade off in delay and memory overhead. Finally, this paper discusses the generalization of the proposed scheme to codes with larger error correction capabilities. Pedro Reviriego, Shanshan Liu 0001, Ori Rottenstreich, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2019 | Enhancing Instruction TLB Resilience to Soft ErrorsabstractA translation lookaside buffer (TLB) is a type of cache used to speed up the virtual to physical memory translation process. Instruction TLBs store virtual page numbers and their related physical page numbers for the last accessed pages of instruction memory. TLBs like other memories suffer soft errors that can corrupt their contents. A false positive due to an error produced in the virtual page number stored in the TLB may lead to a wrong translation and, consequently, the execution of a wrong instruction that can lead to a program hard fault or to data corruption. Parity or error correction codes have been proposed to provide protection for the TLB, but they require additional storage space. This paper presents some schemes to increase the instruction TLB resilience to this type of errors without requiring any extra storage space, by taking advantage of the spatial locality principle that takes place when executing a program. Alfonso Sánchez-Macián, Luis Alberto Aranda, Pedro Reviriego, Vahdaneh Kiani, Juan Antonio Maestro |
IEEE Trans. Computers | 3 |
| 2019 | The Tandem Counting Bloom Filter - It Takes Two Counters to TangoabstractSet representation is a crucial functionality in various areas such as networking and databases. In many applications, memory and time constraints allow only an approximate representation where errors can appear for some queried elements. The Variable-Increment Counting Bloom Filter (VI-CBF) is a popular data structure for the representation of dynamically-changing sets, achieving a good tradeoff between memory efficiency and queries accuracy. For some applications, the required accuracy is higher than that enabled by the VI-CBF. In this paper, we present the Tandem Counting Bloom Filter (T-CBF), a new data structure that relies on the interaction among counters to describe sets with higher accuracy. We analyze its performance and show that by a joint consideration of counters, the T-CBF always performs better than the VI-CBF and it can for some configurations reduce its false positive probability by an order of magnitude. The overhead of such an approach is expressed upon an element insertion or query as read or write operations to a pair of counters rather than a single counter in each hash location. The operations themselves also require considering a larger number of scenarios. Pedro Reviriego, Ori Rottenstreich |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Error Detection and Correction in SRAM Emulated TCAMsabstractTernary content addressable memories (TCAMs) are widely used in network devices to implement packet classification. They are used, for example, for packet forwarding, for security, and to implement software-defined networks (SDNs). TCAMs are commonly implemented as standalone devices or as an intellectual property block that is integrated on networking application-specific integrated circuits. On the other hand, field-programmable gate arrays (FPGAs) do not include TCAM blocks. However, the flexibility of FPGAs makes them attractive for SDN implementations, and most FPGA vendors provide development kits for SDN. Those need to support TCAM functionality and, therefore, there is a need to emulate TCAMs using the logic blocks available in the FPGA. In recent years, a number of schemes to emulate TCAMs on FPGAs have been proposed. Some of them take advantage of the large number of memory blocks available inside modern FPGAs to use them to implement TCAMs. A problem when using memories is that they can be affected by soft errors that corrupt the stored bits. The memories can be protected with a parity check to detect errors or with an error correction code to correct them, but this requires additional memory bits per word. In this brief, the protection of the memories used to emulate TCAMs is considered. In particular, it is shown that by exploiting the fact that only a subset of the possible memory contents are valid, most single-bit errors can be corrected when the memories are protected with a parity bit. Pedro Reviriego, Salvatore Pontarelli, Anees Ullah |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | PR-TCAM: Efficient TCAM Emulation on Xilinx FPGAs Using Partial ReconfigurationabstractModern field-programmable gate arrays (FPGAs) provide a vast amount of logic resources that can be used to implement complex systems while providing the flexibility to modify the design once deployed. This makes them attractive for software-defined networks (SDNs) applications, and, in fact, most vendors provide the building blocks needed for those applications, which include basic packet classification functions such as exact match, longest prefix match, and match with wildcards. Those are needed for different functions such as routing, security filtering, monitoring or quality of service. The match with wildcards can be done using ternary content addressable memories (TCAMs). TCAMs can be implemented as independent standalone devices or as Internet Protocol (IP) blocks that are used inside networking application-specific integrated circuits (ASICs) such as switching ICs. In both cases, the cells of a TCAM are more complex than that of a normal memory and also than that of a binary content addressable memory (CAMs). This is due to the more complex matching that they need to implement. As FPGAs are used in many different applications, it does not make sense to include TCAM blocks inside them as they would be used only in a small fraction of the systems. Therefore, TCAMs are emulated using the logic resources available inside the FPGA. In recent years, a number of schemes to emulate TCAMs on FPGAs have been proposed, some of them based on the use of the logic resources and others on the use of the embedded memory blocks available on the FPGA. In this brief, a technique to efficiently emulate TCAMs on Xilinx FPGAs is presented. The proposed scheme is based on the use of lookup tables (LUTs) and partial reconfiguration to achieve a more effective use of the FPGA resources while supporting the addition and removal of rules. The proposed scheme has been compared to existing implementations and the results show that it can achieve significant savings in resource usage. In addition, it enables the use of all the LUTs in the device for TCAM implementation, something that is not supported by existing approaches that use LUTRAMs. Pedro Reviriego, Anees Ullah, Salvatore Pontarelli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Adaptive Cuckoo FiltersabstractWe introduce the adaptive cuckoo filter (ACF), a data structure for approximate set membership that extends cuckoo filters by reacting to false positives, removing them for future queries. As an example application, in packet processing queries may correspond to flow identifiers, so a search for an element is likely to be followed by repeated searches for that element. Removing false positives can therefore significantly lower the false positive rate. The ACF, like the cuckoo filter, uses a cuckoo hash table to store fingerprints. We allow fingerprint entries to be changed in response to a false positive in a manner designed to minimize the effect on the performance of the filter. We show that the ACF is able to significantly reduce the false positive rate by presenting both a theoretical model for the false positive rate and simulations using both synthetic data sets and real packet traces. Michael Mitzenmacher, Salvatore Pontarelli, Pedro Reviriego |
ALENEX | 3 |
| 2018 | Position-aware cuckoo filtersabstractCuckoo filters have been recently proposed as an efficient structure to perform approximate membership checks. In this paper, it is shown that the false positive rate of a cuckoo filter can be reduced by storing in addition to the fingerprint the information that tell us if a given fingerprint has been inserted in the first or the second bucket. This improvement of the cuckoo filter is denoted as Position Aware (PA) cuckoo filter. Minseok Kwon, Vijay Shankar, Pedro Reviriego |
ANCS | 3 |
| 2018 | Multiple Hash Matching Units (MHMU): An Algorithmic Ternary Content Addressable Memory Design for Field Programmable Gate ArraysabstractAs applications and user requirements are constantly evolving, there is a need to provide flexible networks that are able to process packets at high speed. One of the basic functions used for packet processing is matching a key formed by some fields of the incoming packet header against a set of stored rules. This is done for example to determine the next hop of a packet or to apply security checks on a firewall. In many cases, the stored rules have do not care bits as that enables a more flexible and compact representation of the rules. Therefore, the matching can be done in hardware using Ternary Content Addressable Memories (TCAMs). However, TCAMs pose several problems in many implementations. For example, for ASICs they require much more circuit area and power than standard SRAMs. On the other hand, designs based on programmable logic such as Field Programmable Gate Arrays (FPGAs) can only use the blocks provided by the FPGA that do not typically include TCAMs. In this last case, a TCAM can be emulated using the FPGA logic resources but with a large cost. To reduce the cost of implementing TCAMs, a number of algorithmic solutions have been proposed and are known as Algorithmic TCAMs or A-TCAMs. Most of those schemes target either software or ASIC implementations. In this paper we present Multiple Hash Matching Units (MHMU) an A-TCAM solution targeted towards FPGA implementations. The proposed scheme exploits the massive parallelism of FPGAs to implement many hash based matching units that use the embedded block RAM memories of the FPGA. The proposed MHMU scheme has been mapped to a Xilinx series 7 FPGA to check its efficiency in terms of resource usage and its scalability. To validate the effectiveness of MHMU, a simple configuration has been tested with Classbench generated sets of rules. The results show that the MHMU is able to consistently accommodate sets with several tens of thousands of rules with large keys. Pedro Reviriego, Salvatore Pontarelli, Anees Ullah, Ali Zahir, Giuseppe Bianchi 0001 |
HPSR | 1 |
| 2018 | A Scheme to Design Concurrent Error Detection Techniques for the Fast Fourier Transform Implemented in SRAM-Based FPGAsabstractSoft errors are an important issue for SRAM-based Field Programmable Gate Arrays (FPGAs), since they result in permanent alterations of the mapped circuit when they affect their configuration memory. Concurrent Error Detection (CED) techniques, such as Dual Modular Redundancy (DMR), are usually employed to detect errors that affect the performance of the circuit. When trying to detect errors produced on the complex Fast Fourier Transform (FFT), the Parseval Sum of Squares (SoS) is a widely used technique. In this paper, we present a scheme to implement CED techniques for the complex FFT implemented in SRAM-based FPGAs. These techniques perform checks based on the relationships existing between one or more of the inputs and the outputs of the algorithm. Three examples of these techniques are provided to further clarify how to construct them. These techniques, along with DMR and SoS, have been tested through fault injection. An analysis on their error detection capabilities shows that they achieve high detection rates with much less resource usage than DMR and SoS. In addition, the number of false error detections for these techniques is lower than that of SoS, which leads to less unnecessary reconfigurations of the device. Ricardo Gonzalez-Toral, Pedro Reviriego, Juan Antonio Maestro, Zhen Gao 0005 |
IEEE Trans. Computers | 2 |
| 2018 | Efficient Protection of the Register File in Soft-Processors Implemented on Xilinx FPGAsabstractSoft-processors implemented on SRAM-based FPGAs are increasingly being adopted in on-board computing for space and avionics applications due to their flexibility and ease of integration. However, efficient component-level protection techniques for these processors against radiation-induced upsets are necessary otherwise as system failures could manifest. A register file is one of the critical structures that stores vital information the processor uses related to user computations and program execution. In this paper, we present a fault tolerance technique for the register file of a microprocessor implemented in Xilinx SRAM-based FPGAs. The proposed scheme leverages the inherent implementation redundancy created by the FPGA design automation tools when mapping the register file to on-chip distributed memory. A parity-based error detection and switching logic are added for fault masking against single-bit errors. The proposed scheme has been implemented and evaluated in lowRISC, a RISC-V ISA soft-processor implementation. The effectiveness of the proposed scheme was tested using fault injection. The fault masking overhead required in terms of FPGA resources was much lower than a traditional Triple Modular Redundancy protection. Therefore, the proposed scheme is an interesting option to protect the register file of soft processors that are implemented in Xilinx FPGAs. Alexis Ramos, Anees Ullah, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Computers | 3 |
| 2018 | EMOMA: Exact Match in One Memory AccessabstractAn important function in modern routers and switches is to perform a lookup for a key. Hash-based methods, and in particular cuckoo hash tables, are popular for such lookup operations, but for large structures stored in off-chip memory, such methods have the downside that they may require more than one off-chip memory access to perform the key lookup. Although the number of off-chip memory accesses can be reduced using on-chip approximate membership structures such as Bloom filters, some lookups may still require more than one off-chip memory access. This can be problematic for some hardware implementations, as having only a single off-chip memory access enables a predictable processing of lookups and avoids the need to queue pending requests. We provide a data structure for hash-based lookups based on cuckoo hashing that uses only one off-chip memory access per lookup, by utilizing an on-chip pre-filter to determine which of multiple locations holds a key. We make particular use of the flexibility to move elements within a cuckoo hash table to ensure the pre-filter always gives the correct response. While this requires a slightly more complex insertion procedure and some additional memory accesses during insertions, it is suitable for most packet processing applications where key lookups are much more frequent than insertions. An important feature of our approach is its simplicity. Our approach is based on simple logic that can be easily implemented in hardware, and hardware implementations would benefit most from the single off-chip memory access per lookup. Salvatore Pontarelli, Pedro Reviriego, Michael Mitzenmacher |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | An Efficient Fault-Tolerance Design for Integer Parallel Matrix-Vector MultiplicationsabstractParallel matrix processing is a typical operation in many systems, and in particular matrix-vector multiplication (MVM) is one of the most common operations in the modern digital signal processing and digital communication systems. This paper proposes a fault-tolerant design for integer parallel MVMs. The scheme combines ideas from error correction codes with the self-checking capability of MVM. Field-programmable gate array evaluation shows that the proposed scheme can significantly reduce the overheads compared to the protection of each MVM on its own. Therefore, the proposed technique can be used to reduce the cost of providing fault tolerance in practical implementations. Zhen Gao 0005, Qingqing Jing, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Extending 3-bit Burst Error-Correction Codes With Quadruple Adjacent Error CorrectionabstractThe use of error-correction codes (ECCs) with advanced correction capability is a common system-level strategy to harden the memory against multiple bit upsets (MBUs). Therefore, the construction of ECCs with advanced error correction and low redundancy has become an important problem, especially for adjacent ECCs. Existing codes for mitigating MBUs mainly focus on the correction of up to 3-bit burst errors. As the technology scales and cell interval distance decrease, the number of affected bits can easily extend to more than 3 bit. The previous methods are therefore not enough to satisfy the reliability requirement of the applications in harsh environments. In this paper, a technique to extend 3-bit burst error-correction (BEC) codes with quadruple adjacent error correction (QAEC) is presented. First, the design rules are specified and then a searching algorithm is developed to find the codes that comply with those rules. The ${H}$ matrices of the 3-bit BEC with QAEC obtained are presented. They do not require additional parity check bits compared with a 3-bit BEC code. By applying the new algorithm to previous 3-bit BEC codes, the performance of 3-bit BEC is also remarkably improved. The encoding and decoding procedure of the proposed codes is illustrated with an example. Then, the encoders and decoders are implemented using a 65-nm library and the results show that our codes have moderate total area and delay overhead to achieve the correction ability extension. Jiaqiang Li, Pedro Reviriego, Liyi Xiao, Costas Argyrides, Jie Li 0030 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Combined Modular Key and Data Error Protection for Content-Addressable MemoriesabstractContent-addressable memories (CAMs) are a type of memory that receives an input search key and compares it to every entry of a table of stored keys. If there is a match, they return the corresponding address where the value was found. Alternatively, they can include a related Random Access Memory (RAM) that is accessed with the matching address, returning the corresponding data values. To protect a CAM with associated RAM against errors, parity or error-correction codes (ECCs) are typically used. They usually protect the CAM and the RAM information separately incurring in additional storage needs. This paper proposes a scheme to protect some configurations of CAM with its associated RAM from errors with a single ECC code. This ECC code can be used to provide advanced error correction to the combination of the key and data values stored in the CAM and the RAM, but it can also be applied in a modular way to provide simpler protection to the key or to the values individually. Alfonso Sánchez-Macián, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Computers | 2 |
| 2017 | Single Event Transient Tolerant Bloom Filter ImplementationsabstractBloom filters have been used to reduce the delay in networking and computing applications when a set membership check is to be applied. Error sources can affect the behavior of Bloom filters resulting in a wrong outcome of this membership test and a possible effect in the system's output. Single event transients are a type of temporary errors altering the operation of combinational logic. A single event transient affecting the hash generation logic of a hardware-implemented Bloom filter can produce errors such as false negatives. This paper presents different approaches to build Bloom filters that are tolerant to single event transients occurring in the hash generation circuitry. They are compared to the use of traditional Modular Redundancy approaches. The results show that the new schemes can reduce significantly the circuit area needed to implement the Bloom filter. Alfonso Sánchez-Macián, Pedro Reviriego, Juan Antonio Maestro, Shanshan Liu 0001 |
IEEE Trans. Computers | 2 |
| 2017 | A Scheme to Reduce the Number of Parity Check Bits in Orthogonal Latin Square CodesabstractThe use of error-correcting codes is a common strategy to protect memories from errors. Single-error correction, double-error detection linear block codes have been traditionally utilized. However, there are applications where multiple errors are frequent and more complex codes are needed. Orthogonal Latin square codes are one type of codes with multiple-error-correction capability. They are of interest for memory protection because they can be decoded with low complexity and delay. This paper presents a modification to orthogonal Latin square codes that reduces the number of parity check bits to be stored in memory therefore lowering the memory overhead needed to implement the codes. The proposed codes can also be decoded with low delay and complexity. This paper also presents an evaluation of the encoder and decoder implementations for various word sizes and compares them with the standard orthogonal Latin square implementations. The results show that they are similar in terms of circuit area and introduce only a small penalty in delay. Pedro Reviriego, Shanshan Liu 0001, Alfonso Sánchez-Macián, Liyi Xiao, Juan Antonio Maestro |
IEEE Trans. Reliab. | 1 |
| 2016 | Efficient fault tolerant parallel matrix-vector multiplicationsabstractParallel matrix processing is a typical operation in many systems, and in particular matrix-vector multiplication is one of the most common operations in modern digital signal processing and digital communication systems. This paper proposes a fault tolerant design for parallel matrix-vector multiplications. The scheme combines ideas from Error Correction Codes with the self-checking capability of matrix-vector multiplication. Zhen Gao 0001, Pedro Reviriego, Juan Antonio Maestro |
IOLTS | 2 |
| 2016 | Improving counting Bloom filter performance with fingerprints
Salvatore Pontarelli, Pedro Reviriego, Juan Antonio Maestro |
Inf. Process. Lett. | 2 |
| 2016 | Unequal Error Protection Codes Derived from Double Error Correction Orthogonal Latin Square CodesabstractIn recent years, there has been a growing interest in multi-bit error correction codes (ECCs) to protect SRAM memories. This has been caused by the increased number of multiple errors that memories suffer as technology scales. To be suitable to protect an SRAM memory, an ECC has to be decodable in parallel and with low latency. Among the codes proposed for memory protection are orthogonal latin square (OLS) codes that provide low latency decoding and a modular construction. For some applications, like multimedia or signal processing, the effect of errors on the memory bits can be very different depending on their position on the word. Therefore, in these cases, it is more effective to provide different degrees of error correction for the different bits. This is done with unequal error protection (UEP) codes. In this paper, UEP codes are derived from double error correction (DEC) OLS codes. The derived codes are implemented for an FPGA platform to evaluate the decoder complexity and latency. The results show that the new codes can be implemented with lower decoding delay than traditional SEC-DED codes and with a cost similar to that of both DEC OLS and SEC-DED codes. Mustafa Demirci, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Computers | 2 |
| 2016 | Parallel d-Pipeline: A Cuckoo Hashing Implementation for Increased ThroughputabstractCuckoo hashing has proven to be an efficient option to implement exact matching in networking applications. It provides good memory utilization and deterministic worst case access time. The continuous increase in speed and complexity of networking devices creates a need for higher throughput exact matching in many applications. In this paper, a new Cuckoo hashing implementation named parallel d-pipeline is proposed to increase throughput. The scheme presented is targeted to implementations in which the tables are accessed in parallel. A parallel implementation increases the throughput and therefore is well suited to high speed applications. Parallel schemes are common for ASIC/FPGA implementations in which the tables are stored in several embedded memories. Using the proposed technique, the throughput can be significantly increased with gains that in practical scenarios can reach 60 percent compared to existing parallel implementations. The new scheme has been evaluated using a case study and detailed results for performance and implementation costs are reported. Salvatore Pontarelli, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Computers | 2 |
| 2016 | OMASS: One Memory Access Set SeparationabstractIn many applications, there is a need to identify to which of a group of sets an element x belongs, if any. For example, in a router, this functionality can be used to determine the next hop of an incoming packet. This problem is generally known as set separation and has been widely studied. Most existing solutions make use of hash-based algorithms, particularly when a small percentage of false positives is allowed. A known approach is to use a collection of Bloom filters in parallel. Such schemes can require several memory accesses, a significant limitation for some implementations. We propose an approach using Block Bloom Filters, where each element is first hashed to a single memory block that stores a small Bloom filter that tracks the element and the set or sets the element belongs to. In a naive solution, when an element x in a set S is stored, it necessarily increases the false positive probability for finding that x is in another set T. In this paper, we introduce our One Memory Access Set Separation (OMASS) scheme to avoid this problem. OMASS is designed so that for a given element x, the corresponding Bloom filter bits for each set map to different positions in the memory word. This ensures that the false positive rates for the Bloom filters for element x under other sets are not affected. In addition, OMASS requires fewer hash functions compared to the naive solution. Michael Mitzenmacher, Pedro Reviriego, Salvatore Pontarelli |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | A Comment on "Fast Bloom Filters and Their Generalization"abstractA Bloom filter is a data structure that provides probabilistic membership checking. Bloom filters have many applications in computing and communications systems. The performance of a Bloom filter is measured by false positive rate, memory size requirement, and query (or memory look-up) overhead. A recent paper by Qiao et al. proposes the Fast Bloom Filter, also called Bloom-1, which requires only a single memory look-up for a membership test. Bloom-1 achieves a reduced query overhead at the expense of a slightly higher false positive rate for a given memory size. The false positive rate of Bloom-1 has been analyzed theoretically by Qiao et al. relying on a well-known, but flawed, approximation for the false positive rate for a Bloom filter. In this comment paper we show that the Qiao et al. analysis of Bloom-1 under-estimates the false positive rate for low loads. We provide a correct analysis of Bloom-1 yielding an expression for the exact false positive rate. Pedro Reviriego, Kenneth J. Christensen, Juan Antonio Maestro |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Fault Tolerant Parallel FFTs Using Error Correction Codes and Parseval ChecksabstractSoft errors pose a reliability threat to modern electronic circuits. This makes protection against soft errors a requirement for many applications. Communications and signal processing systems are no exceptions to this trend. For some applications, an interesting option is to use algorithmic-based fault tolerance (ABFT) techniques that try to exploit the algorithmic properties to detect and correct errors. Signal processing and communication applications are well suited for ABFT. One example is fast Fourier transforms (FFTs) that are a key building block in many systems. Several protection schemes have been proposed to detect and correct errors in FFTs. Among those, probably the use of the Parseval or sum of squares check is the most widely known. In modern communication systems, it is increasingly common to find several blocks operating in parallel. Recently, a technique that exploits this fact to implement fault tolerance on parallel filters has been proposed. In this brief, this technique is first applied to protect FFTs. Then, two improved protection schemes that combine the use of error correction codes and Parseval checks are proposed and evaluated. The results show that the proposed schemes can further reduce the implementation cost of protection. Zhen Gao 0001, Pedro Reviriego, Xin Su 0001, Ming Zhao 0001, Jing Wang 0001, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | An Efficient Single and Double-Adjacent Error Correcting Parallel Decoder for the (24, 12) Extended Golay CodeabstractMemories that operate in harsh environments, like for example space, suffer a significant number of errors. The error correction codes (ECCs) are routinely used to ensure that those errors do not cause data corruption. However, ECCs introduce overheads both in terms of memory bits and decoding time that limit speed. In particular, this is an issue for applications that require strong error correction capabilities. A number of recent works have proposed advanced ECCs, such as orthogonal Latin squares or difference set codes that can be decoded with relatively low delay. The price paid for the low decoding time is that in most cases, the codes are not optimal in terms of memory overhead and require more parity check bits. On the other hand, codes like the (24,12) Golay code that minimize the number of parity check bits have a more complex decoding. A compromise solution has been recently explored for Bose-Chaudhuri-Hocquenghem codes. The idea is to implement a fast parallel decoder to correct the most common error patterns (single and double adjacent) and use a slower serial decoder for the rest of the patterns. In this brief, it is shown that the same scheme can be efficiently implemented for the (24,12) Golay code. In this case, the properties of the Golay code can be exploited to implement a parallel decoder that corrects single- and double-adjacent errors that is faster and simpler than a single-error correction decoder. The evaluation results using a 65-nm library show significant reductions in area, power, and delay compared with the traditional decoder that can correct single and double-adjacent errors. In addition, the proposed decoder is also able to correct some triple-adjacent errors, thus covering the most common error patterns. Pedro Reviriego, Shanshan Liu 0001, Liyi Xiao, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Optimizing the Implementation of SEC-DAEC Codes in FPGAsabstractSingle error correction and double-adjacent error correction (SEC-DAEC) codes are a type of error correction codes (ECCs) capable of correcting single and double-adjacent errors. They are useful in applications where multiple adjacent errors may occur, such as space or avionics. ECC encoders and decoders have a regular structure that makes it easier to accommodate them into field-programmable gate arrays (FPGAs). This brief proposes methods to optimize the decoder of SEC-DAEC codes when implemented in an FPGA, reducing the resource utilization when compared with the conventional implementations. Alfonso Sánchez-Macián, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Low Delay Single Symbol Error Correction Codes Based on Reed Solomon CodesabstractTo avoid data corruption, error correction codes (ECCs) are widely used to protect memories. ECCs introduce a delay penalty in accessing the data as encoding or decoding has to be performed. This limits the use of ECCs in high-speed memories. This has led to the use of simple codes such as single error correction double error detection (SEC-DED) codes. However, as technology scales multiple cell upsets (MCUs) become more common and limit the use of SEC-DED codes unless they are combined with interleaving. A similar issue occurs in some types of memories like DRAM that are typically grouped in modules composed of several devices. In those modules, the protection against a device failure rather than isolated bit errors is also desirable. In those cases, one option is to use more advanced ECCs that can correct multiple bit errors. The main challenge is that those codes should minimize the delay and area penalty. Among the codes that have been considered for memory protection are Reed-Solomon (RS) codes. These codes are based on non-binary symbols and therefore can correct multiple bit errors. In this paper, single symbol error correction codes based on Reed-Solomon codes that can be implemented with low delay are proposed and evaluated. The results show that they can be implemented with a substantially lower delay than traditional single error correction RS codes. Salvatore Pontarelli, Pedro Reviriego, Marco Ottavi, Juan Antonio Maestro |
IEEE Trans. Computers | 2 |
| 2015 | Fault Tolerant Parallel Filters Based on Error Correction CodesabstractDigital filters are widely used in signal processing and communication systems. In some cases, the reliability of those systems is critical, and fault tolerant filter implementations are needed. Over the years, many techniques that exploit the filters' structure and properties to achieve fault tolerance have been proposed. As technology scales, it enables more complex systems that incorporate many filters. In those complex systems, it is common that some of the filters operate in parallel, for example, by applying the same filter to different input signals. Recently, a simple technique that exploits the presence of parallel filters to achieve fault tolerance has been presented. In this brief, that idea is generalized to show that parallel filters can be protected using error correction codes (ECCs) in which each filter is the equivalent of a bit in a traditional ECC. This new scheme allows more efficient protection when the number of parallel filters is large. The technique is evaluated using a case study of parallel finite impulse response filters showing the effectiveness in terms of protection and implementation cost. Zhen Gao 0001, Pedro Reviriego, Wen Pan, Ming Zhao 0001, Jing Wang 0001, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Class of SEC-DED-DAEC Codes Derived From Orthogonal Latin Square CodesabstractRadiation-induced soft errors are a major reliability concern for memories. To ensure that memory contents are not corrupted, single error correction double error detection (SEC-DED) codes are commonly used, however, in advanced technology nodes, soft errors frequently affect more than one memory bit. Since SEC-DED codes cannot correct multiple errors, they are often combined with interleaving. Interleaving, however, impacts memory design and performance and cannot always be used in small memories. This limitation has spurred interest in codes that can correct adjacent bit errors. In particular, several SEC-DED double adjacent error correction (SEC-DED-DAEC) codes have recently been proposed. Implementing DAEC has a cost as it impacts the decoder complexity and delay. Another issue is that most of the new SEC-DED-DAEC codes miscorrect some double nonadjacent bit errors. In this brief, a new class of SEC-DED-DAEC codes is derived from orthogonal latin squares codes. The new codes significantly reduce the decoding complexity and delay. In addition, the codes do not miscorrect any double nonadjacent bit errors. The main disadvantage of the new codes is that they require a larger number of parity check bits. Therefore, they can be useful when decoding delay or complexity is critical or when miscorrection of double nonadjacent bit errors is not acceptable. The proposed codes have been implemented in Hardware Description Language and compared with some of the existing SEC-DED-DAEC codes. The results confirm the reduction in decoder delay. Pedro Reviriego, Salvatore Pontarelli, Adrian Evans, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A Synergetic Use of Bloom Filters for Error Detection and CorrectionabstractBloom filters (BFs) provide a fast and efficient way to check whether a given element belongs to a set. The BFs are used in numerous applications, for example, in communications and networking. There is also ongoing research to extend and enhance BFs and to use them in new scenarios. Reliability is becoming a challenge for advanced electronic circuits as the number of errors due to manufacturing variations, radiation, and reduced noise margins increase as technology scales. In this brief, it is shown that BFs can be used to detect and correct errors in their associated data set. This allows a synergetic reuse of existing BFs to also detect and correct errors. This is illustrated through an example of a counting BF used for IP traffic classification. The results show that the proposed scheme can effectively correct single errors in the associated set. The proposed scheme can be of interest in practical designs to effectively mitigate errors with a reduced overhead in terms of circuit area and power. Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | MCU Tolerance in SRAMs Through Low-Redundancy Triple Adjacent Error CorrectionabstractStatic random access memories (SRAMs) are key in electronic systems. They are used not only as standalone devices, but also embedded in application specific integrated circuits. One key challenge for memories is their susceptibility to radiation-induced soft errors that change the value of memory cells. Error correction codes (ECCs) are commonly used to ensure correct data despite soft errors effects in semiconductor memories. Single error correction/double error detection (SEC-DED) codes have been traditionally the preferred choice for data protection in SRAMs. During the last decade, the percentage of errors that affect more than one memory cell has increased substantially, mainly due to multiple cell upsets (MCUs) caused by radiation. The bits affected by these errors are physically close. To mitigate their effects, ECCs that correct single errors and double adjacent errors have been proposed. These codes, known as single error correction/double adjacent error correction (SEC-DAEC), require the same number of parity bits as traditional SEC-DED codes and a moderate increase in the decoder complexity. However, MCUs are not limited to double adjacent errors, because they affect more bits as technology scales. In this brief, new codes that can correct triple adjacent errors and 3-bit burst errors are presented. They have been implemented using a 45-nm library and compared with previous proposals, showing that our codes have better error protection with a moderate overhead and low redundancy. Luis J. Saiz, Pedro Reviriego, Pedro J. Gil, Salvatore Pontarelli, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Dependable reconfigurable space systems: Challenges, new trends and case studiesabstractThe current state-of-the-art radiation tolerant or hardened reconfigurable SRAM-based FGPAs offer qualification for space applications, dynamic partial reconfiguration for in-flight adaptability, high density and performance. These FPGAs are very suitable for implementation of digital signal processing algorithms, providing the possibility for orders of magnitude increases in performance over processor-based implementations and thus are well suitable to establish dependable reconfigurable space systems. Alternatively, reconfigurable Complex Programmable Logic Devices (CPLDs) can be used as low-cost, low-power solutions for non-critical space applications. In this paper, challenges and new trends for dependable reconfigurable space systems are discussed and illustrated with 3 case studies. Antonis M. Paschalis, Harald Michalik, Nektarios Kranitis, Celia López-Ongil, Pedro Reviriego |
IOLTS | 5 |
| 2014 | Exploiting a fast and simple ECC for scaling supply voltage in level-1 cachesabstractScaling supply voltage to near-threshold is a very effective approach in reducing the energy consumption of computer systems. However, executing below the safe operation margin of supply voltage introduces high number of persistent failures, especially in memory structures. Thus, it is essential to provide reliability schemes to tolerate these persistent failures in the memory structures. In this study, we adopt a Single Error Correction Multiple Adjacent Error Correction (SEC-MAEC) code in order to minimize the energy consumption of L1 caches. In our evaluations, we present that the SEC-MAEC code is a fast and energy efficient Error Correcting Code (ECC). It presents 10X less area overhead and 2X less latency for the decoder compared to Orthogonal Latin Square Code, the state-of-the art ECC utilized in the L1 cache under the scaling supply voltage. Gulay Yalcin, Emrah Islek, Oyku Tozlu, Pedro Reviriego, Adrián Cristal, Osman S. Unsal, Oguz Ergin |
IOLTS | 4 |
| 2014 | An experimental power profile of Energy Efficient Ethernet switches
Vijay Sivaraman, Pedro Reviriego, Alfonso Sánchez-Macián, Arun Vishwanath, Juan Antonio Maestro, Craig Russell |
Comput. Commun. | 2 |
| 2014 | Improving the performance of Invertible Bloom Lookup Tables
Salvatore Pontarelli, Pedro Reviriego, Michael Mitzenmacher |
Inf. Process. Lett. | 2 |
| 2014 | A Method to Extend Orthogonal Latin Square CodesabstractError correction codes (ECCs) are commonly used to protect memories from errors. As multibit errors become more frequent, single error correction codes are not enough and more advanced ECCs are needed. The use of advanced ECCs in memories is, however, limited by their decoding complexity. In this context, one-step majority logic decodable (OS-MLD) codes are an interesting option as the decoding is simple and can be implemented with low delay. Orthogonal Latin squares (OLS) codes are OS-MLD and have been recently considered to protect caches and memories. The main advantage of OLS codes is that they provide a wide range of choices for the block size and the error correction capabilities. In this brief, a method to extend OLS codes is presented. The proposed method enables the extension of the data block size that can be protected with a given number of parity bits thus reducing the overhead. The extended codes are also OS-MLD and have a similar decoding complexity to that of the original OLS codes. The proposed codes have been implemented to evaluate the circuit area and delay needed for different block sizes. Pedro Reviriego, Salvatore Pontarelli, Alfonso Sánchez-Macián, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Performance analysis of Energy Efficient Ethernet on video streaming servers
Antonio de la Oliva, Tito R. Vargas, Juan Carlos Guerri, José Alberto Hernández 0001, Pedro Reviriego |
Comput. Networks | 5 |
| 2013 | Low Complexity Concurrent Error Detection for Complex MultiplicationabstractThis paper studies the problem of designing a low complexity Concurrent Error Detection (CED) circuit for the complex multiplication function commonly used in Digital Signal Processing circuits. Five novel CED architectures are proposed and their computational complexity, area, and delay evaluated in several circuit implementations. The most efficient architecture proposed reduces the number of gates required by up to 30 percent when compared with a conventional CED architecture based on Dual Modular Redundancy. Compared to a Residue Code CED scheme, the area of the proposed architectures is larger. However, for some of the proposed CEDs delay is significantly lower with reductions exceeding 30 percent in some configurations. Salvatore Pontarelli, Pedro Reviriego, Chris J. Bleakley, Juan Antonio Maestro |
IEEE Trans. Computers | 2 |
| 2013 | A Method to Construct Low Delay Single Error Correction Codes for Protecting Data Bits OnlyabstractError correction codes (ECCs) have been used for decades to protect memories from soft errors. Single error correction (SEC) codes that can correct 1-bit error per word are a common option for memory protection. In some cases, SEC codes are extended to also provide double error detection and are known as SEC-DED codes. As technology scales, soft errors on registers also became a concern and, therefore, SEC codes are used to protect registers. The use of an ECC impacts the circuit design in terms of both delay and area. Traditional SEC or SEC-DED codes developed for memories have focused on minimizing the number of redundant bits added by the code. This is important in a memory as those bits are added to each word in the memory. However, for registers used in circuits, minimizing the delay or area introduced by the ECC can be more important. In this paper, a method to construct low delay SEC or SEC-DED codes that correct errors only on the data bits is proposed. The method is evaluated for several data block sizes, showing that the new codes offer significant delay reductions when compared with traditional SEC or SEC-DED codes. The results for the area of the encoder and decoder also show substantial savings compared to existing codes. Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Using Single Error Correction Codes to Protect Against Isolated Defects and Soft ErrorsabstractDifferent techniques have been used to deal with defects and soft errors. Repair techniques are commonly used for defects, while error correction codes are used for soft errors. Recently, some proposals have been made to use error correction codes to deal with defects. In this paper, we analyze the impact on reliability of such approaches that use error correction codes, which in addition to soft errors can resolve defects, at the cost of reduced ability to correct soft errors. The results showed that low defect rates or small memory sizes are required to have a low impact on reliability. Additionally, a technique that can improve reliability is proposed and analyzed. The results show that our new approach can achieve a similar reliability in terms of time to failure as that of a defect free memory at the cost of a more complex decoding algorithm. Costas Argyrides, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Reliab. | 2 |
| 2013 | Error Detection in Majority Logic Decoding of Euclidean Geometry Low Density Parity Check (EG-LDPC) CodesabstractIn a recent paper, a method was proposed to accelerate the majority logic decoding of difference set low density parity check codes. This is useful as majority logic decoding can be implemented serially with simple hardware but requires a large decoding time. For memory applications, this increases the memory access time. The method detects whether a word has errors in the first iterations of majority logic decoding, and when there are no errors the decoding ends without completing the rest of the iterations. Since most words in a memory will be error-free, the average decoding time is greatly reduced. In this brief, we study the application of a similar technique to a class of Euclidean geometry low density parity check (EG-LDPC) codes that are one step majority logic decodable. The results obtained show that the method is also effective for EG-LDPC codes. Extensive simulation results are given to accurately estimate the probability of error detection for different code sizes and numbers of errors. Pedro Reviriego, Juan Antonio Maestro, Mark F. Flanagan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Concurrent Error Detection for Orthogonal Latin Squares Encoders and Syndrome ComputationabstractError correction codes (ECCs) are commonly used to protect memories against errors. Among ECCs, orthogonal latin squares (OLS) codes have gained renewed interest for memory protection due to their modularity and the simplicity of the decoding algorithm that enables low delay implementations. An important issue is that when ECCs are used, the encoder and decoder circuits can also suffer errors. In this brief, a concurrent error detection technique for OLS codes encoders and syndrome computation is proposed and evaluated. The proposed method uses the properties of OLS codes to efficiently implement a parity prediction scheme that detects all errors that affect a single circuit node. Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Low Power embedded DRAM caches using BCH code partitioningabstractTechnology advances have recently enabled the use of DRAMs into logic integrated circuits. These embedded DRAMs can be used to efficiently implement caches since DRAMs require substantially less area than SRAMs. One challenge for DRAM based caches is that a small time between refreshes is needed to ensure data retention. These refreshes increase the power consumption even when the cache is idle. To mitigate this issue, the use of longer times between refreshes combined with the use of Error Correction Codes (ECCs) has been recently proposed. The idea is that the time between refreshes can be increased significantly while only causing data retention failures on a small percentage of the cells. Then those errors can be corrected by the ECC. For this scheme to be efficient the number of additional bits required by the ECC should be small. This is achieved by using large data blocks for the ECC which in turns means that a large data block has to be accessed even when only a small portion of it is needed. This has no effect on idle power consumption but increases the dynamic power consumption and reduces the effective memory bandwidth. In this paper, a technique to mitigate this issue is proposed. It enables better granularity in the read data accesses by partitioning the ECC block into two sub-blocks and modifying the error detection and correction processes. This reduces the dynamic power consumption and increases the available memory bandwidth while requiring only a moderate increase in the number of additional bits. Pedro Reviriego, Alfonso Sánchez-Macián, Juan Antonio Maestro |
IOLTS | 1 |
| 2012 | Network monitoring for energy efficiency in large-scale networks: the case of the Spanish Academic Network
José Luis García-Dorado, Eduardo Magaña, Pedro Reviriego, Mikel Izal, Daniel Morató, Juan Antonio Maestro, Javier Aracil 0001, Jorge E. López de Vergara |
J. Supercomput. | 3 |
| 2012 | Efficient Majority Logic Fault Detection With Difference-Set Codes for Memory ApplicationsabstractNowadays, single event upsets (SEUs) altering digital circuits are becoming a bigger concern for memory applications. This paper presents an error-detection method for difference-set cyclic codes with majority logic decoding. Majority logic decodable codes are suitable for memory applications due to their capability to correct a large number of errors. However, they require a large decoding time that impacts memory performance. The proposed fault-detection method significantly reduces memory access time when there is no error in the data read. The technique uses the majority logic decoder itself to detect failures, which makes the area overhead minimal and keeps the extra power consumption low. Shih-Fu Liu, Pedro Reviriego, Juan Antonio Maestro |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Designing ad-hoc scrubbing sequences to improve memory reliability against soft errorsabstractIn this paper, we propose the use of ad-hoc scrubbing sequences to improve memory reliability. The key idea is to exploit the locality of the errors caused by a Multiple Cell Upset (MCU) to make scrubbing more efficient. The starting point is the MCU distributions for a given device. A procedure is presented that uses that information to determine an ad-hoc scrubbing sequence that maximizes reliability. The approach is then applied to a case study and results show a significant increase in the Mean Time To Failure (MTTF) compared with traditional scrubbing. Pedro Reviriego, Juan Antonio Maestro, Sanghyeon Baeg |
DAC | 1 |
| 2011 | Validation and optimization of TMR protections for circuits in radiation environmentsabstractA methodology based on optimization processes and software fault injection is presented to verify and improve TMR protection against SEUs. It allows validating the reliability achieved by the protection, optimizing the solution area cost. Oscar Ruano, Juan Antonio Maestro, Pedro Reviriego |
DDECS | 3 |
| 2011 | Using Coordinated Transmission with Energy Efficient Ethernet
Pedro Reviriego, Kenneth J. Christensen, Alfonso Sánchez-Macián, Juan Antonio Maestro |
Networking (1) | 1 |
| 2011 | Fault Tolerant Single Error Correction Encoders
Juan Antonio Maestro, Pedro Reviriego, Costas Argyrides, Dhiraj K. Pradhan |
J. Electron. Test. | 2 |
| 2011 | Offset DMR: A Low Overhead Soft Error Detection and Correction Technique for Transform-Based ConvolutionabstractA novel concurrent soft error detection and correction scheme is introduced for parallel hardware implementations of transform-based convolution. The proposed technique is based on the structure of radix-2 Fast Fourier Transforms (FFT) of length 2nwhere n is an integer. The scheme can provide up to 100 percent detection and correction of isolated single soft errors in the convolution at the cost of little more than double the system area, rather than triple, as is required when using conventional Triple Modular Redundancy (TMR). Pedro Reviriego, Chris J. Bleakley, Juan Antonio Maestro, Anne O'Donnell |
IEEE Trans. Computers | 1 |
| 2011 | Mitigating the effects of large multiple cell upsets (MCUs) in memoriesabstractReliability is a critical issue for memories. Radiation particles that hit the device can cause errors in some cells, which can lead to data corruption. To avoid this problem, memories are protected with per-word error correction codes (ECCs). Typically, single-error correction and double-error detection (SEC-DED) codes are used. As technology scales, errors caused by radiation particles on memories tend to affect more than one cell—what is known as a multiple cell upset (MCU). To ensure that only a single cell is affected in each word, interleaving is used. With interleaving, cells that belong to the same word are placed at a sufficient distance such that an MCU will only affect a single cell on each word. The use of interleaving significantly increases the cost of the device. Also, determining the interleaving distance (ID) required to avoid MCUs causing double errors is not trivial. Typically, accelerated radiation experiments with a limited number of particle hits are used. They provide a lower bound on the required ID, but larger MCUs may occur with a low probability. But even if the percentage of such large MCUs is very low, the impact on reliability can be significant. This article presents a technique to mitigate the effects of large MCUs that is, those that exceed the ID, on memory reliability. The proposed approach is able to correct most double errors caused by large MCUs by exploiting the locality of the errors within an MCU. Juan Antonio Maestro, Pedro Reviriego, Sanghyeon Baeg, Shi-Jie Wen, Richard Wong |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Reliability analysis of memories protected with BICS and a per-word parity bitabstractThis article presents an analysis of the reliability of memories protected with Built-in Current Sensors (BICS) and a per-word parity bit when exposed to Single Event Upsets (SEUs). Reliability is characterized by Mean Time to Failure (MTTF) for which two analytic models are proposed. A simple model, similar to the one traditionally used for memories protected with scrubbing, is proposed for the low error rate case. A more complex Markov model is proposed for the high error rate case. The accuracy of the models is checked using a wide set of simulations. The results presented in this article allow fast estimation of MTTF enabling design of optimal memory configurations to meet specified MTTF goals at minimum cost. Additionally the power consumption of memories protected with BICS is compared to that of memories using scrubbing in terms of the number of read cycles needed in both configurations. Pedro Reviriego, Juan Antonio Maestro, Chris J. Bleakley |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2009 | Soft error detection and correction for FFT based convolution using different block lengthsabstractThe structure of radix-2 Fast Fourier Transforms of length 2nwhere n is an integer is used to propose a new soft error detection and correction scheme for transform based convolution. The scheme can provide up to 100% detection and correction of isolated soft errors for, in many cases, approximately double the original system cost in terms of area and/or computational complexity. This is a substantial reduction when compared with conventional Triple Modular Redundancy. The method can be used for both hardware and software implementations of transform-based convolution. Pedro Reviriego, Juan Antonio Maestro, Anne O'Donnell, Chris J. Bleakley |
IOLTS | 1 |
| 2009 | Protection against soft errors in the space environment: A finite impulse response (FIR) filter case study
Juan Antonio Maestro, Pedro Reviriego, Pilar Reyes, Oscar Ruano |
Integr. | 2 |
| 2009 | Efficient error detection codes for multiple-bit upset correction in SRAMs with BICSabstractMemories are one of the most widely used elements in electronic systems, and their reliability when exposed to Single Events Upsets (SEUs) has been studied extensively. As transistor sizes shrink, Multiple Bits Upsets (MBUs) are becoming an increasingly important factor in the reliability of memories exposed to radiation effects. To address this issue, Built-in Current Sensors (BICS) have recently been applied in conjunction with Single Error Correction/Double Error Detection (SEC-DED) codes to protect memories from MBUs. In this article, this approach is taken one step further, proposing specific codes optimized to be combined with BICS to provide protection against MBUs in memories. By exploiting the locality of errors within an MBU and the error detection and location capabilities of BICS, the proposed codes result in both a better protection level and a reduced cost compared with the existing SEC-DED approach. Pedro Reviriego, Juan Antonio Maestro |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2009 | Reliability of Single-Error Correction Protected MemoriesabstractReliability is a critical factor for systems operating in radiation environments. Among the different components in a system, memories are one of the parts most sensitive to soft errors due to their relatively large area. Due to their large cost, traditional techniques like triple modular redundancy are not used to protect memories. A typical approach is to apply error correction codes to correct single errors, and detect double errors. This type of codes, for example those based on Hamming, provides an initial level of protection. Detected single errors are usually corrected using scrubbing, by which the memory positions are periodically re-written after a fixed (deterministic scrubbing), or variable period (probabilistic scrubbing). These traditional models usually offer good results when calculating the reliability of memories (e.g. through the mean time to failure). However, there are some particularities that are not modeled through these approaches, to the best of our knowledge. One of these particularities is how double errors are handled. In a traditional approach, two errors in the same word produce always a system failure (only single errors can be corrected). However, if the two (or more) errors affect the same bit, either the second one reinforces the first one (keeping just a single error), or corrects it. In both scenarios, the resulting situation does not trigger a system failure, which has a direct impact on the reliability of the memory. In this paper, traditional reliability models are refined to handle the mentioned scenarios, which produces a more precise analysis in the calculation of mean time to failure for memory systems. Juan Antonio Maestro, Pedro Reviriego |
IEEE Trans. Reliab. | 2 |
| 2008 | Study of the effects of MBUs on the reliability of a 150 nm SRAM deviceabstractSoft errors induced by radiation are an increasing problem in the microelectronic field. Although traditional models estimate the reliability of memories suffering Single Event Upsets (SEUs), Multiple Bit Upsets (MBUs) are becoming more and more important as technology scales. In this paper, a model that deals with MBUs in memory systems, which allows calculating reliability in a fast way similar to the SEU case, has been used to analyze the Mean Time To Failure (MTTF) of a 150 nm device under radiation. This analysis illustrates the importance that physical factors, as the energy, have on the system reliability. Juan Antonio Maestro, Pedro Reviriego |
DAC | 2 |