Salvatore Pontarelli

dblp:72/5613 · DBLP profile ↗
← Back
66ranked-venue papers
20as first author
8since 2021 · last 2025
0000-0002-3626-6404ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 12 first-author · 3 since 2021Software engineering, systems software and programming languages · 15 · 7 first-author · 1 since 2021Computer networks · 11 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorTheory of computation · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Security and privacy · 1
YearPublicationVenuePosition
2025 Sensing at the Edge: Location-Aware Caching
abstract
Sensing has become a fundamental component of modern network infrastructures, powering applications from environmental monitoring to industrial automation, and bridging the gap between digital systems and the physical world. Given the large amount of data generated by these systems, it is important to find strategies that are able to intelligently manage the flow of data between all interconnected devices, in order to reduce the utilized bandwidth to a minimum. In this paper, we study the use of Wireless Edge Caching (WEC) techniques to reduce latency and optimize bandwidth and energy consumption of sensing applications. We study a specific sensing scenario where a network of wireless sensors spread over a vast territory continuously collects data, part of which must be transmitted to a central base station. To aid and enhance the performance of this application, we employ the use of WEC techniques. Since the frequency of data queries is related to the sensors’ location, we introduce a novel caching algorithm, Closest In Farthest Out (CIFO), tailored to this scenario. CIFO is able to capture the characteristics of the sensing application and perform cache eviction decisions accordingly. We demonstrate the performance of our solution by implementing it in a simulated environment and comparing it to traditional caching strategies, showing how our solution is able to outperform the other strategies under multiple settings and different metrics.
Federico Trombetti, Novella Bartolini, Salvatore Pontarelli
CNSM3
2025 NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI
abstract
NET4EXA aims to develop a next-generation high-performance interconnect for HPC and AI systems, addressing the increasing demands of large-scale infrastructures, such as those required for training Large Language Models. Building upon the proven BXI (Bull eXascale Interconnect) European technology used in TOP15 supercomputers, NET4EXA will deliver the new BXI release, BXIv3, a complete hardware and software interconnect solution, including switch and network interface components. The project will integrate a fully functional pilot system at TRL 8, ready for deployment into upcoming exascale and post-exascale systems from 2025 onward. Leveraging prior research from European initiatives like RED-SEA, the previous achievements of consortium partners and over 20 years of expertise from BULL, NET4EXA also lays the groundwork for the future generation of BXI, BXIv4, providing analysis and preliminary design. The project will use a hybrid development and co-design approach, combining commercial switch technology with custom IP and FPGA-based NICs. Performances of NET4EXA BXIv3 interconnect will be evaluated using a broad portfolio of benchmarks, scientific scalable applications, and AI workloads.
Michele Martinelli, Roberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Pier Stanislao Paolucci, Elena Pastorelli, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Francesco Simula, Piero Vicini, David Colin, Gregoire Pichon, Alexandre Louvet, John Gliksberg, Matteo Turisini, Andrea Monterubbiano, Jean-Philippe Nomine, Denis Dutoit, Hugo Taboada, Lilia Zaourar, Mohamed Benazouz, Angelos Bilas, Fabien Chaix, Manolis Katevenis, Nikolaos Chrysos, Evangelos Mageiropoulos, Christos Kozanitis, Thomas Moen, Steffen Persvold, Einar Rustad, Sandro Fiore, Fabrizio Granelli, Simone Pezzuto, Raffaello Potestio, Luca Tubiana, Philippe Velha, Flavio Vella, Daniele De Sensi, Salvatore Pontarelli
DSD44
2023 eHDL: Turning eBPF/XDP Programs into Hardware Designs for the NIC
abstract
Scaling network packet processing performance to meet the increasing speed of network ports requires software programs to carefully leverage the network devices’ hardware features. This is a complex task for network programmers, who need to learn and deal with the heterogeneity of device architectures, and re-think their software to leverage them. In this paper we make first steps to reverse this design process, enabling the automatic generation of tailored hardware designs starting from a network packet processing program. We introduce eHDL, a high-level synthesis tool that automatically generates hardware pipelines from unmodified Linux’s eBPF/XDP programs. eHDL is designed to enable software developers to directly define and implement the hardware functions they need in the NIC. We prototype eHDL targeting a Xilinx Alveo U50 FPGA NIC, and evaluate it with a set of 5 eBPF/XDP programs. Our results show that the generated pipelines are efficient in terms of required hardware resources, using only 6.5%-13.3% of the FPGA, and always achieve the line rate forwarding throughput with about 1 microsecond of per-packet forwarding latency. Compared to other network-specific high-level synthesis tool, eHDL enables software programmers with no hardware expertise to describe stateful functions that operate on the entire packet data. Compared to alternative processor-based solutions that perform eBFP/XDP offloading to a NIC, eHDL provides 10-100x higher throughput.
Alessandro Rivitti, Roberto Bifulco, Angelo Tulumello, Marco Bonola, Salvatore Pontarelli
ASPLOS (3)5
2023 Metronome: Adaptive and Precise Intermittent Packet Retrieval in DPDK
abstract
The increasing performance requirements of modern applications place a significant burden on software-based packet processing. Most of today’s software input/output accelerations achieve high performance at the expense of reserving CPU resources dedicated to continuously poll the Network Interface Card. This is specifically the case with DPDK (Data Plane Development Kit), probably the most widely used framework for software-based packet processing today. The approach presented in this paper, descriptively called Metronome, has the dual goals of providing CPU utilization proportional to the load, and allowing flexible sharing of CPU resources between I/O tasks and applications. Metronome replaces DPDK’s continuous polling with an intermittent sleep&wake mode, and revolves around a new multi-threaded operation, which improves service continuity. Since the proposed operation trades CPU usage with buffering delay, we propose an analytical model devised to dynamically adapt the sleep&wake parameters to the actual traffic load, meanwhile providing a target average latency. Our experimental results show a significant reduction of the CPU cycles, improvements in power usage, and robustness to CPU sharing even when challenged with CPU-intensive applications.
Marco Faltelli, Giacomo Belocchi, Francesco Quaglia, Salvatore Pontarelli, Giuseppe Bianchi 0001
IEEE/ACM Trans. Netw.4
2022 Attacking Adaptive Cuckoo Filters: Too Much Adaptation Can Kill You
abstract
Adaptation has recently been proposed to reduce the false positive rate of approximate membership check filters for applications in which the same elements are checked multiple times. Its operational principle is to adapt the filter when a false positive occurs for a given element, such that subsequent checks of that element do not cause a positive result (as beneficial for example in networking). Security is an important consideration for approximate membership check filters and several attacks have been described in the literature; therefore, it is of interest to study the security of adaptive filters. In this paper, we consider adaptive cuckoo filters and show that an attacker can generate sequences of lookups that cause the filter to continuously adapt and not being able to remove the false positives. This degrades the filter performance due to the adaptation overhead; it also makes it harder for other false positives to be removed, because adaptation can be monopolized by the attacker. This can be done when the attacker has only a black-box access to the filter being able to perform lookups but with no knowledge of the implementation of the filter. The proposed attacks have been implemented and tested to validate their effectiveness in terms of the construction of the attack set and the impact of the attack itself. The evaluation results confirm that adaptation unfortunately increases the attack surface of filters and new mechanisms to protect them should be developed.
Pedro Reviriego, Alfonso Sánchez-Macián, Salvatore Pontarelli, Shanshan Liu 0001, Fabrizio Lombardi
IEEE Trans. Netw. Serv. Manag.3
2022 Processor Security: Detecting Microarchitectural Attacks via Count-Min Sketches
abstract
The continuous quest for performance pushed processors to incorporate elements such as multiple cores, caches, acceleration units, or speculative execution that make systems very complex. On the other hand, these features often expose unexpected vulnerabilities that pose new challenges. For example, the timing differences introduced by caches or speculative execution can be exploited to leak information or detect activity patterns. Protecting embedded systems from existing attacks is extremely challenging, and it is made even harder by the continuous rise of new microarchitectural attacks (e.g., the Spectre and Orchestration attacks). In this article, we present a new approach based on count-min sketches for detecting microarchitectural attacks in the microprocessors featured by embedded systems. The idea is to add to the system a security checking module (without modifying the microprocessor under protection) in charge of observing the fetched instructions and identifying and signaling possible suspicious activities without interfering with the nominal activity of the system. The proposed approach can be programmed at design time (and reprogrammed after deployment) in order to always keep updated the list of the attacks that the checker is able to identify. We integrated the proposed approach in a large RISC-V core, and we proved its effectiveness in detecting several versions of the Spectre, Orchestration, Rowhammer, and Flush + Reload attacks. In its best configuration, the proposed approach has been able to detect 100% of the attacks, with no false alarms and introducing about 10% area overhead, about 4% power increase, and without working frequency reduction.
Kerem Arikan, Alessandro Palumbo, Luca Cassano, Pedro Reviriego, Salvatore Pontarelli, Giuseppe Bianchi 0001, Oguz Ergin, Marco Ottavi
IEEE Trans. Very Large Scale Integr. Syst.5
2021 Perfect cuckoo filters
abstract
Bloom filters and cuckoo filters are used in many applications to reduce the amount of memory needed to check if an element belongs to a set. The main drawback of these filters is that with low probability, a positive is returned for an element that is not in the set. Recently, the concept of Bloom filters with a false positive free zone has been introduced showing that false positives can be avoided when the universe from which elements are taken and the number of elements inserted in the filter are both small. Unfortunately, this limits the use of such false positive free Bloom filters in many practical applications. In this paper, a false positive free, i.e. perfect, cuckoo filter is presented and evaluated. The proposed design supports universe sizes of billions of elements and stores millions of elements, making it practical for a wide range of applications. The perfect cuckoo filter can be also used to perform mapping, further extending the range of scenarios in which can be used. The benefits of the proposed perfect cuckoo filter are illustrated with two case studies: IP address blacklisting and longest prefix match for IP forwarding.
Pedro Reviriego, Salvatore Pontarelli
CoNEXT2
2021 FlowFight: High performance-low memory top-k spreader detection
Valerio Bruschi, Salvatore Pontarelli, Jerome Tollet, David Barach, Giuseppe Bianchi 0001
Comput. Networks2
2020 Metronome: adaptive and precise intermittent packet retrieval in DPDK
abstract
DPDK (Data Plane Development Kit) is arguably today's most employed framework for software packet processing. Its impressive performance however comes at the cost of precious CPU resources, dedicated to continuously poll the NICs. To face this issue, this paper presents Metronome, an approach devised to replace the continuous DPDK polling with a sleep&wake intermittent mode. Metronome revolves around two main innovations. First, we design a microseconds time-scale sleep function, named hr_sleep(), which outperforms Linux' nanosleep() of more than one order of magnitude in terms of precision when running threads with common time-sharing priorities. Then, we design, model, and assess an efficient multi-thread operation which guarantees service continuity and improved robustness against preemptive thread executions, like in common CPU-sharing scenarios, meanwhile providing controlled latency and high polling efficiency by dynamically adapting to the measured traffic load.
Marco Faltelli, Giacomo Belocchi, Francesco Quaglia, Salvatore Pontarelli, Giuseppe Bianchi 0001
CoNEXT4
2020 When filtering is not possible caching negatives with fingerprints comes to the rescue
abstract
Bloom filters are widely used in networking to accelerate checks and in particular to avoid accessing slow memories when a match will not be found. Unfortunately, filtering requires several on-chip memory bits per element and thus when the tables are large and the on-chip memory small is not applicable. In those cases, caching the most frequently accessed elements on-chip seems the only viable option. However, the key is typically formed by several packet header fields, which means that each cache entry consumes a significant amount of on-chip memory bits. In this paper, an efficient scheme to cache negatives, that is elements that will not find a match on the table is presented. In more detail, the proposed scheme enables the caching of negatives using less than 16 bits per entry regardless of the size of the key. This translates to a reduction of at least 6x in the size of the cache when used for example flow tracking.
Pedro Reviriego, Salvatore Pontarelli
CoNEXT2
2020 hXDP: Efficient Software Packet Processing on FPGA NICs
M. Spaziani Brunella, Giacomo Belocchi, Marco Bonola, Salvatore Pontarelli, Giuseppe Siracusano, Giuseppe Bianchi 0001, Aniello Cammarano, Alessandro Palumbo, Luca Petrucci, Roberto Bifulco
OSDI4
2020 Cuckoo Filters and Bloom Filters: Comparison and Application to Packet Classification
abstract
Bloom filters are used to perform approximate membership checking in a wide range of applications in both computing and networking, but the recently introduced cuckoo filter is also gaining popularity. Therefore, it is of interest to compare both filters and provide insights into their features so that designers can make an informed decision when implementing approximate membership checking in a given application. This article first compares Bloom and cuckoo filters focusing on a packet classification application. The analysis identifies a shortcoming of cuckoo filters in terms of false positive rate when they do not operate close to full occupancy. Based on that observation, this article also proposes the use of a configurable bucket to improve the scaling of the false positive rate of the cuckoo filter with occupancy.
Pedro Reviriego, Jorge Martínez 0001, David Larrabeiti, Salvatore Pontarelli
IEEE Trans. Netw. Serv. Manag.4
2019 FlowBlaze: Stateful Packet Processing in Hardware
Salvatore Pontarelli, Roberto Bifulco, Marco Bonola, Carmelo Cascone, M. Spaziani Brunella, Valerio Bruschi, Davide Sanvito, Giuseppe Siracusano, Antonio Capone, Michio Honda, Felipe Huici
NSDI1
2019 High-speed data plane and network functions virtualization by vectorizing packet processing
Leonardo Linguaglossa, Dario Rossi 0001, Salvatore Pontarelli, David Barach, Damjan Marjon, Pierre Pfister
Comput. Networks3
2019 CuCoTrack: Cuckoo filter based connection tracking
Pedro Reviriego, Salvatore Pontarelli, Gil Levy
Inf. Process. Lett.2
2019 Survey of Performance Acceleration Techniques for Network Function Virtualization
abstract
The ongoing network softwarization trend holds the promise to revolutionize network infrastructures by making them more flexible, reconfigurable, portable, and more adaptive than ever. Still, the migration from hard-coded/hard-wired network functions toward their software-programmable counterparts comes along with the need for tailored optimizations and acceleration techniques so as to avoid or at least mitigate the throughput/latency performance degradation with respect to fixed function network elements. The contribution of this paper is twofold. First, we provide a comprehensive overview of the host-based network function virtualization (NFV) ecosystem, covering a broad range of techniques, from low-level hardware acceleration and bump-in-the-wire offloading approaches to high-level software acceleration solutions, including the virtualization technique itself. Second, we derive guidelines regarding the design, development, and operation of NFV-based deployments that meet the flexibility and scalability requirements of modern communication networks.
Leonardo Linguaglossa, Stanislav Lange, Salvatore Pontarelli, Gábor Rétvári, Dario Rossi 0001, Thomas Zinner, Roberto Bifulco, Michael Jarschel, Giuseppe Bianchi 0001
Proc. IEEE3
2019 XTRA: Towards Portable Transport Layer Functions
abstract
XTRA (XFSM for Transport) aims at providing a first attempt towards a “code-once-port-everywhere” platform-agnostic programming abstraction tailored to the deployment of transport layer functions. XTRA's programming abstraction not only fits SW platforms, but is specifically designed to harness, with no re-coding effort, the offloading opportunities offered by CPU-less HW boards or smart NICs. We demonstrate the viability of XTRA with three completely different implementations of the underlying execution engine (HW proof-of-concept on a NetFPGA board, User-space SW over Linux' Open Data Plane, and NS3 emulator). Flexibility is shown via a number of example applications, ranging from a variety of congestion control algorithms, to a middlebox-type TCP proxy functionality, up to a customized “Timer-Based” (TB) TCP which leverages the native reliance of XTRA on timers, so as to produce a loss recovery operation which, despite being formalized only via a handful of code lines, performs almost comparable with the highly optimized Linux and FreeBSD implementations.
Giuseppe Bianchi 0001, Michael Welzl, Angelo Tulumello, Francesco Gringoli, Giacomo Belocchi, Marco Faltelli, Salvatore Pontarelli
IEEE Trans. Netw. Serv. Manag.7
2019 TupleMerge: Fast Software Packet Processing for Online Packet Classification
abstract
Packet classification is an important part of many networking devices, such as routers and firewalls. Software-defined networking (SDN) heavily relies on online packet classification which must efficiently process two different streams: incoming packets to classify and rules to update. This rules out many offline packet classification algorithms that do not support fast updates. We propose a novel online classification algorithm, TupleMerge (TM), derived from tuple space search (TSS), the packet classifier used by Open vSwitch (OVS). TM improves upon TSS by combining hash tables which contain rules with similar characteristics. This greatly reduces classification time preserving similar performance in updates. We validate the effectiveness of TM using both simulation and deployment in a full-fledged software router, specifically within the vector packet processor (VPP). In our simulation results, which focus solely on the efficiency of the classification algorithm, we demonstrate that TM outperforms all other state of the art methods, including TSS, PartitionSort (PS), and SAX-PAC. For example, TM is 34% faster at classifying packets and 30% faster at updating rules than PS. We then experimentally evaluate TM deployed within the VPP framework comparing TM against linear search and TSS, and also against TSS within the OVS framework. This validation of deployed implementations is important as SDN frameworks have several optimizations such as caches that may minimize the influence of a classification algorithm. Our experimental results clearly validate the effectiveness of TM. VPP TM classifies packets nearly two orders of magnitude faster than VPP TSS and at least one order of magnitude faster than OVS TSS.
James Daly, Valerio Bruschi, Leonardo Linguaglossa, Salvatore Pontarelli, Dario Rossi 0001, Jerome Tollet, Eric Torng, Andrew Yourtchenko
IEEE/ACM Trans. Netw.4
2019 Error Detection and Correction in SRAM Emulated TCAMs
abstract
Ternary content addressable memories (TCAMs) are widely used in network devices to implement packet classification. They are used, for example, for packet forwarding, for security, and to implement software-defined networks (SDNs). TCAMs are commonly implemented as standalone devices or as an intellectual property block that is integrated on networking application-specific integrated circuits. On the other hand, field-programmable gate arrays (FPGAs) do not include TCAM blocks. However, the flexibility of FPGAs makes them attractive for SDN implementations, and most FPGA vendors provide development kits for SDN. Those need to support TCAM functionality and, therefore, there is a need to emulate TCAMs using the logic blocks available in the FPGA. In recent years, a number of schemes to emulate TCAMs on FPGAs have been proposed. Some of them take advantage of the large number of memory blocks available inside modern FPGAs to use them to implement TCAMs. A problem when using memories is that they can be affected by soft errors that corrupt the stored bits. The memories can be protected with a parity check to detect errors or with an error correction code to correct them, but this requires additional memory bits per word. In this brief, the protection of the memories used to emulate TCAMs is considered. In particular, it is shown that by exploiting the fact that only a subset of the possible memory contents are valid, most single-bit errors can be corrected when the memories are protected with a parity bit.
Pedro Reviriego, Salvatore Pontarelli, Anees Ullah
IEEE Trans. Very Large Scale Integr. Syst.2
2019 PR-TCAM: Efficient TCAM Emulation on Xilinx FPGAs Using Partial Reconfiguration
abstract
Modern field-programmable gate arrays (FPGAs) provide a vast amount of logic resources that can be used to implement complex systems while providing the flexibility to modify the design once deployed. This makes them attractive for software-defined networks (SDNs) applications, and, in fact, most vendors provide the building blocks needed for those applications, which include basic packet classification functions such as exact match, longest prefix match, and match with wildcards. Those are needed for different functions such as routing, security filtering, monitoring or quality of service. The match with wildcards can be done using ternary content addressable memories (TCAMs). TCAMs can be implemented as independent standalone devices or as Internet Protocol (IP) blocks that are used inside networking application-specific integrated circuits (ASICs) such as switching ICs. In both cases, the cells of a TCAM are more complex than that of a normal memory and also than that of a binary content addressable memory (CAMs). This is due to the more complex matching that they need to implement. As FPGAs are used in many different applications, it does not make sense to include TCAM blocks inside them as they would be used only in a small fraction of the systems. Therefore, TCAMs are emulated using the logic resources available inside the FPGA. In recent years, a number of schemes to emulate TCAMs on FPGAs have been proposed, some of them based on the use of the logic resources and others on the use of the embedded memory blocks available on the FPGA. In this brief, a technique to efficiently emulate TCAMs on Xilinx FPGAs is presented. The proposed scheme is based on the use of lookup tables (LUTs) and partial reconfiguration to achieve a more effective use of the FPGA resources while supporting the addition and removal of rules. The proposed scheme has been compared to existing implementations and the results show that it can achieve significant savings in resource usage. In addition, it enables the use of all the LUTs in the device for TCAM implementation, something that is not supported by existing approaches that use LUTRAMs.
Pedro Reviriego, Anees Ullah, Salvatore Pontarelli
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Adaptive Cuckoo Filters
abstract
We introduce the adaptive cuckoo filter (ACF), a data structure for approximate set membership that extends cuckoo filters by reacting to false positives, removing them for future queries. As an example application, in packet processing queries may correspond to flow identifiers, so a search for an element is likely to be followed by repeated searches for that element. Removing false positives can therefore significantly lower the false positive rate. The ACF, like the cuckoo filter, uses a cuckoo hash table to store fingerprints. We allow fingerprint entries to be changed in response to a false positive in a manner designed to minimize the effect on the performance of the filter. We show that the ACF is able to significantly reduce the false positive rate by presenting both a theoretical model for the false positive rate and simulations using both synthetic data sets and real packet traces.
Michael Mitzenmacher, Salvatore Pontarelli, Pedro Reviriego
ALENEX2
2018 A Programmable Hardware Calendar for High Resolution Pacing
abstract
The challenge addressed in this paper consists in offloading packet-based pacing to a hardware Network Interface Card, while retaining the flexibility of software timers. In this direction, we propose, design, implement, and evaluate a hardware calendar, which can be programmed via a simple yet very flexible programming interface leveraging stateful (adaptive) per-packet timers. We show, for both specific examples (exponential, linearly increasing, etc) as well as for the general case, how to derive such a per-packet timer setting from a high-level desired rate envelope. Further, we describe and evaluate an FPGA implementation which relies on a novel insertion strategy for solving collisions in the calendar's hash table.
Salvatore Pontarelli, Giuseppe Bianchi 0001, Michael Welzl
HPSR1
2018 Multiple Hash Matching Units (MHMU): An Algorithmic Ternary Content Addressable Memory Design for Field Programmable Gate Arrays
abstract
As applications and user requirements are constantly evolving, there is a need to provide flexible networks that are able to process packets at high speed. One of the basic functions used for packet processing is matching a key formed by some fields of the incoming packet header against a set of stored rules. This is done for example to determine the next hop of a packet or to apply security checks on a firewall. In many cases, the stored rules have do not care bits as that enables a more flexible and compact representation of the rules. Therefore, the matching can be done in hardware using Ternary Content Addressable Memories (TCAMs). However, TCAMs pose several problems in many implementations. For example, for ASICs they require much more circuit area and power than standard SRAMs. On the other hand, designs based on programmable logic such as Field Programmable Gate Arrays (FPGAs) can only use the blocks provided by the FPGA that do not typically include TCAMs. In this last case, a TCAM can be emulated using the FPGA logic resources but with a large cost. To reduce the cost of implementing TCAMs, a number of algorithmic solutions have been proposed and are known as Algorithmic TCAMs or A-TCAMs. Most of those schemes target either software or ASIC implementations. In this paper we present Multiple Hash Matching Units (MHMU) an A-TCAM solution targeted towards FPGA implementations. The proposed scheme exploits the massive parallelism of FPGAs to implement many hash based matching units that use the embedded block RAM memories of the FPGA. The proposed MHMU scheme has been mapped to a Xilinx series 7 FPGA to check its efficiency in terms of resource usage and its scalability. To validate the effectiveness of MHMU, a simple configuration has been tested with Classbench generated sets of rules. The results show that the MHMU is able to consistently accommodate sets with several tens of thousands of rules with large keys.
Pedro Reviriego, Salvatore Pontarelli, Anees Ullah, Ali Zahir, Giuseppe Bianchi 0001
HPSR2
2018 EMOMA: Exact Match in One Memory Access
abstract
An important function in modern routers and switches is to perform a lookup for a key. Hash-based methods, and in particular cuckoo hash tables, are popular for such lookup operations, but for large structures stored in off-chip memory, such methods have the downside that they may require more than one off-chip memory access to perform the key lookup. Although the number of off-chip memory accesses can be reduced using on-chip approximate membership structures such as Bloom filters, some lookups may still require more than one off-chip memory access. This can be problematic for some hardware implementations, as having only a single off-chip memory access enables a predictable processing of lookups and avoids the need to queue pending requests. We provide a data structure for hash-based lookups based on cuckoo hashing that uses only one off-chip memory access per lookup, by utilizing an on-chip pre-filter to determine which of multiple locations holds a key. We make particular use of the flexibility to move elements within a cuckoo hash table to ensure the pre-filter always gives the correct response. While this requires a slightly more complex insertion procedure and some additional memory accesses during insertions, it is suitable for most packet processing applications where key lookups are much more frequent than insertions. An important feature of our approach is its simplicity. Our approach is based on simple logic that can be easily implemented in hardware, and hardware implementations would benefit most from the single off-chip memory access per lookup.
Salvatore Pontarelli, Pedro Reviriego, Michael Mitzenmacher
IEEE Trans. Knowl. Data Eng.1
2017 Implementing advanced network functions for datacenters with stateful programmable data planes
abstract
Programmable dataplanes are emerging as a disruptive technology to implement network function virtualization in an SDN environment. This technology can be further enhanced by using data plane abstraction with stateful processing. In this paper we focus on real world use cases with stateful forwarding requirements to validate Open Packet Processor data plane abstraction. We first demonstrate the suitability of OPP for implementing complex stateful network functions by providing the detailed implementation of three use cases. Second, we assess the scalability of the use case implementations in the context of a datacenter deployment.
Marco Bonola, Roberto Bifulco, Luca Petrucci, Salvatore Pontarelli, Angelo Tulumello, Giuseppe Bianchi 0001
LANMAN4
2017 Demo: Implementing advanced network functions with stateful programmable data planes
abstract
Stateful programmable dataplanes are emerging as a disruptive technology and an enabler factor for network function virtualization in SDN environments. In this demo paper we show three real world use cases with stateful forwarding requirements to validate both the hardware and software implementations of the Open Packet Processor data plane.
Marco Bonola, Roberto Bifulco, Luca Petrucci, Salvatore Pontarelli, Angelo Tulumello, Giuseppe Bianchi 0001
LANMAN4
2017 Smashing SDN "built-in" actions: Programmable data plane packet manipulation in hardware
abstract
Recently, with new hardware architectures such as Reconfigurable Match Tables and languages such as P4, the Software Defined Networking community has started to bring linerate data plane programmatic flexibility inside switching chipsets. Starting from the original OpenFlow's match/action abstraction, most of the work has so far focused on key improvements in matching flexibility. Conversely, the "action" part, i.e. the set of operations (such as encapsulation or header manipulation) performed on packets after the forwarding decision, has received way less attention: the OpenFlow community has limited to standardize the set of supported actions, whereas their implementation has been delegated to each specific vendor/device. Goal of this paper is to move beyond the idea of "atomic", pre-implemented, actions, and rather make them programmable while retaining high speed multi-gbps operation. In this work we propose a domain-specific HW architecture, called Packet Manipulation Processor (PMP), able to efficiently support microprograms implementing such actions. We describe three non trivial use cases (tunneling, NAT, and ARP reply generation), and assess the relevant throughput performance.
Salvatore Pontarelli, Marco Bonola, Giuseppe Bianchi 0001
NetSoft1
2017 On offloading programmable SDN controller tasks to the embedded microcontroller of stateful SDN dataplanes
abstract
This paper presents a method to implement complex tasks into stateful SDN programmable dataplanes. In particular, the presented method proposes to use the internal microcontroller typically used to configure the programmable dataplane also to perform some complex operations that do not require to be executed on each packet. These operations can be executed on a set of data gathered by the dataplane and processed in a time scale that is much higher than the time window of a packet, but is much less than time scale needed for an external SDN controller. Moreover, the use of the configuration microcontroller instead of an external SDN controller to avoid the exchange of data on the control links permits a fine grain tuning of the operations to perform from the timing point of view (precise timestamping, low latency etc.). A set of measurements showing the feasibility of the method are presented, and a simple use case is described to show the effectiveness of this method to enhance the capability of stateful programmable dataplanes.
Salvatore Pontarelli, Valerio Bruschi, Marco Bonola, Giuseppe Bianchi 0001
NetSoft1
2017 StreaMon: A Data-Plane Programming Abstraction for Software-Defined Stream Monitoring
abstract
The fast evolving nature of modern cyber threats and network monitoring needs calls for new, “software-defined”, approaches to simplify and quicken programming and deployment of online (stream-based) traffic analysis functions. StreaMon is a carefully designed data-plane abstraction devised to scalably decouple the “programming logic” of a traffic analysis application (tracked states, features, anomaly conditions, etc.) from elementary primitives (counting and metering, matching, events generation, etc), efficiently pre-implemented in the probes, and used as common instruction set for supporting the desired logic. Multi-stage multi-step real-time tracking and detection algorithms are supported via the ability to deploy custom states, relevant state transitions, and associated monitoring actions and triggering conditions. Such a separation entails platform-independent, portable, online traffic analysis tasks written in a high level language, without requiring developers to access the monitoring device internals and program their custom monitoring logic via low level compiled languages (e.g., C, assembly, VHDL). We validate our design by developing a prototype and a set of simple (but functionally demanding) use-case applications and by testing them over real traffic traces.
Marco Bonola, Giuseppe Bianchi 0001, Giulio Picierro, Salvatore Pontarelli, Marco Monaci
IEEE Trans. Dependable Secur. Comput.4
2016 Improving counting Bloom filter performance with fingerprints
Salvatore Pontarelli, Pedro Reviriego, Juan Antonio Maestro
Inf. Process. Lett.1
2016 Guest Editorial: IEEE Transactions on Computers and IEEE Transactions on Nanotechnology Joint Special Section on Defect and Fault Tolerance in VLSI and Nanotechnology Systems
abstract
The papers in this special issue focus on defect and fault tolerance in VLSI and nanotechnology systems. With the increasing demand for ever smaller, portable, energy-efficient and high-performance electronic systems, scaling of CMOS technology continues. As CMOS scaling approaches physical limits, continued innovation in materials, manufacturing processes, device structures and design paradigms have been necessary. High-k oxide and metal-gate stack were introduced to address oxide leakage; thin body undoped channels, and multiple-gate structures were introduced to mitigate subthreshold leakage; restricted design rules were introduced to improve layout efficiency; yet CMOS technology continues to be challenged in the areas of device aging and reliability. While CMOS is expected to be the dominant semiconductor technology for the foreseeable future, for reasons that are both technological and financial, alternatives to CMOS technology are attracting attention from the researchers.
Cristiana Bolchini, Sandip Kundu, Salvatore Pontarelli
IEEE Trans. Computers3
2016 Parallel d-Pipeline: A Cuckoo Hashing Implementation for Increased Throughput
abstract
Cuckoo hashing has proven to be an efficient option to implement exact matching in networking applications. It provides good memory utilization and deterministic worst case access time. The continuous increase in speed and complexity of networking devices creates a need for higher throughput exact matching in many applications. In this paper, a new Cuckoo hashing implementation named parallel d-pipeline is proposed to increase throughput. The scheme presented is targeted to implementations in which the tables are accessed in parallel. A parallel implementation increases the throughput and therefore is well suited to high speed applications. Parallel schemes are common for ASIC/FPGA implementations in which the tables are stored in several embedded memories. Using the proposed technique, the throughput can be significantly increased with gains that in practical scenarios can reach 60 percent compared to existing parallel implementations. The new scheme has been evaluated using a case study and detailed results for performance and implementation costs are reported.
Salvatore Pontarelli, Pedro Reviriego, Juan Antonio Maestro
IEEE Trans. Computers1
2016 OMASS: One Memory Access Set Separation
abstract
In many applications, there is a need to identify to which of a group of sets an element x belongs, if any. For example, in a router, this functionality can be used to determine the next hop of an incoming packet. This problem is generally known as set separation and has been widely studied. Most existing solutions make use of hash-based algorithms, particularly when a small percentage of false positives is allowed. A known approach is to use a collection of Bloom filters in parallel. Such schemes can require several memory accesses, a significant limitation for some implementations. We propose an approach using Block Bloom Filters, where each element is first hashed to a single memory block that stores a small Bloom filter that tracks the element and the set or sets the element belongs to. In a naive solution, when an element x in a set S is stored, it necessarily increases the false positive probability for finding that x is in another set T. In this paper, we introduce our One Memory Access Set Separation (OMASS) scheme to avoid this problem. OMASS is designed so that for a given element x, the corresponding Bloom filter bits for each set map to different positions in the memory word. This ensures that the false positive rates for the Bloom filters for element x under other sets are not affected. In addition, OMASS requires fewer hash functions compared to the naive solution.
Michael Mitzenmacher, Pedro Reviriego, Salvatore Pontarelli
IEEE Trans. Knowl. Data Eng.3
2015 Stateful OpenFlow: Hardware proof of concept
abstract
This paper presents a hardware implementation of Openstate, an extension of OpenFlow that allows performing stateful control functionalities directly inside the switch, without requiring the intervention of an external controller. The paper shows how, with a minimal reworking of the OpenFlow's basic architecture, and reusing the same building blocks, it is possible to greatly extend the intelligence of an OpenFlow switch allowing the offload of many control task directly in the switch. An FPGA based implementation of an Openstate prototype is here presented, the different architectural design choices are discussed, and the performance and limitations of the developed prototype are examinated. Finally, the paper proposes a discussion on the performance achievable by using an ASIC implementation of the OpenState switch1.
Salvatore Pontarelli, Marco Bonola, Giuseppe Bianchi 0001, Antonio Capone, Carmelo Cascone
HPSR1
2015 Low Delay Single Symbol Error Correction Codes Based on Reed Solomon Codes
abstract
To avoid data corruption, error correction codes (ECCs) are widely used to protect memories. ECCs introduce a delay penalty in accessing the data as encoding or decoding has to be performed. This limits the use of ECCs in high-speed memories. This has led to the use of simple codes such as single error correction double error detection (SEC-DED) codes. However, as technology scales multiple cell upsets (MCUs) become more common and limit the use of SEC-DED codes unless they are combined with interleaving. A similar issue occurs in some types of memories like DRAM that are typically grouped in modules composed of several devices. In those modules, the protection against a device failure rather than isolated bit errors is also desirable. In those cases, one option is to use more advanced ECCs that can correct multiple bit errors. The main challenge is that those codes should minimize the delay and area penalty. Among the codes that have been considered for memory protection are Reed-Solomon (RS) codes. These codes are based on non-binary symbols and therefore can correct multiple bit errors. In this paper, single symbol error correction codes based on Reed-Solomon codes that can be implemented with low delay are proposed and evaluated. The results show that they can be implemented with a substantially lower delay than traditional single error correction RS codes.
Salvatore Pontarelli, Pedro Reviriego, Marco Ottavi, Juan Antonio Maestro
IEEE Trans. Computers1
2015 A Class of SEC-DED-DAEC Codes Derived From Orthogonal Latin Square Codes
abstract
Radiation-induced soft errors are a major reliability concern for memories. To ensure that memory contents are not corrupted, single error correction double error detection (SEC-DED) codes are commonly used, however, in advanced technology nodes, soft errors frequently affect more than one memory bit. Since SEC-DED codes cannot correct multiple errors, they are often combined with interleaving. Interleaving, however, impacts memory design and performance and cannot always be used in small memories. This limitation has spurred interest in codes that can correct adjacent bit errors. In particular, several SEC-DED double adjacent error correction (SEC-DED-DAEC) codes have recently been proposed. Implementing DAEC has a cost as it impacts the decoder complexity and delay. Another issue is that most of the new SEC-DED-DAEC codes miscorrect some double nonadjacent bit errors. In this brief, a new class of SEC-DED-DAEC codes is derived from orthogonal latin squares codes. The new codes significantly reduce the decoding complexity and delay. In addition, the codes do not miscorrect any double nonadjacent bit errors. The main disadvantage of the new codes is that they require a larger number of parity check bits. Therefore, they can be useful when decoding delay or complexity is critical or when miscorrection of double nonadjacent bit errors is not acceptable. The proposed codes have been implemented in Hardware Description Language and compared with some of the existing SEC-DED-DAEC codes. The results confirm the reduction in decoder delay.
Pedro Reviriego, Salvatore Pontarelli, Adrian Evans, Juan Antonio Maestro
IEEE Trans. Very Large Scale Integr. Syst.2
2015 A Synergetic Use of Bloom Filters for Error Detection and Correction
abstract
Bloom filters (BFs) provide a fast and efficient way to check whether a given element belongs to a set. The BFs are used in numerous applications, for example, in communications and networking. There is also ongoing research to extend and enhance BFs and to use them in new scenarios. Reliability is becoming a challenge for advanced electronic circuits as the number of errors due to manufacturing variations, radiation, and reduced noise margins increase as technology scales. In this brief, it is shown that BFs can be used to detect and correct errors in their associated data set. This allows a synergetic reuse of existing BFs to also detect and correct errors. This is illustrated through an example of a counting BF used for IP traffic classification. The results show that the proposed scheme can effectively correct single errors in the associated set. The proposed scheme can be of interest in practical designs to effectively mitigate errors with a reduced overhead in terms of circuit area and power.
Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi
IEEE Trans. Very Large Scale Integr. Syst.2
2015 MCU Tolerance in SRAMs Through Low-Redundancy Triple Adjacent Error Correction
abstract
Static random access memories (SRAMs) are key in electronic systems. They are used not only as standalone devices, but also embedded in application specific integrated circuits. One key challenge for memories is their susceptibility to radiation-induced soft errors that change the value of memory cells. Error correction codes (ECCs) are commonly used to ensure correct data despite soft errors effects in semiconductor memories. Single error correction/double error detection (SEC-DED) codes have been traditionally the preferred choice for data protection in SRAMs. During the last decade, the percentage of errors that affect more than one memory cell has increased substantially, mainly due to multiple cell upsets (MCUs) caused by radiation. The bits affected by these errors are physically close. To mitigate their effects, ECCs that correct single errors and double adjacent errors have been proposed. These codes, known as single error correction/double adjacent error correction (SEC-DAEC), require the same number of parity bits as traditional SEC-DED codes and a moderate increase in the decoder complexity. However, MCUs are not limited to double adjacent errors, because they affect more bits as technology scales. In this brief, new codes that can correct triple adjacent errors and 3-bit burst errors are presented. They have been implemented using a 45-nm library and compared with previous proposals, showing that our codes have better error protection with a moderate overhead and low redundancy.
Luis J. Saiz, Pedro Reviriego, Pedro J. Gil, Salvatore Pontarelli, Juan Antonio Maestro
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Complementary resistive switch based stateful logic operations using material implication
abstract
Memristor based logic and memories are increasingly becoming one of the fundamental building blocks for future system design. Hence, it is important to explore various methodologies for implementing these blocks. In this paper, we present a novel Complementary Resistive Switching (CRS) based stateful logic operations using material implication. The proposed solution benefits from exponential reduction in sneak path current in crossbar implemented logic. We validated the effectiveness of our solution through SPICE simulations on a number of logic circuits. It has been shown that only 4 steps are required for implementing N input NAND gate whereas memristor based stateful logic needs N+1 steps.
Yuanfan Yang, Jimson Mathew, Dhiraj K. Pradhan, Marco Ottavi, Salvatore Pontarelli
DATE5
2014 Improving the performance of Invertible Bloom Lookup Tables
Salvatore Pontarelli, Pedro Reviriego, Michael Mitzenmacher
Inf. Process. Lett.1
2014 A Method to Extend Orthogonal Latin Square Codes
abstract
Error correction codes (ECCs) are commonly used to protect memories from errors. As multibit errors become more frequent, single error correction codes are not enough and more advanced ECCs are needed. The use of advanced ECCs in memories is, however, limited by their decoding complexity. In this context, one-step majority logic decodable (OS-MLD) codes are an interesting option as the decoding is simple and can be implemented with low delay. Orthogonal Latin squares (OLS) codes are OS-MLD and have been recently considered to protect caches and memories. The main advantage of OLS codes is that they provide a wide range of choices for the block size and the error correction capabilities. In this brief, a method to extend OLS codes is presented. The proposed method enables the extension of the data block size that can be protected with a given number of parity bits thus reducing the overhead. The extended codes are also OS-MLD and have a similar decoding complexity to that of the original OLS codes. The proposed codes have been implemented to evaluate the circuit area and delay needed for different block sizes.
Pedro Reviriego, Salvatore Pontarelli, Alfonso Sánchez-Macián, Juan Antonio Maestro
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Error detection in ternary CAMs using bloom filters
abstract
This paper presents an innovative approach to detect soft errors in Ternary Content Addressable Memories (TCAMs) based on the use of Bloom Filters. The proposed approach is described in detail and its performance results are presented. The advantages of the proposed method are that no modifications to the TCAM device are required, the checking is done on-line and the approach has low power and area overheads.
Salvatore Pontarelli, Marco Ottavi, Adrian Evans, Shi-Jie Wen
DATE1
2013 Traffic-Aware Design of a High-Speed FPGA Network Intrusion Detection System
abstract
Security of today's networks heavily rely on network intrusion detection systems (NIDSs). The ability to promptly update the supported rule sets and detect new emerging attacks makes field-programmable gate arrays (FPGAs) a very appealing technology. An important issue is how to scale FPGA-based NIDS implementations to ever faster network links. Whereas a trivial approach is to balance traffic over multiple, but functionally equivalent, hardware blocks, each implementing the whole rule set (several thousands rules), the obvious cons is the linear increase in the resource occupation. In this work, we promote a different, traffic-aware, modular approach in the design of FPGA-based NIDS. Instead of purely splitting traffic across equivalent modules, we classify and group homogeneous traffic, and dispatch it to differently capable hardware blocks, each supporting a (smaller) rule set tailored to the specific traffic category. We implement and validate our approach using the rule set of the well-known Snort NIDS, and we experimentally investigate the emerging trade-offs and advantages, showing resource savings up to 80 percent based on real-world traffic statistics gathered from an operator's backbone.
Salvatore Pontarelli, Giuseppe Bianchi 0001, Simone Teofili
IEEE Trans. Computers1
2013 Error Detection and Correction in Content Addressable Memories by Using Bloom Filters
abstract
A content addressable memory (CAM) is an SRAM-based memory that can be accessed in parallel to search for a given search word, providing as a result the address of the matching data. Like conventional memories, a CAM can be affected by the occurrence of single event upsets (SEUs) that can alter the content of one of more memory cells causing different effects such as pseudo-HIT or pseudo-MISS events. It is well known that, because of the parallel search performed by a CAM during the query of a word, a standard error correction code could not defend it against SEU events. In this paper, we propose a method that does not require any modification to a CAM's internal structure and, therefore, can be easily applied at system level. Error detection is performed by using a probabilistic structure called "Bloom filter,” which can signal if a given data is present in the CAM. Bloom filters permit to efficiently store and query the presence of data in a set. But, while a CAM suffers from SEU induced errors, the probabilistic nature of Bloom filters has as a consequence the so called false-positive effect. This paper shows that, by combining the use of a Bloom filter with a CAM, the complementary limitations of these modules can be compensated. The combined use of a CAM and a Bloom filter is analyzed in different cases, showing that the proposed technique can be implemented with a low penalty in terms of area and power consumption.
Salvatore Pontarelli, Marco Ottavi
IEEE Trans. Computers1
2013 Low Complexity Concurrent Error Detection for Complex Multiplication
abstract
This paper studies the problem of designing a low complexity Concurrent Error Detection (CED) circuit for the complex multiplication function commonly used in Digital Signal Processing circuits. Five novel CED architectures are proposed and their computational complexity, area, and delay evaluated in several circuit implementations. The most efficient architecture proposed reduces the number of gates required by up to 30 percent when compared with a conventional CED architecture based on Dual Modular Redundancy. Compared to a Residue Code CED scheme, the area of the proposed architectures is larger. However, for some of the proposed CEDs delay is significantly lower with reductions exceeding 30 percent in some configurations.
Salvatore Pontarelli, Pedro Reviriego, Chris J. Bleakley, Juan Antonio Maestro
IEEE Trans. Computers1
2013 A Method to Construct Low Delay Single Error Correction Codes for Protecting Data Bits Only
abstract
Error correction codes (ECCs) have been used for decades to protect memories from soft errors. Single error correction (SEC) codes that can correct 1-bit error per word are a common option for memory protection. In some cases, SEC codes are extended to also provide double error detection and are known as SEC-DED codes. As technology scales, soft errors on registers also became a concern and, therefore, SEC codes are used to protect registers. The use of an ECC impacts the circuit design in terms of both delay and area. Traditional SEC or SEC-DED codes developed for memories have focused on minimizing the number of redundant bits added by the code. This is important in a memory as those bits are added to each word in the memory. However, for registers used in circuits, minimizing the delay or area introduced by the ECC can be more important. In this paper, a method to construct low delay SEC or SEC-DED codes that correct errors only on the data bits is proposed. The method is evaluated for several data block sizes, showing that the new codes offer significant delay reductions when compared with traditional SEC or SEC-DED codes. The results for the area of the encoder and decoder also show substantial savings compared to existing codes.
Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Concurrent Error Detection for Orthogonal Latin Squares Encoders and Syndrome Computation
abstract
Error correction codes (ECCs) are commonly used to protect memories against errors. Among ECCs, orthogonal latin squares (OLS) codes have gained renewed interest for memory protection due to their modularity and the simplicity of the decoding algorithm that enables low delay implementations. An important issue is that when ECCs are used, the encoder and decoder circuits can also suffer errors. In this brief, a concurrent error detection technique for OLS codes encoders and syndrome computation is proposed and evaluated. The proposed method uses the properties of OLS codes to efficiently implement a parity prediction scheme that detects all errors that affect a single circuit node.
Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Feedback based droop mitigation
abstract
A strong dl/dt event in a VLSI circuit can induce a temporary voltage drop and consequent malfunctioning of logic as for instance failing speed paths. This event, called power droop, usually manifests itself in at-speed scan test where a surge in switching activity (capture phase) follows a period of quiescent circuit state (shift phase). Power droop is also present during mission mode operation. However, because of the less predictable occurrence of the switching events in mission mode, usually the values of power droop measured during test are different from those measured in mission mode. To overcome the power droop problem, different mitigation techniques have been proposed. The goal of these techniques is to create a uniform current demand throughout the test. This paper proposes a feedback based droop mitigation technique which can adapt to the droop by reading the level of VDD and modifying real time the current flowing on ad-hoc droop mitigators. It is shown that the proposed solution not only can compensate for droop events occurring during test mode but also can be used as a method of mission mode droop mitigation and yield enhancement if higher power consumption is acceptable.
Salvatore Pontarelli, Marco Ottavi, Adelio Salsano, Kamran Zarrineh
DATE1
2010 Exploiting Dynamic Reconfiguration for FPGA Based Network Intrusion Detection Systems
abstract
A Network Intrusion Detection System (NIDS) inspects the traffic flowing in a network to detect malicious content such as spam, viruses, and so on. Hardware based solutions appear necessary to face the performance requirements emerging when the goal is to deploy such systems in high speed network scenarios. However, the appropriate choice of the hardware platform is believed to be subject to at least two requirements, usually considered independent each other: i) it needs to be reprogrammable, in order to update the intrusion detection rules each time a new threat arises, and ii) it must be capable of containing the typically very large set of rules of existing NIDSs. The goal of this paper is to show that reprogrammability can be further exploited to reduce the resource requirements for the chosen platform. Specifically, we propose an FPGA-based solution that classifies and dispatches traffic to elastic buffers, connecting one buffer at a time to a dynamically reconfigurable rule matching core. This core supports only the appropriate subset of detection rules. A worst-case analysis shows that the saving in hardware resources is achieved with a relatively small buffer space, currently available in cheap, low end, FPGA boards, with no impairment on the resulting throughput.
Salvatore Pontarelli, Claudio Greco 0004, Enrico Nobile, Simone Teofili, Giuseppe Bianchi 0001
FPL1
2009 Error detection in addition chain based ECC Point Multiplication
abstract
In this paper the problem of error detection in elliptic curve point multiplication is faced. Elliptic Curve Point Multiplication is often used to design cryptographic algorithms that use fewer bits than other methods with the same security level. One of the mode used to break the security of cryptosystem is the injection of a fault in the hardware realizing the cryptographic algorithm. Therefore, to avoid this kind of attack, is very important to develop cryptosystems that are able to detect errors induced by a fault. The paper takes into account the algorithm for elliptic Curve Point Multiplication based on a sequence of additions called ldquoaddition chainrdquo and shows how suitable modifications of the algorithms used for computing the point multiplication adds the error detection property to the algorithm.
Salvatore Pontarelli, Gian Carlo Cardarilli, Marco Re, Adelio Salsano
IOLTS1
2008 Totally Fault Tolerant RNS Based FIR Filters
abstract
In this paper, the design of a finite impulse response (FIR) filter with fault tolerant capabilities based on the residue number system is analyzed. Differently from other approaches that use RNS, the filter implementation is fault tolerant not only with respect to a fault inside the RNS moduli, but also in the reverse converter. An architecture allowing fault masking in the overall RNS FIR filter is presented. It avoids the use of a trivial triple modular redundancy (TMR) to protect the blocks that performs the final stages of the RNS based FIR computation.
Salvatore Pontarelli, Gian Carlo Cardarilli, Marco Re, Adelio Salsano
IOLTS1
2008 Analysis and Evaluations of Reliability of Reconfigurable FPGAs
Salvatore Pontarelli, Marco Ottavi, Vamsi Vankamamidi, Gian Carlo Cardarilli, Fabrizio Lombardi, Adelio Salsano
J. Electron. Test.1
2007 Self Checking Circuit Optimization by means of Fault Injection Analysis: A Case Study on Reed Solomon Decoders
abstract
This paper shows how the use of exhaustive fault injection campaigns in conjunction with the analysis of the property of a circuit, allows to improve the efficiency of the checker of self checking circuits. Experimental results coming from fault injection campaigns on a Reed-Solomon decoder demonstrated that by observing the occurred errors and the correspondent detection module has been possible to reduce the number of detection module, while paying a small reduction of the percentage of SEUs that can be detected.
Salvatore Pontarelli, Luca Sterpone, Gian Carlo Cardarilli, Marco Re, Matteo Sonza Reorda, Adelio Salsano, Massimo Violante
IOLTS1
2007 QCA Circuits for Robust Coplanar Crossing
Sanjukta Bhanja, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli
J. Electron. Test.4
2007 Analysis of Errors and Erasures in Parity Sharing RS Codecs
abstract
Reed Solomon (RS) codes are widely used to protect information from errors in transmission and storage systems. Most of the RS codes are based on GF(28) Galois Fields and use a byte to encode a symbol providing codewords up to 255 symbols. Codewords with more than 255 symbols can be obtained by using GF(2m) Galois fields with m > 8, but this choice increases the complexity of the encoding and decoding algorithms. This limitation can be superseded by introducing parity sharing (PS) RS codes that are characterized by a greater flexibility in terms of design parameters. Consequently, a designer can choose between different PS code implementations in order to meet requirements such as bit error rate (BER), hardware complexity, speed, and throughput. This paper analyzes the performance of PS codes in terms of BER with respect to the code parameters, taking into account either random error or erasure rates as two independent probabilities. This approach provides an evaluation that is independent of the communication channel characteristics and extends the results to memory systems in which permanent faults and transient faults can be modeled, respectively, as erasures and random errors. The paper also provides hardware implementations of the PS encoder and decoder and discusses their performances in terms of hardware complexity, speed, and throughput.
Gian Carlo Cardarilli, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Computers2
2007 Concurrent Error Detection in Reed-Solomon Encoders and Decoders
abstract
Reed-Solomon (RS) codes are widely used to identify and correct errors in transmission and storage systems. When RS codes are used for high reliable systems, the designer should also take into account the occurrence of faults in the encoder and decoder subsystems. In this paper, self-checking RS encoder and decoder architectures are presented. The RS encoder architecture exploits some properties of the arithmetic operations in GF(2m). These properties are related to the parity of the binary representation of the elements of the Galois field. In the RS decoder, the implicit redundancy of the received codeword, under suitable assumptions explained in this paper, allows implementing concurrent error detection schemes useful for a wide range of different decoding algorithms with no intervention on the decoder architecture. Moreover, performances in terms of area and delay overhead for the proposed circuits are presented.
Gian Carlo Cardarilli, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Novel designs for thermally robust coplanar crossing in QCA
abstract
In this paper, different circuit arrangements of quantum-dot cellular automata (QCA) are proposed for the so-called coplanar crossing. These arrangements exploit the majority voting properties of QCA to allow a robust crossing of wires on the Cartesian plane. This is accomplished using enlarged lines and voting. Using a Bayesian network (BN) based simulator, new results are provided to evaluate the robustness to so-called kink of these arrangements to thermal variations. The BN simulator provides fast and reliable computation of the signal polarization versus normalized temperature. It is shown that by modifying the layout, a higher polarization level can be achieved in the routed signal by utilizing the proposed QCA arrangements
Sanjukta Bhanja, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli
DATE4
2006 Localization of Faults in Radix-n Signed Digit Adders
abstract
It is widely known that an adder can be checked by using check symbols that are residues of the numbers modulo some base. This paper extends this characteristic to a radix r signed digit (SD) representation. The confinement of the carry operation can also be exploited to localize the faulty resources in the SD adder and to reconfigure the adder in order to work with a reduced dynamic range. The fault localization procedure is presented in this paper and the reconfiguration the SD adder after the fault localization is discussed
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IOLTS3
2006 Concurrent error detection in Reed Solomon decoders
abstract
Reed Solomon codes are widely used to identify and correct data errors in transmission and storage systems. When Reed Solomon (RS) codes are used for high reliable systems, the designer should take into account also for the occurrence of faults in the encoder and decoder blocks. In this paper a method to obtain a self-checking RS decoder is presented and different architectures for its implementation based on concurrent error detection are provided. The proposed method can be used for a wide range of different decoder algorithms with no intervention on the decoder architecture
Gian Carlo Cardarilli, Salvatore Pontarelli, Marco Re, Adelio Salsano
ISCAS2
2006 Fault tolerant design of signed digit based FIR filters
abstract
This paper proposes a methodology for the development of fault tolerant arithmetic circuits using an r-radix signed digit (SD) representation. A residue checking based technique is applied to detect errors caused by faults belonging to the considered stuck-at fault set. We developed a technique to detect the presence of a fault in the adder by using a completely independent circuit that uses check symbols that are residues of the numbers modulo a suitable base. We show that for any radix r > 3 the correctness of the adder operation can be checked simply by using two check symbols. This property is used to extend the fault detection also to SD constant multipliers and to the rounding operation. The extension to SD constant multiplier can be easily obtained by implementing the multipliers using a shift and add architecture. For the truncation operation we notice that many error detection techniques based on some arithmetic properties of the circuits fails when the truncation operation is performed. Instead, we show how our method can be applied to this operation with low area overhead and therefore it is useful to implement self-checking finite impulse response (FIR) filters
Gian Carlo Cardarilli, Salvatore Pontarelli, Marco Re, Adelio Salsano
ISCAS2
2006 Fault Localization, Error Correction, and Graceful Degradation in Radix 2 Signed Digit-Based Adders
abstract
In this paper, a methodology for the development of fault-tolerant adders based on the radix 2 signed digit (SD) representation is presented. The use of a number representation characterized by a carry propagation confined to neighbor digits implies interesting advantages in terms of error detection, fault localization, and repair. Errors caused by faults belonging to a considered stuck-at fault set can be detected by a parity-based technique. In fact, a carry-free adder preserving the parity of the augends can be implemented allowing fault detection by using a parity checker. Regarding fault localization, the "carry-free" property of the adder ensures the confinement of the error due to a permanent fault to only few digits. The detection of the faulty digit has been obtained by using a recomputation with shifted operands method. Finally, after the fault localization, graceful degradation of the system intended as the reduction of the performances versus a correct output computation can be obtained by using two different procedures. The first one allows obtaining the correct output by recomputing the result performing two different shift operations and using the intersection of the obtained results to recover the correct output, while the second one is based on a reduced dynamic range approach, which allows us to obtain the result in only one step, but with fewer output digits.
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Computers3
2005 On the Analysis of Reed Solomon Coding for Resilience to Transient/Permanent Faults in Highly Reliable Memories
abstract
Single event upsets (SEU), as well as permanent faults, can significantly affect the correct on-line operation of digital systems, such as memories and microprocessors; a memory can be made resilient to permanent and transient faults by using modular redundancy and coding. Different memory systems are compared; these systems utilize simplex and duplex arrangements with a combination of Reed Solomon coding and scrubbing. The memory systems and their operations are analyzed by novel Markov chains to characterize the performance for dynamic reconfiguration as well as error detection and correction under the occurrence of permanent and transient faults. For a specific Reed Solomon code, the duplex arrangement is able to cope efficiently with the occurrence of permanent faults, while the use of scrubbing allows it to cope with transient faults.
Luca Schiano, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli, Adelio Salsano
DATE4
2005 Design of a Self Checking Reed Solomon Encoder
abstract
In this paper, an innovative self-checking Reed Solomon encoder architecture is described. The presented architecture exploits some properties of the arithmetic operations in GF(2/sup 3/) related to the parity of the binary representation of the field elements. Moreover, a method for introducing self-checking capabilities on all the arithmetic structures used in the Reed Solomon encoder is presented. Finally the self-checking encoder architecture has been mapped on a FPGA evaluating its area overhead.
Gian Carlo Cardarilli, Salvatore Pontarelli, Marco Re, Adelio Salsano
IOLTS2
2005 A Comparative Evaluation of Designs for Reliable Memory Systems
Gian Carlo Cardarilli, Fabrizio Lombardi, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
J. Electron. Test.4
2004 A Signed Digit Adder with Error Correction and Graceful Degradation Capabilities
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IOLTS3
2003 Design of a fault tolerant solid state mass memory
abstract
This paper describes a novel architecture of fault tolerant solid state mass memory (SSMM) for satellite applications. Mass memories with low-latency time, high throughput, and storage capabilities cannot be easily implemented using space qualified components, due to the inevitable technological delay of these kind of components. For this reason, the choice of commercial off the shelf (COTS) components is mandatory for this application. Therefore, the design of an electronic system for space applications, based on commercial components, must match the reliability requirements using system level methodologies. In the proposed architecture, error-correcting codes are used to strengthen the commercial dynamic random access memory (DRAM) chips, while the system controller is developed by applying fault tolerant design solutions. The main features of the SSMM are the dynamic reconfiguration capability, and the high performances which can be gracefully reduced in case of permanent faults, maintaining part of the system functionality. The paper shows the system design methodology, the architecture, and the simulation results of the SSMM. The properties of the building blocks are described in detail both in their functionality and fault tolerant capabilities. A detailed analysis of the system reliability and data integrity is reported. The graceful degradation capability of our system allows different levels of acceptable performances, in terms of active I/O link interfaces and storage capability. The results also show that the overall reliability of the SSMM is almost the same using different RS coding schemes, allowing a dynamic reconfiguration of the coding to reduce the latency (shorter codewords), or to improve the data integrity (longer codewords). The use of a scrubbing technique can be useful if a high SEU rate is expected, or if the data must be stored for a long period in the SSMM. The reported simulations show the behavior of the SSMM in presence of permanent and transient faults. In fact, we show that the SCU is able to recover from transient faults. On the other hand, using a spare microcontroller also hard faults can be tolerated. The distributed file system confines the unrecoverable fault effects only in a single I/O Interface. In this way, the SSMM maintains its capability to store and read data. The proposed system allows obtaining SSMM characterized by high reliability and high speed due the intrinsic parallelism of the switching matrix.
Gian Carlo Cardarilli, A. Leandri, P. Marinucci, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Reliab.5