EDBT 2026 Demo / reviewers in the wild / expert
Ricardo Chaves
dblp:39/2903
· DBLP profile ↗
45ranked-venue papers
10as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 8 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AEGIS+AES folded architecture for FPGAabstractCryptography and data security are a main concern on the current digital communication world, but authenticated encryption (AE) protocols have been under performing. The CAESAR competition provided the opportunity for new proposals and discussions, that now reached the finalist stage with the AEGIS-128 and -256 authenticated ciphers. The work herein proposed takes into account the AEGIS implementation on FPGAs considering architectural trade-offs towards compact, efficient, and flexible implementations. A variation of the proposed solution also allows for the computation of both AEGIS and AES. This work considers a round-rolled 5/6-folded structure capable of processing both AEGIS-128 and -256 seamlessly, by carefully scheduling the 5 to 6 128-bit sub-states of the cipher. Experimental results suggest a maximum throughput of 5.9 and 4.9 Gbps for AEGIS-128 and AEGIS-256, respectively, on an AMD/Xilinx Zynq-7020. Furthermore, even though this design can process both AEGIS-128 and -256 variants (the first in the state-of-the-art) its occupation of only 496 Slices, makes it the most compact (and one of the most efficient) solutions on FPGAs, requiring $31.7 \%$ less resources than the smallest AEGIS-128 exclusive competitor. This is achieved with a cost of $16.6 \%$ less efficiency than the most efficient AEGIS-128 only design. João Carlos Resende, Ricardo Chaves |
DSD | 2 |
| 2025 | A Fast Parallel Decoder for LDPC Codes Suitable for Burst Errors in QKD ReconciliationabstractQuantum Key Distribution (QKD) relies on quantum mechanics to enable secure communication, with the reconciliation phase playing an essential role in correcting discrepancies in shared keys caused by noise. Conventional Low-Density Parity-Check (LDPC) codes, while robust against random errors, often struggle with burst errors arising from localized noise sources in practical QKD implementations. To address this challenge, we introduce specialized uniform LDPC codes over Galois fields explicitly designed to correct burst errors effectively. We further present a highly efficient parallel decoder for these LDPC codes, operating within the complexity class NC (Nick’s Class), making it particularly suitable for GPU-based implementation. Our decoder provides guaranteed high-performance reconciliation as long as the number of burst errors remains below a designated threshold. Additionally, we discuss an extension to our decoding approach: when configured to handle higher error rates, the algorithm becomes a Las Vegas algorithm, introducing a small but controlled probability of failure. Finally, we propose a basic algorithm for generating these specialized LDPC codes and highlight prospective optimizations required for constructing larger matrices. Duarte Mateus, Ricardo Chaves |
ITW | 2 |
| 2025 | Energy-Aware Adaptive Security for Smart Farming (EAASF): A Hybrid IDS-IPS Framework with SDN-Orchestrated for Agriculture 4.0abstractSmart farming efficiency has increased through IoT adoption, but the technology raises crucial security and privacy threats that affect resource-limited devices. Smart farming networks experience vulnerability to cyber threats like unauthorized access, data tampering, and Distributed Denial-of-Service attacks because standard security approaches do not provide adaptive protection and energy-efficient security. This paper proposes the Energy-Aware Adaptive Security Framework (EAASF), a multilayered, intelligent security approach designed to safeguard IoTdriven agriculture. The framework is dynamic in terms of security settings, depending on real-time threats and device energy capabilities, and Amount of processing capacity through hybrid Intrusion Detection Systems (IDS), Intrusion Prevention Systems (IPS), Software-Defined Networking (SDN), fog computing. This dynamic security system offers greater protection with minimal waste of resources; thus, it is suitable in rural and energy-limited settings. Focusing on the combined problem of security and energy efficiency, EAASF will increase the cyber resilience of smart farming, the integrity of data and network security, and contribute to sustainable agriculture. Seyed Jamal Mirsadri, Ricardo Chaves, Luis Pedrosa |
NCA | 2 |
| 2024 | Security Layers and Related Services within the Horizon Europe NEUROPULS ProjectabstractIn the contemporary security landscape, the incorporation of photonics has emerged as a transformative force, unlocking a spectrum of possibilities to enhance the resilience and effectiveness of security primitives. This integration represents more than a mere technological augmentation; it signifies a paradigm shift towards innovative approaches capable of delivering security primitives with key properties for low-power systems. This not only augments the robustness of security frameworks, but also paves the way for novel strategies that adapt to the evolving challenges of the digital age. This paper discusses the security layers and related services that will be developed, modeled, and evaluated within the Horizon Europe NEUROPULS project. These layers will exploit novel implementations for security primitives based on physical un-clonable functions (PUFs) using integrated photonics technology. Their objective is to provide a series of services to support the secure operation of a neuromorphic photonic accelerator for edge comnuting applications. Fabio Pavanello, Cédric Marchand 0002, Paul Jiménez, Xavier Letartre, Ricardo Chaves, Niccolò Marastoni, Alberto Lovato, Mariano Ceccato, George Papadimitriou 0001, Vasileios Karakostas, Dimitris Gizopoulos, Roberta Bardini, Tzamn Melendez Carmona, Stefano Di Carlo, Alessandro Savino 0001, Laurence Lerch, Ulrich Rührmair, Sergio Vinagrero Gutierrez, Giorgio Di Natale, Elena I. Vatajelu |
DATE | 5 |
| 2024 | External Memory Protection on FPGA-Based Embedded SystemsabstractWith the proliferation and increased capabilities of embedded systems, they become more exposed and easier targets to attacks. This includes the external components such as DRAM, particularly vulnerable to attacks, especially regarding unauthorized access to the stored data. With the goal of increasing storage security, this paper proposes a memory bridge evaluation platform and an improvement to the authenticated-encryption of off-chip memory, minimizing critical memory access latency. Most state-of-the-art works either use custom high-latency solutions, or frequently recur exclusively to AES-GCM standard for Authenticated Encryption with Associated Data. Besides AES-GCM, this work explores and implements different protection solutions, including NOEKEON-GCM and AEGIS-128L. This work also explores how much of the algorithms can be pre-computed between memory transmissions, by removing the address and data from critical path computations, placing it in the AAD field. The presented prototypes were evaluated on a Xilinx Zynq-7100, showing that the proposed platform allows to analyse and assess different approaches. The obtained experimental results suggest that with a careful selection of algorithms and implementations, memory access latency improvements up to 56% can be achieved in regard to equivalent AES-GCM designs. João Carlos Resende, Aleksandar Ilic, Ricardo Chaves |
DSD | 3 |
| 2023 | EUROPULS: NEUROmorphic energy-efficient secure accelerators based on Phase change materials aUgmented siLicon photonicSabstractThis special session paper introduces the Horizon Europe NEUROPULS project, which targets the development of secure and energy-efficient RISC-V interfaced neuromorphic accelerators using augmented silicon photonics technology. Our approach aims to develop an augmented silicon photonics platform, an FPGA-powered RISC-V-connected computing platform, and a complete simulation platform to demonstrate the neuromorphic accelerator capabilities. In particular, their main advantages and limitations will be addressed concerning the underpinning technology for each platform. Then, we will discuss three targeted use cases for edge-computing applications: Global National Satellite System (GNSS) anti-jamming, autonomous driving, and anomaly detection in edge devices. Finally, we will address the reliability and security aspects of the stand-alone accelerator implementation and the project use cases. Fabio Pavanello, Cédric Marchand 0002, Ian O'Connor, Régis Orobtchouk, Fabien Mandorlo, Xavier Letartre, Sébastien Cueff, Elena I. Vatajelu, Giorgio Di Natale, Benoit Cluzel, Aurelien Coillet, Benoît Charbonnier, Pierre Noe, Frantisek Kavan, Martin Zoldak, Michal Szaj, Peter Bienstman, Thomas Van Vaerenbergh, Ulrich Rührmair, Paulo F. Flores, Luís Guerra e Silva, Ricardo Chaves, Luís Miguel Silveira, Mariano Ceccato, Dimitris Gizopoulos, George Papadimitriou 0001, Vasileios Karakostas, Axel Brando, Francisco J. Cazorla, Ramon Canal, Pau Closas, Adria Gusi-Amigo, Paolo Crovetti, Alessio Carpegna, Tzamn Melendez Carmona, Stefano Di Carlo, Alessandro Savino 0001 |
ETS | 22 |
| 2023 | Lightweight Network-Based IoT Device Authentication in Cloud ServicesabstractInternet of Things (IoT) devices can be divided into two main categories: resource-rich devices such as smart TVs, fitness machines and connected vehicles; and resource-constrained such as low-battery body sensors, pacemakers or bridge monitoring sensors. This work focuses on the second type of devices, which usually have low computing power and/or demanding energy-consumption restrictions (e.g. unable to be charged frequently). It presents three lightweight approaches to provide constrained IoT devices with strong authentication by leveraging the trust in the cellular network and associated standard authentication mechanisms. Furthermore, leveraging the secure communication between cellular devices and the core network, two proposed solutions introduce as well a core network broker that secures the communication of constrained IoT devices that don't have capabilities establish/use secure channels. Tomás Silva, João Casal, Ricardo Chaves |
ICNP | 3 |
| 2023 | Content distribution in a VANET using InterPlanetary file system
Ricardo Chaves, Carlos R. Senna, Miguel Luís, Susana Sargento, Ricardo Matos, Diogo Recharte |
Wirel. Networks | 1 |
| 2022 | SmartFusion2 SoC as a security module for the IoT worldabstractDedicated computational devices such as HSMs and FPGAs are frequently used to provide data security and privacy. However, these options have several drawbacks, particularly when considering IoT environments. HSMs offer high-grade services but are costly and lack application flexibility, while FPGAs, in general, are cheaper and adaptable, but lack security services and protection. Herein, the SmartFusion2 SoC FPGA, a security-oriented system, is evaluated as a possible low-cost and flexible platform for security modules for the IoT. This work analyzes the several security services of the SmartFusion2 SoC, their advantages, and possible trade-offs. To demonstrate the SoC viability as a security module and/or a more adaptable HSM alternative, several case study applications are considered and analyzed to elaborate on the potential, limitations, and mitigations of the latter. Alexandre Rodrigues 0006, João Carlos Resende, Ricardo Chaves |
CF | 3 |
| 2020 | TBOX-Based Mask Scrambling Against SCAabstractIn the last years Side-Channel Attacks have become a significant threat against security devices. Given this, several countermeasures have been proposed, ranging from reducing the leaked power consumption to masking schemes. However, these solutions imply a cost, typically in terms of resources, performance, and power consumption. This work re-adapts the masking scheme of Block Memory Content Scrambling (BMS) to the AES Look-up tables for System-on-Chip $(\mathrm {S}\mathrm {o}\mathrm {C})$ FPGAs, namely the SmartFusion 2 FPGA. The solution is further improved resource-wise by making use of the embedded ARM Cortex-M3 processor for updating the masks. João Carlos Resende, Ricardo J. R. Maçãs, Ricardo Chaves |
FCCM | 3 |
| 2020 | Mask Scrambling Against SCA on Reconfigurable TBOX-Based AESabstractIn the last years Side-Channel Attacks have become a significant threat against security devices. Given this, several countermeasures have been proposed, ranging from reducing the leaked power consumption to masking schemes. However, these solutions imply a cost, typically in terms of resources, performance, and power consumption. This paper focuses on the deployment of masking to the AES computation supported on re-configurable technologies, in this particular case on a SmartFusion 2 SoC and its FPGA fabric and embedded ARM Cortex-M3 processor. This work proposes a novel masking scheme using Auxiliary Random Tables (RBoxes) to further harden the protection against SCA by not only extending the set of used random masks, but also by improving the update frequency of the mask sets. The implementation results suggest that the existing related masking schemes can be deployed at a cost of 645 additional LUTs, 16 μSRAMs, and no additional Large SRAMs, whilst achieving the same operating frequency. João Carlos Resende, Ricardo J. R. Maçãs, Ricardo Chaves |
FPL | 3 |
| 2020 | Hamming-Code Based Fault Detection Design Methodology for Block CiphersabstractFault injection, in particular Differential Fault Analysis (DFA), has become one of the main methods for exploiting vulnerabilities into the block ciphers currently used in a multitude of applications. In order to minimize this type of vulnerabilities, several mechanisms have been proposed to detect this type of attacks. However, these mechanisms can have a significant cost or not adequately cover the implementations against fault attacks. In this paper a novel approach is proposed, consisting in generating the signatures of the internal state using a Hamming code. This allows to cover a larger amount of faults allowing to detect even or odd bit changes, as well as multi-bit and multi-byte changes, the ones that make ciphers more vulnerable to DFA attacks. As case of study, this approach has been applied to the Advanced Encryption Standard (AES) block cipher implemented on FPGA using T-boxes. The results suggest a higher fault coverage with an overhead of 16% of resource consumption and without any penalty in the frequency degradation. Francisco Eugenio Potestad-Ordóñez, Erica Tena, Ricardo Chaves, Manuel Valencia-Barrero, Antonio J. Acosta 0001, Carlos Jesús Jiménez-Fernández |
ISCAS | 3 |
| 2017 | Area-optimized montgomery multiplication on IGLOO 2 FPGAsabstractThis paper presents the first area-optimized Montgomery modular multiplication module on low-power reconfigurable IGLOO® 2 FPGAs, from Microsemi. In order to obtain a good response time with few resources, the FPGA pipelined Math blocks and the embedded memory blocks are fully leveraged. As a result, 256-bit modular multiplications can be done in 2.33 μs, at a cost of 505 LUT4 cells, 257 Flip Flops, 1 Math block and 1 64×18 RAM block. If more area resources are considered, a modular multiplication can be performed in 1.25 μ8 at a cost of 680 LUT4s, 341 Flip Flops, 2 Math blocks and 2 64×18 RAM blocks. This work is the first fundamental step towards area-efficient public-key cryptography on the Microsemi IGLOO® 2 FPGAs. Pedro Maat Costa Massolino, Lejla Batina, Ricardo Chaves, Nele Mentens |
FPL | 3 |
| 2016 | Storekeeper: A Security-Enhanced Cloud Storage Aggregation ServiceabstractCloud storage services are currently a commodity that allows users to store data persistently, access the data from everywhere, and share it with friends or co-workers. However, due to the proliferation of cloud storage accounts and lack of interoperability between cloud services, managing and sharing cloud-hosted files is a nightmare for many users. To address this problem, specialized cloud aggregator systems emerged that provide users a global view of all files in their accounts and enable file sharing between users from different clouds. Such systems, however, have limited security: not only they fail to provide end-to-end privacy from cloud providers, but they require users to grant full access privileges to individual cloud storage accounts. In this paper, we present Storekeeper, a privacy-preserving cloud aggregation service that enables file sharing on multi-user multi-cloud storage platforms while preserving data confidentiality from cloud providers and from the cloud aggregator service. To provide this property, Storekeeper decentralizes most of the cloud aggregation logic to the client side enabling security sensitive functions to be performed only on the trusted client endpoints. This decentralization brings new challenges related with file update propagation, access control, user authentication, and key management that are addressed by Storekeeper. This is provided at a low cost (7% on average) when compared with the underlining cloud providers. Sancha Pereira, André Alves, Nuno Santos 0001, Ricardo Chaves |
SRDS | 4 |
| 2016 | Method for designing two levels RNS reverse converters for large dynamic ranges
Héctor Pettenghi, Ricardo Chaves, Roberto de Matos, Leonel Sousa |
Integr. | 2 |
| 2016 | Compact and On-the-Fly Secure Dynamic Reconfiguration for Volatile FPGAsabstractThe dynamic partial reconfiguration functionality of FPGAs can be attacked, particularly when the FPGA is remotely located or the configuration bitstreams are sent through insecure networks. The existing FPGA technologies provide some built-in security mechanisms; however, these are often inadequate. The existing solutions still impose a significant impact on the reconfiguration process and on the available resources. This article proposes a solution to improve the security of dynamic partial reconfiguration of FPGAs, without significantly affecting the reconfiguration performance. The proposed solution changes the encryption key of the remotely received bitstream by a randomly generated key, unique for each configuration, when storing them in the external unsecured memory. The native frame-wise error detection mechanism combined with an additional CBC-MAC authentication mechanism, allows for an improved countermeasure against replay attack and wrongful bitstream usage. The proposed solution introduces an overhead of 1% of the available resources on the target FPGA and provides the lowest impact on the reconfiguration process when compared to the state of the art, achieving a reconfiguration throughput of 2.5Gbps. Regarding the built-in security mechanism provided by the Xilinx FPGAs, the solution herein proposed provides better security and improves the reconfiguration performance by more than 3 times. Hirak J. Kashyap, Ricardo Chaves |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | CLEFIA Implementation with Full Key ExpansionabstractIn this paper a compact and high throughput hardware structure is proposed allowing for the computation of the novel 128-bit CLEFIA encryption algorithm and its associated full key expansion. In the existing state of the art only the 128-bit key schedule is supported, given the needed modification to the CLEFIA Feistel network. This work shows that with a small area cost and with no performance impact, full key expansion can be supported. This is achieved by using addressable shift registers, available in modern FPGAs, and adaptable scheduling, allowing to compute the 4 and 8 branch CLEFIA Feistel network within the same structure. The obtained experimental results suggest that throughputs above 1 Gbps can be achieved with a low area cost, while achieving efficiency metrics above those of the restricted state of the art. João Carlos Bittencourt, João Carlos Resende, Wagner Luiz Alves de Oliveira, Ricardo Chaves |
DSD | 4 |
| 2015 | Compact dual block AES core on FPGA for CCM ProtocolabstractThis paper presents a compact and FPGA based implementation of the AES encryption standard, specifically designed for processing two independent 128-bit input blocks in feedback modes. This configuration is particularly focused on the Counter with CBC-MAC Protocol, but can also be adapted to other AES based encryption-authentication protocols requiring the processing of two independent data streams. Most of the state of the art solutions implementing CCMP consider large datapaths, sometimes with separated encryption datapaths for the different data streams, leading to low resource efficiency. The work herein proposed suggests that with adequate FPGA component usage and with proper data scheduling a very compact and efficient dual AES core can be derived particularly on FPGAs. Overall, the proposed structure allows for a throughput of 1.7Gbps while achieving a Throughput/Slice efficiency of 24.22 Mbps/Slice, 47% higher than the existing related state of the art. João Carlos Resende, Ricardo Chaves |
FPL | 2 |
| 2015 | Challenges in designing trustworthy cryptographic co-processorsabstractSecurity is becoming ubiquitous in our society. However, the vulnerability of electronic devices that implement the needed cryptographic primitives has become a major issue. This paper starts by presenting a comprehensive overview of the existing attacks to cryptography implementations. Thereafter, the state-of-the-art on some of the most critical aspects of designing cryptographic co-processors are presented. This analysis starts by considering the design of asymmetrical and symmetrical cryptographic primitives, followed by the discussion on the design and online testing of True Random Number Generation. To conclude, techniques for the detection of Hardware Trojans are also discussed. Ricardo Chaves, Giorgio Di Natale, Lejla Batina, Shivam Bhasin, Baris Ege, Apostolos P. Fournaris, Nele Mentens, Stjepan Picek, Francesco Regazzoni 0001, Vladimir Rozic, Nicolas Sklavos 0001, Bohan Yang 0001 |
ISCAS | 1 |
| 2015 | Arithmetic-Based Binary-to-RNS Converter Modulo {2n±k} for jn-bit Dynamic RangeabstractIn this brief, a read-only-memoryless structure for binaryto-residue number system (RNS) conversion modulo (2n±k} is proposed. This structure is based only on adders and constant multipliers. This brief is motivated by the existing (2n± k} binary-to-RNS converters, which are particular inefficient for larger values of n. The experimental results obtained for 4n and 8n bits of dynamic range suggest that the proposed conversion structures are able to significantly improve the forward conversion efficiency, with an AT metric improvement above 100%, regarding the related state of the art. Delay improvements of 2.17 times with only 5% area increase can be achieved if a proper selection of the (2n± k} moduli is performed. Pedro Miguens Matutino, Ricardo Chaves, Leonel Sousa |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Secure partial dynamic reconfiguration with unsecured external memoryabstractThis paper proposes a solution to improve the security of the partial dynamic reconfiguration of FPGA, without significantly affecting the reconfiguration performance. The existing solutions for secure partial dynamic reconfiguration on SRAM based FPGAs impact the reconfiguration process and the available resources due to their complex multi-layered partial bitstream validation process. This adversely affects the performance of applications using reconfigurable hardware. The proposed solution uses high performance encryption engines to change the encryption key of the remotely received bitstream by a randomly generated key, unique to each configuration, when storing the bitstream in the external unsecured memory. An additional CBC-MAC authentication mechanism is also considered that combined with the frame-wise error detection mechanism of the configuration port, allows for an improved countermeasure against replay attack and wrongful bitstream usage. The proposed solution introduces a resource overhead of 1.1% in regard to the base reconfigurable system and provides the lowest impact on the reconfiguration process when compared to the related state of the art, achieving a reconfiguration throughput of 2.5 Gbps. Hirak J. Kashyap, Ricardo Chaves |
FPL | 2 |
| 2014 | Method for designing multi-channel RNS architectures to prevent power analysis SCAabstractPower analysis attacks are one of the most common Side-Channel Attacks (SCAs), proven to be extremely successful even on protected embedded devices. This paper proposes the use of a Residue Number System (RNS) architecture with randomly permuted moduli sets to implement the Double-and-Add computation, which is proven as the most susceptible operation in Elliptic Curve Cryptography (ECC). The proposed solution randomly permutes the moduli sets, allowing randomized power traces, significantly removing the correlation between the power dissipation and the secret key and eliminating the need for the intermediate conversion to binary required in the state-of-the-art. Architectures obtained for a 90nm standard cell technology suggest that a significant power analysis resistance is achieved for the Double-and-Add circuitry, incurring an extra performance cost of 3 times compared to the related state-of-the-art. Héctor Pettenghi, Jude Angelo Ambrose, Ricardo Chaves, Leonel Sousa |
ISCAS | 3 |
| 2013 | A compact and scalable RNS architectureabstractThis paper proposes a unified architecture for designing Residue Number System (RNS) based processors for moduli sets with an arbitrary number of channels. Recently, new RNS moduli sets have been proposed in order to increase the dynamic range and reduce the width of the channels. The proposed architecture allows designing forward and reverse RNS converters, as well as the arithmetic operators of each modulo channel. The forward and reverse conversions are implemented using channel arithmetic units, resulting in a very compact architecture. Moreover, the arithmetic operations supported at the channel level include addition, subtraction, and multiplication with accumulation capability. The presented results suggest that the proposed RNS architecture leads to compact and scalable implementations, with competitive, or even better, performance when compared with the related state of the art, considering fixed moduli sets. Experimental results suggest gains of 17% in the delay of arithmetic operations, with an area reduction of 23% regarding the state of the art. Pedro Miguens Matutino, Ricardo Chaves, Leonel Sousa |
ASAP | 2 |
| 2013 | Scalable and high throughput biosensing platformabstractA novel multi-channel high performance embedded system capable of high throughput biological analysis is proposed in this paper. Despite other integrated lab-on-chip solutions based on magnetoresistive biochips have already been developed, they lack the scalability and computational resources to cope with new biochip designs featuring more than 1000 sensors. A new configurable acquisition and processing architecture is proposed, combining dedicated coprocessors to perform signal filtering and other computational demanding tasks, with a central processor controlling the whole system. The mapping of the architecture into a Zynq SoC demonstrated its ability to support 8 times more sensors, while ensuring a sampling frequency 1000+ times higher than the previous platforms. Furthermore, the Zynq reconfiguration abilities provide a mechanism to adapt the processing and maximize the biological sensitivity. José M. Leitão, José A. Germano, Nuno Roma, Ricardo Chaves, Pedro Tomás |
FPL | 4 |
| 2013 | HotStream: Efficient Data Streaming of Complex Patterns to Multiple Accelerating KernelsabstractDesigning accelerating kernels is a comprehensive task that requires efficient coupling of hardware and software. In particular, the structures responsible for handling data transfers in multi-core accelerator-based systems play a crucial role in the resulting performance. This paper proposes a data streaming accelerator framework that provides efficient data management facilities that are easily tailored for any application and data pattern. This is achieved through an innovative and fully programmable data management structure, implemented with two granularity levels. The obtained results show that the proposed framework is capable of efficient address generation and data fetch for complex streaming data patterns, while significantly reducing the size occupied by the pattern description. A large matrices multiplication case-study, based on a streaming architecture with four sub-block multiplication cores, demonstrates that, by enabling data re-use, the proposed framework increases the available bandwidth by 4.2x, resulting in a performance speedup of 2.1x. Furthermore, it reduces the Host memory requirements and its intervention by more than 40x. Sergio Paiagua, Frederico Pratas, Pedro Tomás, Nuno Roma, Ricardo Chaves |
SBAC-PAD | 5 |
| 2013 | On the Design of RNS Reverse Converters for the Four-Moduli Set ${\bf\{2^{\mmb n}+1, 2^{\mmb n}-1, 2^{\mmb n}, 2^{{\mmb n}+1}+1\}}$abstractIn this brief, we propose a method to design efficient adder-based converters for the four-moduli set {2n+1, 2n-1, 2n, 2n+1+1} with n odd, which provides a dynamic range of 4n+1 bits for the residue number system (RNS). This method hierarchically applies the mixed radix approach to balanced pairs of residues in two levels. With the proposed method, only simple binary and modulo 2k-1 additions are required, fully avoiding the usage of modulo 2k+1 arithmetic operations, which is a significant advantage over the currently available RNS reverse converters for this type of moduli set. Experimental results show that the delay of the proposed converters is significantly reduced when compared with the related state of the art; for example, for a 65-nm CMOS ASIC technology and a dynamic range of 21 bits, the conversion time and the circuit area are reduced by about 44% and 30%, respectively, while the conversion time is reduced by 34% for a dynamic range of 37 bits with the circuit area increasing only by 25%. Moreover, the proposed reverse converters outperform the related state of the art for any value of n by up to 70%, according to the figure-of-merit energy per conversion. Leonel Sousa, Samuel Antão, Ricardo Chaves |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | RNS Arithmetic Units for Modulo {2^n+-k}abstractRecently new Residue Number Systems (RNS) moduli sets have been proposed in order to increase the dynamic range and reduce the width of channels, therefore, reducing the processing time and further exploiting the carry-free characteristic of the modular arithmetic. In this paper we propose improved units for addition, subtraction, and multiplication in RNS for modulo {2n±k}. With this work, the somewhat disregarded field of RNS unit design is covered, encouraging the development of moduli sets with channels other than the traditional {2n}, {2n±1} modulo. In order to evaluate the performance of the proposed structures, they are compared with the well known units for modulo {2n}, {2n±1}, and {2n±3}. These structures allow to implement generic units for modulo {2n±k}, in the case of modular multiplication it is achieved the same critical path delay and merely 4% of increase on area resources, when compared with the dedicated structure presented in the state-of-art for modulo {2n±3}. Pedro Miguens Matutino, Héctor Pettenghi, Ricardo Chaves, Leonel Sousa |
DSD | 3 |
| 2011 | Binary-to-RNS Conversion Units for moduli {2^n ± 3}abstractIn this paper Residue Number Systems (RNS) conversion structures from Binary to RNS modulo {2n± 3} are proposed. These structures are based on arithmetic calculations without the need for Lookup Tables as in the related art. Additionally, the required 4:2 and 3:2 Carry-Save Adders (CSA) modulo {2n± 3} are also proposed. Experimental results obtained for an ASIC technology suggest that the presented CSAs, needed in the conversion, improve the related art by reducing the required area resources by 33% and achieving a 1.49× speedup. Experimental results for the proposed conversion units suggest that improvements in performance up to 3 times can be achieved, while reducing the required area resources by 85%. Pedro Miguens Matutino, Ricardo Chaves, Leonel Sousa |
DSD | 2 |
| 2011 | Compact CLEFIA Implementation on FPGASabstractIn this paper two compact hardware structures for the computation of the CLEFIA encryption algorithm are presented. One structure based on the existing state of the art and a novel structure with a more compact organization. This paper shows that, with the use of the existing embedded FPGA components and a careful scheduling, throughputs above 1Gbit/s can be achieved with a resource usage as low as 86 LUTs and 3 BRAMs on a VIRTEX 5 FPGA. Implementation results suggest that a LUT reduction up to 67% can be achieved at a performance cost of 17% on a VIRTEX 4 FPGA, resulting in Throughput/Slice efficiency gains up to 2.5 times, when compared with the related state of the art. Paulo Proenca, Ricardo Chaves |
FPL | 2 |
| 2010 | Arithmetic Units for RNS Moduli {2n-3} and {2n+3} OperationsabstractA new moduli set {2n- 1, 2n+ 3, 2n+ 1, 2n- 3} has recently been proposed to represent numbers in Residue Number Systems (RNS), increasing the number of channels. With this, the processing time can be reduced by simultaneously exploiting the carry-free characteristic of the modular arithmetic and improving the parallelism. In this paper, hardware structures for addition and multiplication operation in RNS for the moduli {2n- 3} and {2n+ 3} are proposed and analyzed. In order to evaluate the performance of the proposed units they were implemented on an ASIC technology. The obtained experimental results suggest that the performance of the moduli {2n± 3} are acceptable but demand more area resource and impose a larger delay than the typically used {2n± 1} arithmetic units. Addition units require at least 42% more area for a performance identical to the {2n+ 1} modulo adder. The multiplication units require up to 37% more area and impose a delay 25% higher. This paper also suggests that more balanced moduli sets should be developed in order to achieve more efficient RNS. Pedro Miguens Matutino, Ricardo Chaves, Leonel Sousa |
DSD | 2 |
| 2010 | An improved RNS reverse converter for the {22n+1-1, 2n, 2n-1} moduli setabstractIn this paper, we propose a novel high speed memoryless reverse converter for the moduli set {22n+1-1, 2n, 2n-1}. First, we simplify the traditional Chinese Remainder Theorem in order to obtain a reverse converter that only requires arithmetic mod-(22n+l-1). Second, we further improve the resulting architecture to obtain a purely adder based reverse converter. The proposed converter has a critical path delay of (7n + 7) Full Adders (FA) while the best state of the art converter for this moduli set requires (10n + 5) FA on the critical path. To validate these results, the converters are implemented in a Standard Cell 0.18-μm CMOS technology and the results assert that, on average, the proposed converter achieves about 19% delay reduction at the expense of less than 3% area increase. Kazeem Alagbe Gbolagade, Ricardo Chaves, Leonel Sousa, Sorin Cotofana |
ISCAS | 2 |
| 2010 | An improved RNS generator 2n +/- k based on threshold logicabstractThis paper presents a new scheme for designing residue generators using threshold logic. This approach is based on the periodicity of the series of powers of 2 taken modulo 2n± k. In addition, a new algorithm is proposed to obtain a new set of partitions which are more advantageous in terms of area and delay for the presented topology. Experimental results in the analized range of k and n show that new proposed circuits using the novel partitioning are 70% faster and provide area savings of 64%, when compared with similar circuits using the partitioning methods presented to date. Héctor Pettenghi, Ricardo Chaves, Leonel Sousa, Maria J. Avedillo |
VLSI-SoC | 2 |
| 2009 | Compact and Flexible Microcoded Elliptic Curve Processor for Reconfigurable DevicesabstractThis paper presents a very compact and flexible processor to support Elliptic Curve (EC) cryptosystems based on GF(2^m) finite fields. This processor can be customized with a two-level microinstruction hierarchy that allows for customization of both field operations and EC algorithms. It was specially designed to benefit from reconfiguration capabilities to scale arithmetic units for different sizes and to replicate processing units to enhance performance. The flexibility resulting from these characteristics was not found in the related art. The proposed processor was implemented and thoroughly tested in a Xilinx Virtex XC4VSX35, supporting a real EC algorithm for point multiplication for a GF(2^163) field, requiring 1.35ms, and using up to 15 times less area than related implementations. Samuel Antão, Ricardo Chaves, Leonel Sousa |
FCCM | 2 |
| 2008 | Merged Computation for Whirlpool HashingabstractThis paper presents an improved hardware structure for the computation of the Whirlpool hash function. By merging the round key computation with the data compression and by using embedded memories to perform part of the Galois Field (2s) multiplication, a core can be implemented in just 43% of the area of the best current related art while achieving a 12% higher throughput. The proposed core improves the Throughput per Slice compared to the state of the art by 160%, achieving a throughput of 5.47 Gbit/s with 2110 slices and 32 BRAMs on a VIRTEX II Pro FPGA. Results for a real application are also presented by considering a polymorphic computational approach. Ricardo Chaves, Georgi Kuzmanov, Leonel Sousa, Stamatis Vassiliadis |
DATE | 1 |
| 2008 | On-the-fly attestation of reconfigurable hardwareabstractThis paper presents a novel method to perform on-the-fly attestation of hardware structures loaded to reconfigurable devices. Given that a loadable hardware structure to a reconfigurable device is described by a binary bitstream, the hash value of this bitstream can be calculated to validate the hardware structure. To optimize this attestation, the hash value computation is implemented in hardware on the FPGA itself. To guarantee the integrity of the existing computation architecture, the proposed hardware module also enforces region delimitation. With the region delimitation, only the regions intended to be reconfigured can be modified. Implementation results suggest that this bitstream attestation can be performed without imposing an extra delay to the reconfigurable process and at an area cost of less that 10% of a Virtex II Pro 30 FPGA device. Ricardo Chaves, Georgi Kuzmanov, Leonel Sousa |
FPL | 1 |
| 2008 | Efficient FPGA elliptic curve cryptographic processor over GF(2m)abstractIn this paper a processor that supports elliptic curve cryptographic applications over GF (2m) is proposed. The proposed structure is capable of calculating point multiplication and addition using a single coordinate to contain the point information. This compression allows for a better usage of the bandwidth resources. For the point multiplication procedure, all coordinate pre-calculations are completely avoided. This design was successful prototyped on a reconfigurable device for the field GF (2163). Experimental results suggest that point multiplication can be performed in 144 mus and point affine addition in 1.02 mus. Comparing with the related work, a 5 times speedup is obtained for point addition and multiplication. The presented design offers a well balanced area-time performance when compared with existent elliptic curve point multiplication specific processors. Samuel Antão, Ricardo Chaves, Leonel Sousa |
FPT | 2 |
| 2008 | BRAM-LUT Tradeoff on a Polymorphic DES Design
Ricardo Chaves, Blagomir Donchev, Georgi Kuzmanov, Leonel Sousa, Stamatis Vassiliadis |
HiPEAC | 1 |
| 2008 | Cost-Efficient SHA Hardware AcceleratorsabstractThis paper presents a new set of techniques for hardware implementations of secure hash algorithm (SHA) hash functions. These techniques consist mostly in operation rescheduling and hardware reutilization, therefore, significantly decreasing the critical path and required area. Throughputs from 1.3 Gbit/s to 1.8 Gbit/s were obtained for the SHA implementations on a Xilinx VIRTEX II Pro. Compared to commercial cores and previously published research, these figures correspond to an improvement in throughput/slice in the range of 29% to 59% for SHA-1 and 54% to 100% for SHA-2. Experimental results on hybrid hardware/software implementations of the SHA cores, have shown speedups up to 150 times for the proposed cores, compared to pure software implementations. Ricardo Chaves, Georgi Kuzmanov, Leonel Sousa, Stamatis Vassiliadis |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Improving SHA-2 Hardware Implementations
Ricardo Chaves, Georgi Kuzmanov, Leonel Sousa, Stamatis Vassiliadis |
CHES | 1 |
| 2006 | Reconfigurable memory based AES co-processorabstractWe consider the AES encryption/decryption algorithm and propose a memory based hardware design to support it. The proposed implementation is mapped on the Xilinx Virtex II Pro technology. Both the byte substitution and the polynomial multiplication of the AES algorithm are implemented in a single dual port on-chip memory block (BRAM). Two AES encryption/decryption cores have been designed and implemented on a prototyping XC2VP20-7 FPGA: a completely unrolled loop structure capable of achieving a throughput above 34 Gbits/s, with an implementation cost of 3513 slices and 80 BRAMs; and a fully folded structure, requiring only 515 slices and 12 BRAMs, capable of a throughput above 2 Gbits/s. To evaluate the proposed AES design, it has been embedded in a polymorphic processor organization, as a reconfigurable co-processor. Comparisons to state-of-the-art AES cores indicate that the proposed unfolded core outperforms the most recent works by 34% in throughput and requires 68% less reconfigurable area. Experimental results of both folded and unfolded AES cores suggest over 560% improvement in the throughput/slice metric when compared to the recent AES related art Ricardo Chaves, Georgi Kuzmanov, Stamatis Vassiliadis, Leonel Sousa |
IPDPS | 1 |
| 2004 | {2n+1, sn+k, sn-1}: A New RNS Moduli Set ExtensionabstractThe increasing usage of residual number system (RNS) in signal processing applications demands the development of new and more adaptable RNS moduli sets and arithmetic units. This paper presents a new adaptable moduli set extension for the traditional moduli set {2/sup n/ + 1, 2/sup n/, 2/sup n/ - 1}. As it will be shown, this new moduli set extension ({2/sup n/ + 1, 2/sup n+k/, 2/sup n/ - 1}) allows the balancing of the binary channel (2/sup n+k/) in relation to the other two channels. Moreover, it does not require the development of new addition and multiplication units, since it is possible to reuse the already developed and well studied units for these moduli operations. Ricardo Chaves, Leonel Sousa |
DSD | 1 |
| 2003 | RDSP: A RISC DSP based on Residue Number SystemabstractThis paper is focused on low power programmable fast digital signal processors (DSP) design based on a configurable 5-stage RISC core architecture and on residue number systems (RNS). Several innovative aspects are introduced at the control and datapath architecture levels, which support both the binary system and the RNS. A new moduli set {2/sup n/-1, 2/sup 2n/, 2/sup n/+1} is also proposed for balancing the processing time in the different RNS channels. Experimental results, obtained trough RDSP implementation on FPGA and ASIC, show that not only a significant reduction in circuit area and power consumption but also a speedup may be achieved with RNS when compared with a binary DSP. Ricardo Chaves, Leonel Sousa |
DSD | 1 |
| 2003 | Towards tangibility in gameplay: building a tangible affective interface for a computer gameabstractIn this paper we describe a way of controlling the emotional states of a synthetic character in a game (FantasyA) through a tangible interface named SenToy. SenToy is a doll with sensors in the arms, legs and body, allowing the user to influence the emotions of her character in the game. The user performs gestures and movements with SenToy, which are picked up by the sensors and interpreted according to a scheme found through an initial Wizard of Oz study. Different gestures are used to express each of the following emotions: anger, fear, happiness, surprise, sadness and gloating. Depending upon the expressed emotion, the synthetic character in FantasyA will, in turn, perform different actions. The evaluation of SenToy acting as the interface to the computer game FantasyA has shown that users were able to express most of the desired emotions to influence the synthetic characters, and that overall, players, especially children, really liked the doll as an interface. Ana Paiva 0001, Rui Prada, Ricardo Chaves, Marco Vala, Adrian Bullock, Gerd Andersson, Kristina Höök |
ICMI | 3 |
| 2003 | Demo: playingfFantasyA with senToyabstractGame development is an emerging area of development for new types of interaction between computers and humans. New forms of communication are now being explored there, influenced not only by face to face communication but also by recent developments in multi-modal communication and tangible interfaces. This demo will feature a computer game, FantasyA, where users can play the game by interacting with a tangible interface, SenToy (see Figure 1). The main idea is to involve objects and artifacts from real life into ways to interact with systems, and in particular with games. So, SenToy is an interface for users to project some of their emotional gestures through moving the doll in certain ways. This device would establish a link between the users (holding the physical device) and a controlled avatar (embodied by that physical device) of the computer game, FantasyA. Ana Paiva 0001, Rui Prada, Ricardo Chaves, Marco Vala, Adrian Bullock, Gerd Andersson, Kristina Höök |
ICMI | 3 |
| 2003 | SenToy: an affective sympathetic interface
Ana Paiva 0001, Marco Costa 0002, Ricardo Chaves, Moisés Simões Piedade, Dário Mourão, Daniel Sobral, Kristina Höök, Gerd Andersson, Adrian Bullock |
Int. J. Hum. Comput. Stud. | 3 |