EDBT 2026 Demo / reviewers in the wild / expert
José Luis Imaña
dblp:92/6884
· DBLP profile ↗
18ranked-venue papers
14as first author
7since 2021 · last 2024
0000-0002-4220-4111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 12 first-author · 4 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Integrating Post-Quantum Cryptography Plugins for IPsec Offloads to Data Processing Units in the Cloud-Edge ContinuumabstractThe imminent advent of Quantum Computers poses a significant threat to the cryptographic algorithms supporting the public key infrastructure (PKI) of widely used communication protocols. High Performance Computing (HPC) data centers among other interested parties are well aware of the catastrophic consequences quantum attacks could have on their PKI and are consequently transitioning to Post-Quantum Cryptographic (PQC) methods, despite the substantial overhead this introduces for handling incoming network packets. This work addresses the transition to PQC within the context of the Cloud-Edge Continuum by integrating the Open Quantum Safe (OQS) library into the accelerated strongSwan developed by Mellanox for Data Processing Units (DPUs). This integration offloads cryptographic operations from central servers to data DPUs distributed across the cloud-edge continuum. Our solution ensures quantum security by providing PQ authentication through CRYSTALS-Dilithium or CRYSTALS-FALCON, PQ key exchanges via CRYSTALS-Kyber, and confidential data transmission using AES-256. Additionally, the deployment of this implementation on DPUs helps reduce the computational load on both HPC data centers and edge devices, promoting more efficient and secure operations across the entire cloud-edge continuum. Abraham Cano Aguilera, Carlos Rubio Garcia, Raphael Frantz, Idelfonso Tafur Monroy, José Luis Imaña, Juan Jose Vegas Olmos |
ICNP | 5 |
| 2024 | Low-Complexity Hardware Architecture of APN Permutations Using TU-DecompositionabstractFunctions with good cryptographic properties which are used as S-boxes in the design of block ciphers have a fundamental importance to the security of these ciphers since they determine the resistance to various kinds of cryptanalytic attacks. Almost Perfect Nonlinear (APN) functions provide the best possible resistance to differential cryptanalysis, which is one of the most efficient cryptographic attacks against block ciphers known to date. Furthermore, APN permutations are of particular interest in practice since many cipher designs require the S-box to be a permutation. In this paper, we present a low-complexity hardware architecture for the TU-decomposition of APN permutations, showing how Dillon’s APN permutation can be decomposed in this way as a practically relevant example. The TU-decomposition of an m-bit permutation is based on the use of two$m/2$-bit keyed permutations (T and U) to reduce the complexity of the original permutation. Dillon’s permutation on 6 bits is the only known APN permutation on an even number of bits, so its study is of fundamental interest. We present hardware theoretical complexities and experimental results obtained from FPGA and ASIC implementations for the proposed TU-decomposition hardware architecture. These complexities and results are compared with other hardware architectures given in the literature for the same function. From the comparisons, it can be observed that the TU-decomposition architecture presented here greatly outperforms other hardware approaches with respect to area, delay and area$\times $delay complexities. Lilya Budaghyan, José Luis Imaña, Nikolay S. Kaleyski |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Domain-oriented masked bit-parallel finite-field multiplier against side-channel attacksabstractSide-Channel Analysis (SCA) constitutes a serious threat to the security of implemented cryptosystems. In SCA, the attacker can obtain information leakage from a device executing cryptographic algorithms by means of the measure of side-channels such as power consumption, electromagnetic radiation and execution time. For this reason, effective countermeasures against SCA are indispensable in implemented cryptographic devices. The use of masking schemes (in which intermediate computations are independent from the sensible input data) constitutes the most effective approach to achieve resistance against physical attacks. Among the different masking methods proposed for hardware, domain-oriented masking is one of the most promising due to its lower implementation costs, level of security and glitch resistance. In this paper, a new bit-parallel first-order domain-oriented masked finite field multiplier is presented which incorporates the addition of fresh random values without increasing the computation delay. Explicit expressions for the computation of the new masked multiplier for the binary extension field used in the Advanced Encryption Standard (AES) are also given. José Luis Imaña, Siemen Dhooghe |
Inf. Process. Lett. | 1 |
| 2022 | Work-in-Progress: High-Performance Systolic Hardware Accelerator for RBLWE-based Post-Quantum CryptographyabstractRing-Binary-Learning-with-Errors (RBLWE)-based post-quantum cryptography (PQC) is a promising scheme suitable for lightweight applications. This paper presents an efficient hardware systolic accelerator for RBLWE-based PQC, targeting high-performance applications. We have briefly given the algorithmic background for the proposed design. Then, we have transferred the proposed algorithmic operation into a new systolic accelerator. Lastly, field-programmable gate array (FPGA) implementation results have confirmed the efficiency of the proposed accelerator. Tianyou Bao, José Luis Imaña, Pengzhou He, Jiafeng Xie |
CODES+ISSS | 2 |
| 2022 | Decomposition of Dillon's APN Permutation with Efficient Hardware Implementation
José Luis Imaña, Lilya Budaghyan, Nikolay S. Kaleyski |
WAIFI | 1 |
| 2022 | Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum CryptographyabstractRing learning-with-errors(RLWE)-based encryption scheme is a lattice-based cryptographic algorithm that constitutes one of the most promising candidates for Post-Quantum Cryptography (PQC) standardization due to its efficient implementation and low computational complexity.Binary Ring-LWE (BRLWE) is a new optimized variant of RLWE, which achieves smaller computational complexity and higher efficient hardware implementations. In this paper, two efficient architectures based onLinear-Feedback Shift Register(LFSR) for the arithmetic used inInverted Binary Ring-LWE (InvBRLWE)-based encryption scheme are presented, namely the operation of$A\cdot B+C$over the polynomial ring$\mathbb {Z}_{q}/(x^{n}+1)$. The first architecture optimizes the resource usage for major computation and has a novel input processing setup to speed up the overall processing latency with minimized input loading cycles. The second architecture deploys an innovative serial-in serial-out processing format to reduce the involved area usage further yet maintains a regular input loading time-complexity. Experimental results show that the architectures presented here improve the complexities obtained by competing schemes found in the literature, e.g., involving 71.23% less area-delay product than recent designs. Both architectures are highly efficient in terms of area-time complexities and can be extended for deploying in different lightweight application environments. José Luis Imaña, Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | LFSR-Based Bit-Serial GF(2m) Multipliers Using Irreducible TrinomialsabstractIn this article, a new architecture of bit-serial polynomial basis (PB) multipliers over the binary extension field GF(2m) generated by irreducible trinomials is presented. Bit-serial GF(2m) PB multiplication offers a performance/ area trade-off that is very useful in resource constrained applications. The architecture here proposed is based on LFSR (Linear-Feedback Shift Register) and can perform a multiplication in m clock cycles with a constant propagation delay of TA þ TX. These values match the best time results found in the literature for bit-serial PB multipliers with a slight reduction of the space complexity. Furthermore, the proposed architecture can perform the multiplication of two operands fort different finite fields GF(2m) generated by t irreducible trinomials simultaneously in m clock cycles with the inclusion of t(m - 1Þ flipflops and tm XOR gates. José Luis Imaña |
IEEE Trans. Computers | 1 |
| 2020 | FPGA Implementation of Post-Quantum DME CryptosystemabstractThe rapid development of quantum computing constitutes a significant threat to modern Public-Key Cryptography (PKC). The use of Shor's algorithm with potential powerful quantum computers could easily break the two most widely used public key cryptosystems, namely, RSA and Elliptic Curve Cryptography (ECC), based on integer factorization and discrete logarithm problems. For this reason, Post-Quantum Cryptography (PQC) based on alternative mathematical features has become a fundamental research topic due to its resistance against quantum computers. The National Institute of Standards and Technology (NIST) has even opened a call for proposals of quantum-resistant PKC algorithms in order to standardize one or more PQC algorithms. Cryptographic systems that appear to be extremely difficult to break with large quantum computers are hash -based cryptography, lattice -based cryptography, code -based cryptography, and multivariate -quadratic cryptography. Furthermore, efficient hardware implementations are highly required for these alternative quantum -resistant cryptosystems. José Luis Imaña, Ignacio Luengo |
FCCM | 1 |
| 2020 | High-throughput architecture for post-quantum DME cryptosystem
José Luis Imaña, Ignacio Luengo |
Integr. | 1 |
| 2018 | Reconfigurable implementation of GF(2m) bit-parallel multipliersabstractHardware implementations of arithmetic operations over binary finite fields GF(2m) are widely used in several important applications, such as cryptography, digital signal processing and error-control codes. In this paper, efficient reconfigurable implementations of bit-parallel canonical basis multipliers over binary fields generated by type II irreducible pentanomials f (y) = ym+ yn+2+ yn+1+ yn+ 1 are presented. These pentanomials are important because all five binary fields recommended by NIST for ECDSA can be constructed using such polynomials. In this work, a new approach for GF(2m) multiplication based on type II pentanomials is given and several post-place and route implementation results in Xilinx Artix-7 FPGA are reported. Experimental results show that the proposed multiplier implementations improve the area χ time parameter when compared with similar multipliers found in the literature. José Luis Imaña |
DATE | 1 |
| 2018 | Efficient FPGA Implementation of Binary Field Multipliers Based on Irreducible TrinomialsabstractBinary extension (or Galois) fields GF(2m) have been widely studied due to their use in several important applications, such as cryptography, error control codes and digital signal processing. These applications require efficient hardware implementations of GF(2m) arithmetic operations, particularly multiplication, which is considered the most important and complex one. The complexity of GF(2m) multiplication depends on the representation basis and on the defining irreducible polynomial f(y) selected for the finite field. For efficient hardware implementations, polynomial basis and irreducible trinomials or pentanomials are normally used. Any element A ϵ GF(2m) can be represented in the polynomial basis {1,x,...,xm-1} as A = Σi-0m-1,αixiwith αiϵ GF(2), where x is a root of the irreducible polynomial f(y) = Σi=0mfiyi. Polynomial basis multiplication C = A · B requires a polynomial multiplication followed by a reduction modulo an irreducible polynomial. Mastrovito [1] proposed an efficient bit-parallel polynomial basis multiplier in which a product matrix was introduced to combine the above two steps together. A new polynomial basis multiplication method applied to irreducible trinomials was proposed in [2], where the functions S_i and T_i given by the addition of terms xk= (akbk) and zij= (aibj+ ajbi), with ai, biϵ GF(2), were obtained from the decomposition of a product matrix. The addition of these functions is used for the computation of the product of two GF(2m) elements. In [3], the above method was applied to type II irreducible pentanomials and the functions S_i and T_i were split in the form Si= skiSki+ ... + s0iS0iand Ti= tkiTik+ ... + t0iT0i, with sji, tjiϵ GF(2) and k = log2m. The terms Sijand Tijrepresent the sum of 2jproducts a_kb_l and therefore can be implemented as a j-level complete binary tree of XOR gates. The addition in pairs of binary trees with the same depth leads to a reduction of the multiplication delay. However, splitting method imposes hard restrictions (given by the use of parenthesis in the expressions of the coordinates) for the addition of Sijand Tijterms in order to reduce the number of XOR levels. These restrictions could not be efficient for a synthesis tool in order to map that expressions into FPGA's logic blocks. If parenthesized restrictions are removed, more freedom could be given for the synthesizer to find an optimized implementation of the multiplier. In this work, efficient Xilinx FPGA implementations of GF(2m) bit-parallel polynomial basis multipliers for irreducible trinomials are presented. Based on [2], a new general algorithm for multiplication over irreducible trinomials f(y) = ym+ yn+1, with 1 ≤ n ≤ (m+1)/2, is proposed and the splitting method given in [3] is applied to these irreducible polynomials. Furthermore, in order to optimize the synthesis of the multipliers, a new approach for the computation of the product is used where the splitting of S_i and T_i terms is performed, but the restriction given by the addition in pairs of binary trees with the same depth has been removed. In this way, Xilinx tools are free to optimize the synthesis of the multiplier. Several GF(2m) multipliers for different binary fields have been described in VHDL and their post-place and route implementation results in Xilinx Artix-7 have been reported. Experimental results show that the multiplier here proposed exhibits the best delay and Area×Time complexities when it is compared with similar multipliers found in the literature. Moreover, the new approach also achieves the lowest number of slices in most of the implemented multipliers. José Luis Imaña |
FCCM | 1 |
| 2018 | Fast Bit-Parallel Binary Multipliers Based on Type-I PentanomialsabstractIn this paper, a fast implementation of bit-parallel polynomial basis (PB) multipliers over the binary extension field GF(2m) generated by type-I irreducible pentanomials is presented. Explicit expressions for the coordinates of the multipliers and a detailed example are given. Complexity analysis shows that the multipliers here presented have the lowest delay in comparison to similar bit-parallel PB multipliers found in the literature based on this class of irreducible pentanomials. In order to prove the theoretical complexities, hardware implementations over Xilinx FPGAs have also been performed. Experimental results show that the approach here presented exhibits the lowest delay with a balanced Area x Time complexity when it is compared with similar multipliers. José Luis Imaña |
IEEE Trans. Computers | 1 |
| 2013 | Low complexity bit-parallel polynomial basis multipliers over binary fields for special irreducible pentanomials
José Luis Imaña, Román Hermida, Francisco Tirado |
Integr. | 1 |
| 2010 | Efficient FPGA Modular Multiplication and Exponentiation Architectures Using Digit Serial ComputationabstractModular exponentiation with large modulus and exponent has been widely used in public key cryptosystems. Montgomery's modular multiplication algorithm is normally used since no trial division is necessary and the critical path is reduced by using carry-save addition (CSA). In this paper, the Montgomery multiplication is greatly optimized and architectures are proposed to perform the Least-Significant-Bit (LSB) first and the Most-Significant-Bit (MSB) first algorithms. The architecture here presented has the following distinctive characteristics: 1) Use of digit-serial approach for Montgomery multiplication. 2) Conversion of the CSA representation of intermediate multiplication using carry-skip addition which reduces the critical path with a small area-speed penalty. 3) Precompute quotient value in Montgomery iteration in order to speed up operation frequency. In this work, implementation results in Xilinx Virtex 5 and Virtex 2 are reported. Experimental results show that the proposed modular exponentiation and modular multiplication design obtains the best delay performance compared with previous published works and outperforms them in terms of area-time complexity. Gustavo Sutter 0001, Jean-Pierre Deschamps, José Luis Imaña |
FPL | 3 |
| 2006 | Bit-Parallel Finite Field Multipliers for Irreducible TrinomialsabstractA new formulation for the canonical basis multiplication in the finite fields GF(2/sup m/) based on the use of a triangular basis and on the decomposition of a product matrix is presented. From this algorithm, a new method for multiplication (named transpositional) applicable to general irreducible polynomials is deduced. The transpositional method is based on the computation of 1-cycles and 2-cycles given by a permutation defined by the coordinate of the product to be computed and by the cardinality of the field GF(2/sup m/). The obtained cycles define groups corresponding to subexpressions that can be shared among the different product coordinates. This new multiplication method is applied to five types of irreducible trinomials. These polynomials have been widely studied due to their low-complexity implementations. The theoretical complexity analysis of the corresponding bit-parallel multipliers shows that the space complexities of our multipliers match the best results known to date for similar canonical GF(2/sup m/) multipliers. The most important new result is the reduction, in two of the five studied trinomials, of the time complexity with respect to the best known results. José Luis Imaña, Juan Manuel Sánchez, Francisco Tirado |
IEEE Trans. Computers | 1 |
| 2006 | Low Complexity Bit-Parallel Multipliers Based on a Class of Irreducible PentanomialsabstractIn this paper, we consider the design of bit-parallel canonical basis multipliers over the finite field$GF(2^{m})$generated by a special type ofirreducible pentanomialthat is used as an irreducible polynomial in theAdvanced Encryption Standard(AES). Explicit formulas for the coordinates of the multiplier are given. The main advantage of our design is that some of the expressions obtained are common toanyirreducible polynomial, so our multiplier can be generalized to perform the multiplication overgeneral irreducible polynomials. Moreover, the obtained expressions can be easily converted to parameterizable code using hardware description languages. The theoretical complexity analysis also shows that our bit-parallel multipliers present a reduced number ofxorgates with respect to the best known results found in the literature. José Luis Imaña, Román Hermida, Francisco Tirado |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Reconfigurable implementation of bit-parallel multipliers over GF(2m) for two classes of finite fieldsabstractGalois fields GF(2/sup m/) are used in a wide number of applications such as cryptography, digital signal processing and error-control codes. The multiplication is considered the most important and one of the most complex GF(2/sup m/) operations, so efficient multiplier architectures are highly desired. A new construction method of bit-parallel multipliers over GF(2/sup m/) for two classes of finite fields is presented. Our approach determines groups of subexpressions that can be shared among the product coordinates. General expressions are given, and the theoretical complexity analysis proves that our multipliers reduce the best time complexities known to date. The multipliers have been implemented on Xilinx Virtex FPGAs. The experiments prove that our method reduces the area requirements of the multipliers with respect to other similar multipliers. José Luis Imaña |
FPT | 1 |
| 2003 | A New Reconfigurable-Oriented Method for Canonical Basis Multiplication over a Class of Finite Fields GF(2m)
José Luis Imaña, Juan Manuel Sánchez |
FPL | 1 |