Rei Ueno

dblp:147/7755 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0002-9754-6792ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 4 first-author · 9 since 2021Systems, architecture and hardware · 10 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Formal Security Proof of Masking - Reduction from Strong Noisy Leakage to Probing Model Without Random Probing and Application to LR Primitive
Rei Ueno, Akiko Inoue, Kazuhiko Minematsu, Akira Ito 0002, Naofumi Homma
CRYPTO (7)1
2024 Comparative Analysis and Implementation of Jump Address Masking for Preventing TEE Bypassing Fault Attacks
abstract
Attacks on embedded devices continue to evolve with the increasing number of applications in actual products. A trusted execution environment (TEE) enhances the security of embedded devices by isolating and protecting sensitive applications such as cryptography from malicious or vulnerable applications. However, the emergence of TEE bypass attacks using faults exposes TEEs to threats. In CHES’22, jump address masking (JAM) was proposed as a countermeasure against TEE bypass attacks, specifically targeting RISC-V. JAM prevents modifications of protected data by calculating jump addresses using the protected data, and is expected to provide promising resistance to TEE bypass attacks, for which traditional countermeasures are ineffective. However, JAM was originally proposed for bare metal applications. Therefore, its application to TEEs that operate with an OS presents technical and security challenges. This study proposes a method for applying JAM to Keystone, a major TEE framework for RISC-V, and validates its practical effectiveness and performance through a comparative evaluation with existing countermeasures such as memory encryption, random delays, and instruction duplication. Our evaluation reveals that the proposed JAM implementation is the first countermeasure that achieves complete resistance to TEE bypass attacks with an execution time overhead of approximately 340% for context switches and 1.0% across the entire program, which is acceptable compared with other countermeasures.
Shoei Nashimoto, Rei Ueno, Naofumi Homma
ARES2
2024 Crystalor: Recoverable Memory Encryption Mechanism with Optimized Metadata Structure
Rei Ueno, Hiromichi Haneda, Naofumi Homma, Akiko Inoue, Kazuhiko Minematsu
CCS1
2023 SCARF - A Low-Latency Block Cipher for Secure Cache-Randomization
Federico Canale, Tim Güneysu, Gregor Leander, Jan Philipp Thoma, Yosuke Todo, Rei Ueno
USENIX Security Symposium6
2022 On the Success Rate of Side-Channel Attacks on Masked Implementations: Information-Theoretical Bounds and Their Practical Usage
abstract
This study derives information-theoretical bounds of the success rate (SR) of side-channel attacks on masked implementations. We first develop a communication channel model representing side-channel attacks on masked implementations. We then derive two SR bounds based on the conditional probability distribution and mutual information of shares. The basic idea is to evaluate the upper-bound of the mutual information between the non-masked secret value and the side-channel trace by the conditional probability distribution of shares given its leakage, with a help of the Walsh-Hadamard transform. With the derived theorems, we also prove the security of masking schemes: the SR decreases exponentially with an increase in the number of masking shares, under a much more relaxed condition than the previous proof. To validate and utilize our theorems in practice, we propose a deep-learning-based profiling method for approximating the conditional probability distribution of shares to estimate the SR bound and the number of traces required for attacking a given device. We experimentally confirm that our bounds are much stronger than the conventional bounds on masked implementations, which validates the relevance of our theorems to practice.
Akira Ito 0002, Rei Ueno, Naofumi Homma
CCS2
2022 Efficient Modular Polynomial Multiplier for NTT Accelerator of Crystals-Kyber
abstract
This paper presents a hardware design that efficiently performs the number theoretic transform (NTT) for lattice-based cryptography. First, we propose an efficient modular multiplication method for lattice-based cryptography defined over Proth numbers. The proposed method is based on a K-RED technique specific to Proth numbers. In particular, we divide the intermediate result into the sign bit and the other absolute value bits and handle them separately to significantly reduce implementation costs. Then, we show a butterfly unit datapath of NTT and inverse INTT equipped with the proposed modular multiplier. We apply the proposed NTT accelerator to Crystals-Kyber, which is lattice-based cryptography, and evaluate its performance on Xilinx Artix-7. The results show that the proposed NTT accelerators achieve up-to 3% and 33% higher area-time efficiency in terms of LUTs and FFs, respectively, than conventional best methods. In addition, the low-latency version of the proposed NTT accelerators achieves a 18% lower-latency with an area-time efficiency (in terms of LUTs, FFs, and DSPs) than the existing fastest method.
Yuma Itabashi, Rei Ueno, Naofumi Homma
DSD2
2022 High-Speed Hardware Architecture for Post-Quantum Diffie-Hellman Key Exchange Based on Residue Number System
abstract
This paper presents a hardware architecture for a post-quantum key exchange protocol, named super-singular isogeny Diffie-Hellman (SIDH). The proposed hardware employs residue number system (RNS) and is optimized to reduce the latency of $\mathbb{F}_{p^{2}}$ multiplication and RNS Montgomery reduction, which are major time-consuming procedures in SIDH. The performance of the proposed hardware is validated and evaluated through an experimental implementation on Xilinx Kintex7 Ultrascale+. As a result, we confirm that the proposed hardware can perform an SIDH computation 34% faster than the state-of-the-art existing one on the same device at a resource overhead.
Rei Ueno, Naofumi Homma
ISCAS1
2022 Efficient Formal Verification of Galois-Field Arithmetic Circuits Using ZDD Representation of Boolean Polynomials
abstract
In this study, we present a new formal method for verifying the functionality of Galois-field (GF) arithmetic circuits. Assuming that the input–output relation (i.e., the specification of a GF arithmetic circuit) can be represented as polynomials over 2, the proposed method formally checks the equivalence between GF polynomials derived from a netlist and the specification. To efficiently verify the equivalence, we employ a zero-suppressed binary decision diagram (ZDD) to represent polynomials over 2. Even though polynomial reduction is the most time-consuming process of verification (i.e., equivalence checking), our new algorithm can efficiently reduce the GF polynomials in the form of a zero-suppressed binary decision diagram derived from the target netlist. The proposed algorithm derives the polynomials representing all intermediate nodes (i.e., the outputs of all gates) in the order from primary inputs to those primary outputs that are in accordance with the reverse topological traversal order. We demonstrated the efficiency and effectiveness of the proposed method via a set of experimental verifications. In particular, we confirmed that the proposed method can verify practical GF multipliers (including those used in standardized elliptic curve cryptography) approximately 30 times faster on average and at most 170 times faster than the best conventional method.
Akira Ito 0002, Rei Ueno, Naofumi Homma
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 ELM: A Low-Latency and Scalable Memory Encryption Scheme
abstract
Memory encryption (ME) with authentication is becoming a key security feature of modern processors, as evident by the adoption of ME by Intel’s SGX. Recently ME is actively studied from the viewpoint of system architecture. This paper studies ME from the viewpoint of symmetric-key cryptographic designs, with a primal focus on latency. A significant progress in such a direction can be observed in the SGX Integrity Tree (SIT). Using a variant of AES-GCM, SIT achieves an excellent latency. However, it has a scalability issue. By carefully examining SIT, we develop a new ME scheme dubbed ELM. We present an AES-based instantiation of ELM, and show that ELM significantly reduces latency from SIT for large memories, and achieves the provable security and equivalent hardware-protected (on-chip) area. We also present preliminary hardware implementations to substantiate our advantages.
Akiko Inoue, Kazuhiko Minematsu, Maya Oda, Rei Ueno, Naofumi Homma
IEEE Trans. Inf. Forensics Secur.4
2021 Fault-Injection Attacks Against NIST's Post-Quantum Cryptography Round 3 KEM Candidates
Keita Xagawa, Akira Ito 0002, Rei Ueno, Junko Takahashi, Naofumi Homma
ASIACRYPT (2)3
2021 Imbalanced Data Problems in Deep Learning-Based Side-Channel Attacks: Analysis and Solution
abstract
In recent years, the threat of profiling attacks using deep learning has emerged. Successful attacks have been demonstrated against various types of cryptographic modules. However, the application of deep learning to side-channel attacks (SCAs) is often not adequately assessed because the labels that are widely used in SCAs, such as the Hamming weight (HW) and Hamming distance (HD), follow an imbalanced distribution. This study analyzes and solves the problems caused by dataset imbalance during training and inference. First, we state the reasons for the negative effect of data imbalance in classification for deep-learning-based SCAs and introduce the Kullback-Leibler (KL) divergence as a metric to measure this effect. Using the KL divergence, we demonstrate through analysis how the recently reported cross-entropy ratio loss function can solve the problem of imbalanced data. We further propose a method to solve dataset imbalance at the inference phase, which utilizes a likelihood function based on the key value instead of the HW/HD. The proposed method can be easily applied in deep-learning-based SCAs because it only needs an extra multiplication of the inverted binomial coefficients and inference results (i.e., the output probabilities) from the conventionally trained model. The proposed solution corresponds to data-augmentation techniques at the training phase, and furthermore, it better estimates the keys because the probability distributions of the training and test data are preserved. We demonstrate the validity of our analysis and the effectiveness of our solution through extensive experiments on two public databases.
Akira Ito 0002, Kotaro Saito, Rei Ueno, Naofumi Homma
IEEE Trans. Inf. Forensics Secur.3
2021 Diffusional Side-Channel Leakage From Unrolled Lightweight Block Ciphers: A Case Study of Power Analysis on PRINCE
abstract
This study investigates a new side-channel leakage observed in the inner rounds of an unrolled hardware implementation of block ciphers in a chosen-input attack scenario. The side-channel leakage occurs in the first round and it can be observed in the later inner rounds because it arises from path activation bias caused by the difference between two consecutive inputs. Therefore, a new attack that exploits the leakage is possible even for unrolled implementations equipped with countermeasures (masking and/or deglitchers that separate the circuit in terms of glitch propagation) in the round involving the leakage. We validate the existence of such a unique side-channel leakage through a set of experiments with a fully unrolled PRINCE cipher hardware, implemented on a field-programmable gate array (FPGA). In addition, we verify the validity and evaluate the hardware cost of a countermeasure for the unrolled implementation, namely the Threshold Implementation (TI) countermeasure.
Ville Yli-Mäyry, Rei Ueno, Noriyuki Miura, Makoto Nagata, Shivam Bhasin, Yves Mathieu, Tarik Graba, Jean-Luc Danger, Naofumi Homma
IEEE Trans. Inf. Forensics Secur.2
2020 Machine Learning and Hardware security: Challenges and Opportunities -Invited Talk-
abstract
Machine learning techniques have significantly changed our lives. They helped improving our everyday routines, but they also demonstrated to be an extremely helpful tool for more advanced and complex applications. However, the implications of hardware security problems under a massive diffusion of machine learning techniques are still to be completely understood. This paper first highlights novel applications of machine learning for hardware security, such as evaluation of post quantum cryptography hardware and extraction of physically unclonable functions from neural networks. Later, practical model extraction attack based on electromagnetic side-channel measurements are demonstrated followed by a discussion of strategies to protect proprietary models by watermarking them.
Francesco Regazzoni 0001, Shivam Bhasin, Amir Ali Pour, Ihab Alshaer, Furkan Aydin, Aydin Aysu, Vincent Beroulle, Giorgio Di Natale, Paul D. Franzon, David Hély, Naofumi Homma, Akira Ito 0002, Dirmanto Jap, Priyank Kashyap, Ilia Polian, Seetal Potluri, Rei Ueno, Elena I. Vatajelu, Ville Yli-Mäyry
ICCAD17
2020 PMAC++: Incremental MAC Scheme Adaptable to Lightweight Block Ciphers
abstract
This paper presents a new incremental parallelizable message authentication code (MAC) scheme adaptable to lightweight block ciphers for memory integrity verification. The highlight of the proposed scheme is to achieve both incremental update capability and sufficient security bound with lightweight block ciphers, which is a novel feature. We extend the conventional parallelizable MAC to realize the incremental update capability while keeping the original security bound. We prove that a comparable security bound can be obtained even if this change is incorporated. We also present a hardware architecture for the proposed MAC scheme with lightweight block ciphers and demonstrate the effectiveness through FPGA implementation. The evaluation results indicate that the proposed MAC hardware achieves 3.4 times improvement in the latency-area product for the tag update compared with the conventional MAC.
Maya Oda, Rei Ueno, Akiko Inoue, Kazuhiko Minematsu, Naofumi Homma
ISCAS2
2020 High Throughput/Gate AES Hardware Architectures Based on Datapath Compression
abstract
This article proposes highly efficient Advanced Encryption Standard (AES) hardware architectures that support encryption and both encryption and decryption. New operation-reordering and register-retiming techniques presented in this article allow us to unify the inversion circuits in SubBytes and InvSubBytes without any delay overhead. In addition, a new optimization technique for minimizing linear mappings, named multiplicative-offset, further enhances the hardware efficiency. We also present a shared key scheduling datapath that can work on-the-fly in the proposed architecture. To the best of our knowledge, the proposed architecture has the shortest critical path delay and is the most efficient in terms of throughput per area among conventional AES encryption/decryption and encryption architectures with tower-field S-boxes. The proposed round-based architecture can perform AES encryption where block-wise parallelism is unavailable (e.g., cipher block chaining (CBC) mode); thus, our techniques can be globally applied to any type of architecture including pipelined ones. We evaluated the performance of the proposed and some conventional datapaths by logic synthesis with the NanGate 45-nm open-cell library. As a result, we can confirm that our proposed architectures achieve approximately 51-64 percent higher efficiency (i.e., higher bps/GE) and lower power/energy consumption than the other conventional counterparts.
Rei Ueno, Naofumi Homma, Sumio Morioka, Noriyuki Miura, Kohei Matsuda, Makoto Nagata, Shivam Bhasin, Yves Mathieu, Tarik Graba, Jean-Luc Danger
IEEE Trans. Computers1
2019 High Throughput/Gate FN-Based Hardware Architectures for AES-OTR
abstract
This paper presents high throughput/gates Feistel network (FN)-based AES-OTR hardware architectures. AES-OTR is an authenticated encryption (AE) scheme as a block cipher mode of operation using AES. While AES-OTR is one of the most theoretically efficient AEs using AES and has superior features, its practical efficiency in hardware is unclear due to no known reports of its hardware implementation. In this paper, we present efficient AES-OTR hardware architectures. In contrast to conventional AE architectures, our architecture forms the 2-round FN of OTR, which makes it easy to integrate the peripheral into hardware for OTR operations. The proposed architectures had 2.4 and 13.5 times higher throughput/gates than the de facto standard AE (i.e., AES-GCM) core on FPGA and ASIC, respectively, through logic syntheses.
Rei Ueno, Naofumi Homma, Tomonori Iida, Kazuhiko Minematsu
ISCAS1
2019 Tackling Biased PUFs Through Biased Masking: A Debiasing Method for Efficient Fuzzy Extractor
abstract
This paper presents an efficient fuzzy extractor (FE) design for biased physically unclonable functions (PUFs). To remove entropy leak from helper data in an efficient manner, we propose a new debiasing method, namely biased masking (BM). The proposed scheme removes the entropy leak by applying artificial noise (i.e., biased mask) such that the resulting response is uniform, and the added noise is removed by ECC decoding at the reconstruction as well as PUF noise. In addition, BM-based debiasing can be easily implemented with only additional random number generator and bit-parallel AND or OR operation in an enrollment server. Client devices with PUF, which are sometimes resource-constrained, require no additional operation. Furthermore, we show that the BM-based FE is reusable as well as the conventional code-offset FE. We evaluate the efficiency and effectiveness of the BM-based FE compared with the conventional debiasing-based FEs. Consequently, we confirm that the BM-based FE can achieve approximately 20 percent lower PUF size for nonnegligible biases (e.g., 60 percent) by just increasing the length of repetition code, which indicates that the BM-based FE is suitable for resource-constrained devices in terms of hardware cost for implementing PUF and computational cost at the reconstruction.
Rei Ueno, Manami Suzuki, Naofumi Homma
IEEE Trans. Computers1
2017 Automatic generation of formally-proven tamper-resistant Galois-field multipliers based on generalized masking scheme
abstract
In this study, we propose a formal design system for tamper-resistant cryptographic hardwares based on Generalized Masking Scheme (GMS). The masking scheme, which is a state-of-the-art masking-based countermeasure against higher-order differential power analyses (DPAs), can securely construct any kind of Galois-field (GF) arithmetic circuits at the register transfer level (RTL) description, while most other ones require specific physical design. In this study, we first present a formal design methodology of GMS-based GF arithmetic circuits based on a hierarchical dataflow graph, called GF arithmetic circuit graph (GF-ACG), and present a formal verification method for both functionality and security property based on Gröbner basis. In addition, we propose an automatic generation system for GMS-based GF multipliers, which can synthesize a fifth-order 256-bit multiplier (whose input bit-length is 256 × 77) within 15 min.
Rei Ueno, Naofumi Homma, Sumio Morioka, Takafumi Aoki
DATE1
2017 Formal Approach for Verifying Galois Field Arithmetic Circuits of Higher Degrees
abstract
This paper presents an efficient approach to verifying higher-degree Galois-field (GF) arithmetic circuits. The proposed method describes GF arithmetic circuits using a mathematical graph-based representation and verifies them by a combination of algebraic transformations and a new verification method based on natural deduction for first-order predicate logic with equal sign. The natural deduction method can verify one type of higher-degree GF arithmetic circuit efficiently while the existing methods require an enormous amount of time, if they can verify them at all. In this paper, we first apply the proposed method to the design and verification of various Reed-Solomon (RS) code decoders. We confirm that the proposed method can verify RS decoders with higher-degree functions while the existing method needs a lot of time or fail. In particular, we show that the proposed method can be applied to practical decoders with 8-bit symbols, which are performed with up to 2,040-bit operands. We then demonstrate the design and verification of the Advanced Encryption Standard (AES) encryption and decryption processors. As a result, the proposed method successfully verifies the AES decryption datapath while an existing method fails.
Rei Ueno, Naofumi Homma, Yukihiro Sugawara, Takafumi Aoki
IEEE Trans. Computers1
2016 A High Throughput/Gate AES Hardware Architecture by Compressing Encryption and Decryption Datapaths - Toward Efficient CBC-Mode Implementation
Rei Ueno, Sumio Morioka, Naofumi Homma, Takafumi Aoki
CHES1
2015 Highly Efficient GF(28) Inversion Circuit Based on Redundant GF Arithmetic and Its Application to AES Design
Rei Ueno, Naofumi Homma, Yukihiro Sugawara, Yasuyuki Nogami, Takafumi Aoki
CHES1