EDBT 2026 Demo / reviewers in the wild / expert
Naofumi Homma
dblp:72/699
· DBLP profile ↗
53ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0003-0864-3126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 3 first-author · 4 since 2021Security and privacy · 18 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Formal Security Proof of Masking - Reduction from Strong Noisy Leakage to Probing Model Without Random Probing and Application to LR Primitive
Rei Ueno, Akiko Inoue, Kazuhiko Minematsu, Akira Ito 0002, Naofumi Homma |
CRYPTO (7) | 5 |
| 2024 | Comparative Analysis and Implementation of Jump Address Masking for Preventing TEE Bypassing Fault AttacksabstractAttacks on embedded devices continue to evolve with the increasing number of applications in actual products. A trusted execution environment (TEE) enhances the security of embedded devices by isolating and protecting sensitive applications such as cryptography from malicious or vulnerable applications. However, the emergence of TEE bypass attacks using faults exposes TEEs to threats. In CHES’22, jump address masking (JAM) was proposed as a countermeasure against TEE bypass attacks, specifically targeting RISC-V. JAM prevents modifications of protected data by calculating jump addresses using the protected data, and is expected to provide promising resistance to TEE bypass attacks, for which traditional countermeasures are ineffective. However, JAM was originally proposed for bare metal applications. Therefore, its application to TEEs that operate with an OS presents technical and security challenges. This study proposes a method for applying JAM to Keystone, a major TEE framework for RISC-V, and validates its practical effectiveness and performance through a comparative evaluation with existing countermeasures such as memory encryption, random delays, and instruction duplication. Our evaluation reveals that the proposed JAM implementation is the first countermeasure that achieves complete resistance to TEE bypass attacks with an execution time overhead of approximately 340% for context switches and 1.0% across the entire program, which is acceptable compared with other countermeasures. Shoei Nashimoto, Rei Ueno, Naofumi Homma |
ARES | 3 |
| 2024 | Crystalor: Recoverable Memory Encryption Mechanism with Optimized Metadata Structure
Rei Ueno, Hiromichi Haneda, Naofumi Homma, Akiko Inoue, Kazuhiko Minematsu |
CCS | 3 |
| 2022 | On the Success Rate of Side-Channel Attacks on Masked Implementations: Information-Theoretical Bounds and Their Practical UsageabstractThis study derives information-theoretical bounds of the success rate (SR) of side-channel attacks on masked implementations. We first develop a communication channel model representing side-channel attacks on masked implementations. We then derive two SR bounds based on the conditional probability distribution and mutual information of shares. The basic idea is to evaluate the upper-bound of the mutual information between the non-masked secret value and the side-channel trace by the conditional probability distribution of shares given its leakage, with a help of the Walsh-Hadamard transform. With the derived theorems, we also prove the security of masking schemes: the SR decreases exponentially with an increase in the number of masking shares, under a much more relaxed condition than the previous proof. To validate and utilize our theorems in practice, we propose a deep-learning-based profiling method for approximating the conditional probability distribution of shares to estimate the SR bound and the number of traces required for attacking a given device. We experimentally confirm that our bounds are much stronger than the conventional bounds on masked implementations, which validates the relevance of our theorems to practice. Akira Ito 0002, Rei Ueno, Naofumi Homma |
CCS | 3 |
| 2022 | Efficient Modular Polynomial Multiplier for NTT Accelerator of Crystals-KyberabstractThis paper presents a hardware design that efficiently performs the number theoretic transform (NTT) for lattice-based cryptography. First, we propose an efficient modular multiplication method for lattice-based cryptography defined over Proth numbers. The proposed method is based on a K-RED technique specific to Proth numbers. In particular, we divide the intermediate result into the sign bit and the other absolute value bits and handle them separately to significantly reduce implementation costs. Then, we show a butterfly unit datapath of NTT and inverse INTT equipped with the proposed modular multiplier. We apply the proposed NTT accelerator to Crystals-Kyber, which is lattice-based cryptography, and evaluate its performance on Xilinx Artix-7. The results show that the proposed NTT accelerators achieve up-to 3% and 33% higher area-time efficiency in terms of LUTs and FFs, respectively, than conventional best methods. In addition, the low-latency version of the proposed NTT accelerators achieves a 18% lower-latency with an area-time efficiency (in terms of LUTs, FFs, and DSPs) than the existing fastest method. Yuma Itabashi, Rei Ueno, Naofumi Homma |
DSD | 3 |
| 2022 | High-Speed Hardware Architecture for Post-Quantum Diffie-Hellman Key Exchange Based on Residue Number SystemabstractThis paper presents a hardware architecture for a post-quantum key exchange protocol, named super-singular isogeny Diffie-Hellman (SIDH). The proposed hardware employs residue number system (RNS) and is optimized to reduce the latency of $\mathbb{F}_{p^{2}}$ multiplication and RNS Montgomery reduction, which are major time-consuming procedures in SIDH. The performance of the proposed hardware is validated and evaluated through an experimental implementation on Xilinx Kintex7 Ultrascale+. As a result, we confirm that the proposed hardware can perform an SIDH computation 34% faster than the state-of-the-art existing one on the same device at a resource overhead. Rei Ueno, Naofumi Homma |
ISCAS | 2 |
| 2022 | Efficient Formal Verification of Galois-Field Arithmetic Circuits Using ZDD Representation of Boolean PolynomialsabstractIn this study, we present a new formal method for verifying the functionality of Galois-field (GF) arithmetic circuits. Assuming that the input–output relation (i.e., the specification of a GF arithmetic circuit) can be represented as polynomials over 2, the proposed method formally checks the equivalence between GF polynomials derived from a netlist and the specification. To efficiently verify the equivalence, we employ a zero-suppressed binary decision diagram (ZDD) to represent polynomials over 2. Even though polynomial reduction is the most time-consuming process of verification (i.e., equivalence checking), our new algorithm can efficiently reduce the GF polynomials in the form of a zero-suppressed binary decision diagram derived from the target netlist. The proposed algorithm derives the polynomials representing all intermediate nodes (i.e., the outputs of all gates) in the order from primary inputs to those primary outputs that are in accordance with the reverse topological traversal order. We demonstrated the efficiency and effectiveness of the proposed method via a set of experimental verifications. In particular, we confirmed that the proposed method can verify practical GF multipliers (including those used in standardized elliptic curve cryptography) approximately 30 times faster on average and at most 170 times faster than the best conventional method. Akira Ito 0002, Rei Ueno, Naofumi Homma |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | ELM: A Low-Latency and Scalable Memory Encryption SchemeabstractMemory encryption (ME) with authentication is becoming a key security feature of modern processors, as evident by the adoption of ME by Intel’s SGX. Recently ME is actively studied from the viewpoint of system architecture. This paper studies ME from the viewpoint of symmetric-key cryptographic designs, with a primal focus on latency. A significant progress in such a direction can be observed in the SGX Integrity Tree (SIT). Using a variant of AES-GCM, SIT achieves an excellent latency. However, it has a scalability issue. By carefully examining SIT, we develop a new ME scheme dubbed ELM. We present an AES-based instantiation of ELM, and show that ELM significantly reduces latency from SIT for large memories, and achieves the provable security and equivalent hardware-protected (on-chip) area. We also present preliminary hardware implementations to substantiate our advantages. Akiko Inoue, Kazuhiko Minematsu, Maya Oda, Rei Ueno, Naofumi Homma |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Fault-Injection Attacks Against NIST's Post-Quantum Cryptography Round 3 KEM Candidates
Keita Xagawa, Akira Ito 0002, Rei Ueno, Junko Takahashi, Naofumi Homma |
ASIACRYPT (2) | 5 |
| 2021 | Extraction of Binarized Neural Network Architecture and Secret Parameters Using Side-Channel InformationabstractIn recent years, neural networks have been applied to various applications. To speed up the evaluation, a method using binarized network weights has been introduced, facilitating extremely efficient hardware implementation. Using electromagnetic (EM) side-channel analysis techniques, this study presents a framework of model extraction from practical binarized neural network (BNN) hardware. The target BNN hardware is generated and synthesized using open-source and commercial high-level synthesis tools GUINNESS and Xilinx SDSoC, respectively. With the hardware implemented on an up-to-date FPGA chip, we demonstrate how the layers can be identified from a single EM trace measured during the network evaluation, and we also demonstrate how an attacker may use side-channel attacks to recover secret weights used in the network. Ville Yli-Mäyry, Akira Ito 0002, Naofumi Homma, Shivam Bhasin, Dirmanto Jap |
ISCAS | 3 |
| 2021 | Imbalanced Data Problems in Deep Learning-Based Side-Channel Attacks: Analysis and SolutionabstractIn recent years, the threat of profiling attacks using deep learning has emerged. Successful attacks have been demonstrated against various types of cryptographic modules. However, the application of deep learning to side-channel attacks (SCAs) is often not adequately assessed because the labels that are widely used in SCAs, such as the Hamming weight (HW) and Hamming distance (HD), follow an imbalanced distribution. This study analyzes and solves the problems caused by dataset imbalance during training and inference. First, we state the reasons for the negative effect of data imbalance in classification for deep-learning-based SCAs and introduce the Kullback-Leibler (KL) divergence as a metric to measure this effect. Using the KL divergence, we demonstrate through analysis how the recently reported cross-entropy ratio loss function can solve the problem of imbalanced data. We further propose a method to solve dataset imbalance at the inference phase, which utilizes a likelihood function based on the key value instead of the HW/HD. The proposed method can be easily applied in deep-learning-based SCAs because it only needs an extra multiplication of the inverted binomial coefficients and inference results (i.e., the output probabilities) from the conventionally trained model. The proposed solution corresponds to data-augmentation techniques at the training phase, and furthermore, it better estimates the keys because the probability distributions of the training and test data are preserved. We demonstrate the validity of our analysis and the effectiveness of our solution through extensive experiments on two public databases. Akira Ito 0002, Kotaro Saito, Rei Ueno, Naofumi Homma |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Diffusional Side-Channel Leakage From Unrolled Lightweight Block Ciphers: A Case Study of Power Analysis on PRINCEabstractThis study investigates a new side-channel leakage observed in the inner rounds of an unrolled hardware implementation of block ciphers in a chosen-input attack scenario. The side-channel leakage occurs in the first round and it can be observed in the later inner rounds because it arises from path activation bias caused by the difference between two consecutive inputs. Therefore, a new attack that exploits the leakage is possible even for unrolled implementations equipped with countermeasures (masking and/or deglitchers that separate the circuit in terms of glitch propagation) in the round involving the leakage. We validate the existence of such a unique side-channel leakage through a set of experiments with a fully unrolled PRINCE cipher hardware, implemented on a field-programmable gate array (FPGA). In addition, we verify the validity and evaluate the hardware cost of a countermeasure for the unrolled implementation, namely the Threshold Implementation (TI) countermeasure. Ville Yli-Mäyry, Rei Ueno, Noriyuki Miura, Makoto Nagata, Shivam Bhasin, Yves Mathieu, Tarik Graba, Jean-Luc Danger, Naofumi Homma |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2020 | Machine Learning and Hardware security: Challenges and Opportunities -Invited Talk-abstractMachine learning techniques have significantly changed our lives. They helped improving our everyday routines, but they also demonstrated to be an extremely helpful tool for more advanced and complex applications. However, the implications of hardware security problems under a massive diffusion of machine learning techniques are still to be completely understood. This paper first highlights novel applications of machine learning for hardware security, such as evaluation of post quantum cryptography hardware and extraction of physically unclonable functions from neural networks. Later, practical model extraction attack based on electromagnetic side-channel measurements are demonstrated followed by a discussion of strategies to protect proprietary models by watermarking them. Francesco Regazzoni 0001, Shivam Bhasin, Amir Ali Pour, Ihab Alshaer, Furkan Aydin, Aydin Aysu, Vincent Beroulle, Giorgio Di Natale, Paul D. Franzon, David Hély, Naofumi Homma, Akira Ito 0002, Dirmanto Jap, Priyank Kashyap, Ilia Polian, Seetal Potluri, Rei Ueno, Elena I. Vatajelu, Ville Yli-Mäyry |
ICCAD | 11 |
| 2020 | PMAC++: Incremental MAC Scheme Adaptable to Lightweight Block CiphersabstractThis paper presents a new incremental parallelizable message authentication code (MAC) scheme adaptable to lightweight block ciphers for memory integrity verification. The highlight of the proposed scheme is to achieve both incremental update capability and sufficient security bound with lightweight block ciphers, which is a novel feature. We extend the conventional parallelizable MAC to realize the incremental update capability while keeping the original security bound. We prove that a comparable security bound can be obtained even if this change is incorporated. We also present a hardware architecture for the proposed MAC scheme with lightweight block ciphers and demonstrate the effectiveness through FPGA implementation. The evaluation results indicate that the proposed MAC hardware achieves 3.4 times improvement in the latency-area product for the tag update compared with the conventional MAC. Maya Oda, Rei Ueno, Akiko Inoue, Kazuhiko Minematsu, Naofumi Homma |
ISCAS | 5 |
| 2020 | High Throughput/Gate AES Hardware Architectures Based on Datapath CompressionabstractThis article proposes highly efficient Advanced Encryption Standard (AES) hardware architectures that support encryption and both encryption and decryption. New operation-reordering and register-retiming techniques presented in this article allow us to unify the inversion circuits in SubBytes and InvSubBytes without any delay overhead. In addition, a new optimization technique for minimizing linear mappings, named multiplicative-offset, further enhances the hardware efficiency. We also present a shared key scheduling datapath that can work on-the-fly in the proposed architecture. To the best of our knowledge, the proposed architecture has the shortest critical path delay and is the most efficient in terms of throughput per area among conventional AES encryption/decryption and encryption architectures with tower-field S-boxes. The proposed round-based architecture can perform AES encryption where block-wise parallelism is unavailable (e.g., cipher block chaining (CBC) mode); thus, our techniques can be globally applied to any type of architecture including pipelined ones. We evaluated the performance of the proposed and some conventional datapaths by logic synthesis with the NanGate 45-nm open-cell library. As a result, we can confirm that our proposed architectures achieve approximately 51-64 percent higher efficiency (i.e., higher bps/GE) and lower power/energy consumption than the other conventional counterparts. Rei Ueno, Naofumi Homma, Sumio Morioka, Noriyuki Miura, Kohei Matsuda, Makoto Nagata, Shivam Bhasin, Yves Mathieu, Tarik Graba, Jean-Luc Danger |
IEEE Trans. Computers | 2 |
| 2019 | High Throughput/Gate FN-Based Hardware Architectures for AES-OTRabstractThis paper presents high throughput/gates Feistel network (FN)-based AES-OTR hardware architectures. AES-OTR is an authenticated encryption (AE) scheme as a block cipher mode of operation using AES. While AES-OTR is one of the most theoretically efficient AEs using AES and has superior features, its practical efficiency in hardware is unclear due to no known reports of its hardware implementation. In this paper, we present efficient AES-OTR hardware architectures. In contrast to conventional AE architectures, our architecture forms the 2-round FN of OTR, which makes it easy to integrate the peripheral into hardware for OTR operations. The proposed architectures had 2.4 and 13.5 times higher throughput/gates than the de facto standard AE (i.e., AES-GCM) core on FPGA and ASIC, respectively, through logic syntheses. Rei Ueno, Naofumi Homma, Tomonori Iida, Kazuhiko Minematsu |
ISCAS | 2 |
| 2019 | Tackling Biased PUFs Through Biased Masking: A Debiasing Method for Efficient Fuzzy ExtractorabstractThis paper presents an efficient fuzzy extractor (FE) design for biased physically unclonable functions (PUFs). To remove entropy leak from helper data in an efficient manner, we propose a new debiasing method, namely biased masking (BM). The proposed scheme removes the entropy leak by applying artificial noise (i.e., biased mask) such that the resulting response is uniform, and the added noise is removed by ECC decoding at the reconstruction as well as PUF noise. In addition, BM-based debiasing can be easily implemented with only additional random number generator and bit-parallel AND or OR operation in an enrollment server. Client devices with PUF, which are sometimes resource-constrained, require no additional operation. Furthermore, we show that the BM-based FE is reusable as well as the conventional code-offset FE. We evaluate the efficiency and effectiveness of the BM-based FE compared with the conventional debiasing-based FEs. Consequently, we confirm that the BM-based FE can achieve approximately 20 percent lower PUF size for nonnegligible biases (e.g., 60 percent) by just increasing the length of repetition code, which indicates that the BM-based FE is suitable for resource-constrained devices in terms of hardware cost for implementing PUF and computational cost at the reconstruction. Rei Ueno, Manami Suzuki, Naofumi Homma |
IEEE Trans. Computers | 3 |
| 2017 | Automatic generation of formally-proven tamper-resistant Galois-field multipliers based on generalized masking schemeabstractIn this study, we propose a formal design system for tamper-resistant cryptographic hardwares based on Generalized Masking Scheme (GMS). The masking scheme, which is a state-of-the-art masking-based countermeasure against higher-order differential power analyses (DPAs), can securely construct any kind of Galois-field (GF) arithmetic circuits at the register transfer level (RTL) description, while most other ones require specific physical design. In this study, we first present a formal design methodology of GMS-based GF arithmetic circuits based on a hierarchical dataflow graph, called GF arithmetic circuit graph (GF-ACG), and present a formal verification method for both functionality and security property based on Gröbner basis. In addition, we propose an automatic generation system for GMS-based GF multipliers, which can synthesize a fifth-order 256-bit multiplier (whose input bit-length is 256 × 77) within 15 min. Rei Ueno, Naofumi Homma, Sumio Morioka, Takafumi Aoki |
DATE | 2 |
| 2017 | Design Methodology and Validity Verification for a Reactive Countermeasure Against EM Attacks
Naofumi Homma, Yuichi Hayashi, Noriyuki Miura, Daisuke Fujimoto, Makoto Nagata, Takafumi Aoki |
J. Cryptol. | 1 |
| 2017 | Formal Approach for Verifying Galois Field Arithmetic Circuits of Higher DegreesabstractThis paper presents an efficient approach to verifying higher-degree Galois-field (GF) arithmetic circuits. The proposed method describes GF arithmetic circuits using a mathematical graph-based representation and verifies them by a combination of algebraic transformations and a new verification method based on natural deduction for first-order predicate logic with equal sign. The natural deduction method can verify one type of higher-degree GF arithmetic circuit efficiently while the existing methods require an enormous amount of time, if they can verify them at all. In this paper, we first apply the proposed method to the design and verification of various Reed-Solomon (RS) code decoders. We confirm that the proposed method can verify RS decoders with higher-degree functions while the existing method needs a lot of time or fail. In particular, we show that the proposed method can be applied to practical decoders with 8-bit symbols, which are performed with up to 2,040-bit operands. We then demonstrate the design and verification of the Advanced Encryption Standard (AES) encryption and decryption processors. As a result, the proposed method successfully verifies the AES decryption datapath while an existing method fails. Rei Ueno, Naofumi Homma, Yukihiro Sugawara, Takafumi Aoki |
IEEE Trans. Computers | 2 |
| 2016 | A High Throughput/Gate AES Hardware Architecture by Compressing Encryption and Decryption Datapaths - Toward Efficient CBC-Mode Implementation
Rei Ueno, Sumio Morioka, Naofumi Homma, Takafumi Aoki |
CHES | 3 |
| 2015 | A DPA/DEMA/LEMA-resistant AES cryptographic processor with supply-current equalizer and micro EM probe sensorabstractCombination of a supply-current equalizer (EQ) and a micro EM probe sensor (EMS) exhibits strong resiliency against major three DPA/DEMA/LEMA low-cost side-channel attacks on a cryptographic processor. Test-chip measurements with 128bit AES cryptographic processor in 0.18μm CMOS successfully demonstrate the secret key protection from all three attacks. A digital-oriented circuit implementation together with a careful design optimization minimize the hardware overhead of EQ and EMS to +33%, +1.6% in area, +7.6%, +0.15% in power, and ~0%, -0.2% in performance of an unprotected AES, respectively. Daisuke Fujimoto, Noriyuki Miura, Yuichi Hayashi, Naofumi Homma, Takafumi Aoki, Makoto Nagata |
ASP-DAC | 4 |
| 2015 | Highly Efficient GF(28) Inversion Circuit Based on Redundant GF Arithmetic and Its Application to AES Design
Rei Ueno, Naofumi Homma, Yukihiro Sugawara, Yasuyuki Nogami, Takafumi Aoki |
CHES | 2 |
| 2015 | EM attack sensor: concept, circuit, and design-automation methodologyabstractA side-channel attack exploiting EM-field leakage from a cryptographic processor IC is an existing serious threat to our information society. EM radiation during the IC operation is captured by an EM probe and the correlation to the crypto processing is statistically analyzed to reveal the secret information although it is protected in a software (algorithm) domain. This paper presents a reactive hardware (implementation) domain countermeasure against this EM attack, namely EM attack sensor. An on-chip sensor coil detects EM probe approach and reacts to protect the secret information from the tamper attack. The sensor concept and low-cost digital circuit implementation are reviewed, and the detail of the design-automation methodology highly-compatible to standard EDA tools is presented. A small hardware overhead of the sensor is silicon-proven in an actual 0.18μm CMOS test-chip implementation together with a 128bit AES crypto core. The test-chip measurements demonstrate successful sensor operation against the actual EM probe attack. Noriyuki Miura, Daisuke Fujimoto, Makoto Nagata, Naofumi Homma, Yuichi Hayashi, Takafumi Aoki |
DAC | 4 |
| 2015 | A Silicon-Level Countermeasure Against Fault Sensitivity Analysis and Its EvaluationabstractIn this paper, we present an efficient countermeasure against fault sensitivity analysis (FSA) based on configurable delay blocks (CDBs). FSA is a new type of fault attack, which exploits the relationship between fault sensitivity (FS) and secret information. Previous studies reported that it could break cryptographic modules equipped with conventional countermeasures against differential fault analysis (DFA), such as redundancy calculation, masked and-or, and wave dynamic differential logic. The proposed countermeasure can thwart both DFA and FSA attacks based on setup time violation faults. The proposed ideas are to use a CDB as a time base for detection and to combine the technique with Li's countermeasure concept that removes the dependency between FSs and secret data. The postmanufacture configuration of the CDBs allows minimization of the overhead in operating frequency that comes from manufacture variability. In this paper, we also present an implementation of the proposed countermeasure in application-specified integrated circuit, and describe its configuration method. We then investigate the hardware overhead of the proposed countermeasure for an advanced encryption standard processor and demonstrate its validity through an experiment. Sho Endo, Yang Li 0001, Naofumi Homma, Kazuo Sakiyama, Kazuo Ohta, Daisuke Fujimoto, Makoto Nagata, Toshihiro Katashita, Jean-Luc Danger, Takafumi Aoki |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | A Threat for Tablet PCs in Public Space: Remote Visualization of Screen Images Using EM EmanationabstractThe use of tablet PCs is spreading rapidly, and accordingly users browsing and inputting personal information in public spaces can often be seen by third parties. Unlike conventional mobile phones and notebook PCs equipped with distinct input devices (e.g., keyboards), tablet PCs have touchscreen keyboards for data input. Such integration of display and input device increases the potential for harm when the display is captured by malicious attackers. This paper presents the description of reconstructing tablet PC displays via measurement of electromagnetic (EM) emanation. In conventional studies, such EM display capture has been achieved by using non-portable setups. Those studies also assumed that a large amount of time was available in advance of capture to obtain the electrical parameters of the target display. In contrast, this paper demonstrates that such EM display capture is feasible in real time by a setup that fits in an attaché case. The screen image reconstruction is achieved by performing a prior course profiling and a complemental signal processing instead of the conventional fine parameter tuning. Such complemental processing can eliminate the differences of leakage parameters among individuals and therefore correct the distortions of images. The attack distance, 2 m, makes this method a practical threat to general tablet PCs in public places. This paper discusses possible attack scenarios based on the setup described above. In addition, we describe a mechanism of EM emanation from tablet PCs and a countermeasure against such EM display capture. Yuichi Hayashi, Naofumi Homma, Mamoru Miura, Takafumi Aoki, Hideaki Sone |
CCS | 2 |
| 2014 | EM Attack Is Non-invasive? - Design Methodology and Validity Verification of EM Attack Sensor
Naofumi Homma, Yuichi Hayashi, Noriyuki Miura, Daisuke Fujimoto, Daichi Tanaka, Makoto Nagata, Takafumi Aoki |
CHES | 1 |
| 2014 | Toward Formal Design of Practical Cryptographic Hardware Based on Galois Field ArithmeticabstractThis paper presents a formal method for designing cryptographic processor datapaths on the basis of arithmetic circuits over Galois fields (GFs). The proposed method describes GF arithmetic circuits in the form of hierarchical graph structures, where nodes represent sub-circuits whose functions are defined by arithmetic formulae over GFs, and edges represent data dependency between nodes. In this paper, we first introduce the application of graph representation to arithmetic circuits over extension fields of${\mbi {GF}}({{\mbi {p}}^{\mbi {m}}})$$({\mbi {p}} \geq {\bf 2})$and composite fields, which are commonly used in the design of cryptographic processors. The newly proposed graph representation can be formally verified through symbolic computation techniques based on polynomial reduction and Gröbner basis. We then demonstrate the capabilities of the proposed approach through an experimental design of a 128-bit AES (Advanced Encryption Standard) datapath including multiplicative inversion circuits over the composite field${\mbi {GF}}{(((2^2)^2)^2})$. The results show that the proposed method can describe such practical datapaths, as well as that complete verification of such a datapath can be carried out within a short period of time. Naofumi Homma, Kazuya Saito, Takafumi Aoki |
IEEE Trans. Computers | 1 |
| 2012 | An Efficient Countermeasure against Fault Sensitivity Analysis Using Configurable Delay BlocksabstractIn this paper, we present an efficient countermeasure against Fault Sensitivity Analysis (FSA) based on a configurable delay blocks (CDBs). FSA is a new type of fault attack which exploits the relationship between fault sensitivity and secret information. Previous studies reported that it could break cryptographic modules equipped with conventional countermeasures against Differential Fault Analysis (DFA) such as redundancy calculation, Masked AND-OR and Wave Dynamic Differential Logic (WDDL). The proposed countermeasure can detect both DFA and FSA attacks based on setup time violation faults. The proposed ideas are to use a CDB as a time base for detection and to combine the technique with Li's countermeasure concept which removes the dependency between fault sensitivities and secret data. Post-manufacture configuration of the delay blocks allows minimization of the overhead in operating frequency which comes from manufacture variability. In this paper, we present an implementation of the proposed countermeasure, and describe its configuration method. We also investigate the hardware overhead of the proposed countermeasure implemented in ASIC for an AES module and demonstrate its validity through an experiment using a prototype FPGA implementation. Sho Endo, Yang Li 0001, Naofumi Homma, Kazuo Sakiyama, Kazuo Ohta, Takafumi Aoki |
FDTC | 3 |
| 2012 | A Formal Approach to Designing Cryptographic Processors Based on $GF(2^m)$ Arithmetic CircuitsabstractThis paper proposes a formal approach to designing Galois-field (GF) arithmetic circuits, which are widely used in modern cryptographic processors. Our method describes GF arithmetic circuits in a hierarchical manner with high-level directed graphs associated with specific GFs and arithmetic functions. The proposed circuit description can be effectively verified by symbolic computations based on polynomial reduction using Grobner bases. The verified description is then translated into the equivalent hardware description language (HDL) codes, which are available for the conventional design flow. We first describe the proposed graph representation and present an example of the description and verification. The significant advantage of the proposed approach is demonstrated through experimental designs of parallel multipliers over GF(2m) for different word lengths and irreducible polynomials. The result shows that the proposed approach has a definite capability of formally verifying practical GF arithmetic circuits for which the conventional techniques fail. We also propose an application of this approach to cryptographic processor design. The target considered here is a 128-bit advanced encryption standard (AES) data path with a loop architecture. To the best of the authors' knowledge, this is the first verification of this type of practical AES data path. We present a detailed description of the AES data path and its verification. The proposed approach successfully verifies the AES data path description within 800 s. Naofumi Homma, Kazuya Saito, Takafumi Aoki |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Fair and Consistent Hardware Evaluation of Fourteen Round Two SHA-3 CandidatesabstractThe first contribution of our paper is that we propose a platform, a design strategy, and evaluation criteria for a fair and consistent hardware evaluation of the second-round SHA-3 candidates. Using a SASEBO-GII field-programmable gate array (FPGA) board as a common platform, combined with well defined hardware and software interfaces, we compare all 256-bit version candidates with respect to area, throughput, latency, power, and energy consumption. Our approach defines a standard testing harness for SHA-3 candidates, including the interface specification for the SHA-3 module on our testing platform. The second contribution is that we provide both FPGA and 90-nm CMOS application-specific integrated circuit (ASIC) synthesis results and thereby are able to compare the results. Our third contribution is that we release the source code of all the candidates and by using a common, fixed, publicly available platform, our claimed results become reproducible and open for a public verification. Miroslav Knezevic, Kazuyuki Kobayashi, Jun Ikegami, Shin'ichiro Matsuo, Akashi Satoh, Ünal Koçabas, Junfeng Fan, Toshihiro Katashita, Takeshi Sugawara 0001, Kazuo Sakiyama, Ingrid Verbauwhede, Kazuo Ohta, Naofumi Homma, Takafumi Aoki |
IEEE Trans. Very Large Scale Integr. Syst. | 13 |
| 2011 | Enhancement of simple electro-magnetic attacks by pre-characterization in frequency domain and demodulation techniques
Olivier Meynard, Denis Réal, Florent Flament, Sylvain Guilley, Naofumi Homma, Jean-Luc Danger |
DATE | 5 |
| 2011 | Systematic Design of RSA Processors Based on High-Radix Montgomery MultipliersabstractThis paper presents a systematic design approach to provide the optimized Rivest-Shamir-Adleman (RSA) processors based on high-radix Montgomery multipliers satisfying various user requirements, such as circuit area, operating time, and resistance against side-channel attacks. In order to involve the tradeoff between the performance and the resistance, we apply four types of exponentiation algorithms: two variants of the binary method with/without Chinese Remainder Theorem (CRT). We also introduces three multiplier-based datapath-architectures using different intermediate data forms: 1) single form, 2) semi carry-save form, and 3) carry-save form, and combined them with a wide variety of arithmetic components. Their radices are parameterized from 28to 2128. A total of 242 datapaths for 1024-bit RSA processors were obtained for each radix. The potential of the proposed approach is demonstrated through an experimental synthesis of all possible processors with a 90-nm CMOS standard cell library. As a result, the smallest design of 861 gates with 118.47 ms/RSA to the fastest design of 0.67 ms/RSA at 153\thinspace 862 gates were obtained. In addition, the use of the CRT technique reduced the RSA operation time of the fastest design to 0.24 ms. Even if we employed the exponentiation algorithm resistant to typical side-channel attacks, the fastest design can perform the RSA operation in less than 1.0 ms. Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Comparative Power Analysis of Modular Exponentiation AlgorithmsabstractThis paper proposes new chosen-message power-analysis attacks for public-key cryptosystems based on modular exponentiation, where specific input pairs are used to generate collisions between squaring operations at different locations in the two power traces. Unlike previous attacks of this kind, the new attack can be applied to all standard implementations of the exponentiation process, namely binary (left-to-right and right-to-left), m-ary, and sliding window methods. The proposed attack can also circumvent typical countermeasures, such as the Montgomery powering ladder and the double-add algorithm. The effectiveness of the attack is demonstrated in experiments with hardware and software implementations of RSA on an FPGA and a PowerPC processor, respectively. In addition to the new collision generation methods, a highly accurate waveform matching technique is introduced for detecting the collisions even when the recorded signals are noisy and there is a certain amount of clock jitter. Naofumi Homma, Atsushi Miyamoto, Takafumi Aoki, Akashi Satoh, Adi Shamir |
IEEE Trans. Computers | 1 |
| 2009 | Evaluation of Simple/Comparative Power Analysis against an RSA ASIC ImplementationabstractSimple power analysis attacks with chosen-message techniques were applied to an RSA processor implemented with standard CMOS technology on ASIC, and the different characteristics of power waveforms caused by two types of implementation (ASIC and FPGA) were investigated in detail. We also applied comparative power analysis an advanced power analysis attack in which a pair of input data was used to enhance the waveform pattern for modular exponentiation. The power dissipation of modular squaring in the difference waveform was greatly reduced when compared to modular multiplication, allowing all of the secret key bits to be successfully revealed. Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2008 | Collision-Based Power Analysis of Modular Exponentiation Using Chosen-Message Pairs
Naofumi Homma, Atsushi Miyamoto, Takafumi Aoki, Akashi Satoh, Adi Shamir |
CHES | 1 |
| 2008 | High-Performance Concurrent Error Detection Scheme for AES Hardware
Akashi Satoh, Takeshi Sugawara 0001, Naofumi Homma, Takafumi Aoki |
CHES | 3 |
| 2008 | Chosen-message SPA attacks against FPGA-based RSA hardware implementationsabstractThis paper presents SPA (Simple Power Analysis) attacks against public-key cryptosystems implemented on an FPGA platform. The SPA attack investigates a power waveform generated by a cryptographic module, and reveals a secret key in the module. We focus on chosen-message SPA attacks, which enhances the differences of operating waveforms between multiplication and squaring correlated to the secret key by using the input of particular messages. In particular, Yen showed a unique SPA attack against RSA cryptosystem, but no verification experiment using actual software or hardware was performed. In this paper, we implemented four-types of RSA processors on an FPGA platform in combination with two variants of the Montgomery multiplication algorithm and two different types of multipliers for SPA attacks experiments. Then we demonstrated effectiveness of various chosen-message attacks as well as Yen’s method, and investigated the characteristics of the attacks depending on the hardware architectures. Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
FPL | 2 |
| 2008 | Systematic design of high-radix Montgomery multipliers for RSA processorsabstractThe present paper proposes a systematic design approach to provide the optimal high-radix Montgomery multipliers for an RSA processor satisfying user requirements. We introduces three multiplier-based architectures using different intermediate-data forms ((i) single form, (ii) semi carry-save form, and (iii) carry-save form), and combined them with a wide variety of arithmetic components. Their radices are also parameterized from 28to 264. A total of 202 designs for 1,024-bit RSA processors were obtained for each radix, and were synthesized using a 90-nm CMOS standard cell library. The smallest design of 0.9 Kgates with 137.8 ms/RSA to the fastest design of 1.8 ms/RSA at 74.7 Kgates were then obtained. In addition, the optimal design to meet the user requirements can be easily obtained from all the combinations. In addition to choosing the datapath architecture, the arithmetic component, and the radix parameters, the proposed systematic approach can also adopt other process technologies. Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ICCD | 2 |
| 2008 | Enhanced power analysis attack using chosen message against RSA hardware implementationsabstractSPA (Simple Power Analysis) attacks against RSA cryptosystems are enhanced by using chosen-message scenarios. One of the most powerful chosen-message SPA attacks was proposed by Yen et. al. in 2005, which can be applied to various algorithms and architectures, and can defeat the most popular SPA countermeasure using dummy multiplication. Special input values of −1 and a pair of −X and X can be used to identify squaring operations performed depending on key bit stream. However, no experimental result on actual implementation was reported. In this paper, we implemented some RSA processors on an FPGA platform and demonstrated that Yen’s attack with a signal filtering technique clearly reveal the secret key information in the actual power waveforms. Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2008 | High-performance ASIC implementations of the 128-bit block cipher CLEFIAabstractIn the present paper, we introduce high-performance hardware architectures for the 128-bit block cipher CLEFIA and evaluate their ASIC performances in comparison with the ISO/IEC 18033-3 standard block ciphers (AES, Camellia, SEED, CAST-128, MISTY1, and TDEA). We designed five types of hardware architectures for CLEFIA, combining two loop structures and three F-functions. These designs were synthesized with a 90-nm CMOS standard cell library, and size and speed performances were evaluated. The highest hardware efficiency (defined as throughput/gates) obtained was 400.96 Kbps/gates, which is 1.5 times higher than previously achieved efficiencies. Takeshi Sugawara 0001, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2008 | Arithmetic module generator with algorithm optimization capabilityabstractThis paper presents an arithmetic module generator based on an arithmetic description language called ARITH. The use of ARITH makes it possible to describe a wide variety of arithmetic algorithms in a unified manner. The ARITH descriptions are formally verified in the generator even if the arithmetic algorithms include unconventional number systems for operands or internal variables. The proposed generator also optimizes arithmetic algorithms by using performance profiles derived from the previous generation. From these features, we can obtain high-performance arithmetic modules whose functions are completely verified at the algorithm level. In this paper, we demonstrate that the optimal prefix adders improved the performance of generated arithmetic modules such as multipliers in comparison with the standard prefix adders. Yuki Watanabe, Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
ISCAS | 2 |
| 2008 | A Systematic Approach for Designing Redundant Arithmetic Adders Based on Counter Tree DiagramsabstractThis paper introduces a systematic approach to designing high-performance parallel adders based on Counter Tree Diagrams (CTDs). By using CTDs, we can describe addition algorithms at various levels of abstraction. A high-level CTD represents a network of coarse-grained components associated with word-level operands, whereas a low-level CTD represents a network of primitive components that can be directly mapped onto physical devices. The level of abstraction in circuit representation can be changed by decomposition of CTDs. We can derive possible variations of adder structures by decomposing a high-level CTD into low-level CTDs in a formal manner. In this paper, we focus on an application of CTDs to the design of redundant arithmetic adders with limited carry propagation. For any redundant number representation, we can obtain the optimal adder structure by trying every possible CTD decomposition and CTD-variable encoding. The potential of the proposed approach is demonstrated through an experimental synthesis of Redundant-Binary (RB) adders with CMOS standard cell libraries. We can successfully obtain RB adders that achieve an about 30-40% improvement in terms of power-delay product compared with conventional designs. Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
IEEE Trans. Computers | 1 |
| 2007 | Application of symbolic computer algebra to arithmetic circuit verificationabstractThis paper presents a formal approach to verify arithmetic circuits using symbolic computer algebra. Our method describes arithmetic circuits directly with high-level mathematical objects based on weighted number systems and arithmetic formulae. Such circuit description can be effectively verified by polynomial reduction techniques using Grobner Bases. In this paper, we describe how the symbolic computer algebra can be used to describe and verify arithmetic circuits. The advantageous effects of the proposed approach are demonstrated through experimental verification of some arithmetic circuits such as multiply-accumulator and FIR filter. The result shows that the proposed approach has a definite possibility of verifying practical arithmetic circuits where the conventional techniques failed. Yuki Watanabe, Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
ICCD | 2 |
| 2007 | SPA against an FPGA-Based RSA Implementation with a High-Radix Montgomery MultiplierabstractSimple power analysis (SPA) was applied to an RSA processor with a high-radix Montgomery multiplier on an FPGA platform, and the different characteristics of power waveforms caused by two types of multiplier (built-in and custom) were investigated in detail. The authors also applied an active attack where input data was set to a specific pattern to control the modular multiplication. The power dissipation for the multiplication was greatly reduced in comparison with modular squaring, resulting in success in revealing all of the secret key bits Atsushi Miyamoto, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2007 | DPA Using Phase-Based Waveform Matching against Random-Delay CountermeasureabstractWe propose differential power analysis (DPA) with a phase-based waveform matching technique. Conventionally, a trigger signal and a system clock are used to capture the waveform traces, but the signals always contain jitter-related deviations, and this degrades the accuracy of the statistical analysis. Our method can adjust for this timing deviation with a higher resolution than the sampling rate by post-processing on the measured waveforms. Therefore, no modification of the measuring equipment is required. Our method can also defeat DPA countermeasures creating distorted waveforms with random delays or dummy cycles. We implemented Data Encryption Standard (DES) software with and without the countermeasure on a Z80 microprocessor, and demonstrated the advantages of our method in comparison with a conventional attack. Sei Nagashima, Naofumi Homma, Yuichi Imai, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2007 | A High-Performance ASIC Implementation of the 64-bit Block Cipher CAST-128abstractThe authors propose a compact hardware architecture for the 64-bit block cipher CAST-128, which is one of the ISO/IEC 18033-3 standard algorithms. Part of the complexity of CAST-128 is its use of various S-boxes in various sequences, and three types of f-function are switched depending on the round numbers. Therefore a large amount of hardware resources are required for a straight-forward implementation. In order to create compact CAST-128 hardware, the authors minimized the number of S-box components, and merged the three f-functions into one arithmetic component. The CAST-128 hardware based on the proposed architecture was synthesized using 0.13μm and 0.18-μm CMOS standard cell libraries and small, practical circuits of 26.4-39.5 Kgates and 189.9-614.7 Mbps were obtained. Takeshi Sugawara 0001, Naofumi Homma, Takafumi Aoki, Akashi Satoh |
ISCAS | 2 |
| 2006 | High-Resolution Side-Channel Attack Using Phase-Based Waveform Matching
Naofumi Homma, Sei Nagashima, Yuichi Imai, Takafumi Aoki, Akashi Satoh |
CHES | 1 |
| 2004 | Topology-Oriented Design of Analog Circuits Based on Evolutionary Graph Generation
Masanori Natsui, Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
PPSN | 2 |
| 2003 | VLSI circuit design using an object-oriented framework of evolutionary graph generation systemabstractThis paper presents a generic objected-oriented framework of evolutionary graph generation (EGG) for automated circuit synthesis. The EGG system can be systematically implemented for different design problems by inheriting the framework class templates. The potential capability of EGG framework is demonstrated through experimental synthesis of both digital and analogue circuits. Design examples discussed in this paper are: (i) bit-serial multipliers using bit-level arithmetic components; and (ii) current mirrors using transistor-level components. Naofumi Homma, Masanori Natsui, Takafumi Aoki, Tatsuo Higuchi 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2002 | Graph-based individual representation for evolutionary synthesis of arithmetic circuitsabstractThis paper presents a graph-based evolutionary optimization technique, called evolutionary graph generation (EGG), to synthesize arithmetic circuits. The potential capability of EGG has been investigated through an experiment of synthesizing fast constant-coefficient multipliers. Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2002 | Evolutionary Graph Generation System and Its Application to Bit-Serial Arithmetic Circuit Synthesis
Makoto Motegi, Naofumi Homma, Takafumi Aoki, Tatsuo Higuchi 0001 |
PPSN | 2 |
| 2002 | Graph-based evolutionary design of arithmetic circuitsabstractWe present an efficient graph-based evolutionary optimization technique, called evolutionary graph generation (EGG), and the proposed approach is applied to the design of combinational and sequential arithmetic circuits based on parallel counter-tree architecture. The fundamental idea of EGG is to employ general circuit graphs as individuals and manipulate the circuit graphs directly using new evolutionary graph operations without encoding the graphs into other indirect representations, such as the bit strings used in genetic algorithm (GA) proposed by Holland (1992) and trees used in genetic programming (GP) proposed by Koza et al. (1997). In this paper, the EGG system is applied to the design of constant-coefficient multipliers and the design of bit-serial data-parallel adders. The results demonstrate the potential capability of EGG to solve the practical design problems for arithmetic circuits with limited knowledge of computer arithmetic algorithms. The proposed EGG system can help to simplify and speed up the process of designing arithmetic circuits and can produce better solutions to the given problem. Dingjun Chen, Takafumi Aoki, Naofumi Homma, Toshiki Terasaki, Tatsuo Higuchi 0001 |
IEEE Trans. Evol. Comput. | 3 |