Lu Li 0006

dblp:72/2266-6 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-6947-654XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Fast Fourier Transform and Gaussian Sampling Instructions Designed for FALCON Digital Signature
Tao-Yun Wang, Shuai-Yu Chen, Lu Li 0006, Weijia Wang 0003
Inscrypt (2)3
2024 Compact Instruction Set Extensions for Kyber
abstract
Kyber is the only post-quantum cryptography (PQC) key encapsulation mechanism in the National Institute of Standards and Technology PQC project. This brief investigates the design of compact instruction set extensions (ISEs) for Kyber. We focus on implementing number-theoretic transform (NTT) and propose a hardware design of the modular multiplication based on an optimized$k^{2}$-reduction. Compared to other works, our design is more compact since the optimized$k^{2}$-reduction comprises multiplications with significantly smaller multipliers than Montgomery reduction and Barrett reduction. Then, we integrate the$k^{2}$-reduction into an instruction for the butterfly transformation. We also propose auxiliary instructions that can switch the half words between two registers to facilitate the rearranging coefficients in NTT. To showcase the advantage of the instructions, we implement the ISEs in a chip design for the Hummingbird E203 core. Compared to the software implementation on RISC-V with assembly code, our co-design implementations for NTT show a speedup by a factor of 2.6. Besides, the area overhead is 93 LUTs and 1 DSP without any additional resources of FFs and RAMs using Artix-7 FPGA, which is more compact than previous software–hardware co-designs of Kyber.
Lu Li 0006, Guofeng Qin, Yang Yu 0008, Weijia Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 ISA Extensions of Shuffling Against Side-Channel Attacks
abstract
Shuffling is a time-randomized countermeasure against side-channel attacks. To achieve effective protections, shuffling is usually combined with other countermeasures, such as the masking. It requires the shuffling to be as efficient as possible. In this work, we describe an instruction set extensions (ISEs) for shuffling countermeasure. Our ISEs focuses on the generation of random permutations, which is the most difficult part to deploy the shuffling in microprocessors. The Thorp shuffling is implemented in hardware, enabling the instruction to generate random permutations. We design new ISEs compatible to the RISC-V standard instruction set format. Then, we present applications of our ISEs by giving two combinations of shuffling and masking, which can be regarded as promising software–hardware co-designs of side-channel countermeasures. At last, we embed the ISEs to the RISC-V core called tinyriscv, and evaluate the silicon overhead and the side-channel security of the shuffled masked AND operation. The evaluation shows that the new instruction can significantly improve the security of masking countermeasures.
Jiayun Zhou, Guofeng Qin, Lu Li 0006, Chun Guo 0002, Weijia Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Compact Instruction Set Extensions for Dilithium
abstract
Post-quantum cryptography is considered to provide security against both traditional and quantum computer attacks. Dilithium is a digital signature algorithm that derives its security from the challenge of finding short vectors in lattices. It has been selected as one of the standardizations in the NIST post-quantum cryptography project. Hardware-software co-design is a commonly adopted implementation strategy to address various implementation challenges, including limited resources, high performance, and flexibility requirements. In this study, we investigate using compact instruction set extensions (ISEs) for Dilithium, aiming to improve software efficiency with low hardware overheads. To begin with, we propose tightly coupled accelerators that are deeply integrated into the RISC-V processor. These accelerators target the most computationally demanding components in resource-constrained processors, such as polynomial generation, Number Theoretic Transform (NTT), and modular arithmetic. Next, we design a set of custom instructions that seamlessly integrate with the RISC-V base instruction formats, completing the accelerators in a compact manner. Subsequently, we implement our ISEs in a chip design for the Hummingbird E203 core and conduct performance benchmarks for Dilithium utilizing these ISEs. Additionally, we evaluate the resource consumption of the ISEs on FPGA and ASIC technologies. Compared to the reference software implementation on the RISC-V core, our co-design demonstrates a remarkable speedup factor ranging from 6.95 to 9.96. This significant improvement in performance is achieved by incorporating additional hardware resources, specifically, a 35% increase in LUTs, a 14% increase in FFs, 7 additional DSPs, and no additional RAM. Furthermore, compared to the state-of-the-art approach, our work achieves faster speed performance with a reduced circuit cost. Specifically, the usage of additional LUTs, FFs, and RAMs is reduced by 47.53%, 50.43%, and 100%, respectively. On ASIC technology, our approach demonstrates 12, 412 cell counts. Our co-design provides a better tradeoff implementation on speed performance and circuit overheads.
Lu Li 0006, Guofeng Qin, Shuaiyu Chen, Weijia Wang 0003
ACM Trans. Embed. Comput. Syst.1
2023 Bit-Sliced Implementation of SM4 and New Performance Records
abstract
SM4 is a popular block cipher issued by the Office of State Commercial Cryptography Administration (OSCCA) of China. In this paper, we use the bit‐slicing technique that has been shown as a powerful strategy to achieve very fast software implementations of SM4. We investigate optimizations on two frontiers. First, we present a more efficient bit‐sliced representation for SM4, which enables running 64 blocks in parallel with 256‐bit registers. Second, we describe an optimized algorithm for data form transformations, also allowing efficient implementations of SM4 under Counter (CTR) mode and Galois/Counter mode. The above optimizations contribute to a significant performance gain on one core compared with the state‐of‐the‐art results. This work is an extension of the conference paper at Inscrypt 2022, awarded the best paper award.
Lu Li 0006, Chun Guo 0002, Meiqin Wang 0001, Weijia Wang 0003
IET Inf. Secur.2
2018 Conditional cube attack on round-reduced River Keyak
Wenquan Bi, Zheng Li 0008, Xiaoyang Dong 0001, Lu Li 0006, Xiaoyun Wang 0001
Des. Codes Cryptogr.4
2018 Improved integral attacks without full codebook
abstract
The integral attack, exploits the balanced property of the output in the distinguisher. Usually, adversaries append some rounds after the distinguisher, guess the corresponding key bits and check whether the target bits are balanced. Few works add rounds before the distinguisher to make the key recovery attack. In the first full‐round attack on MISTY1, Todo adds one FL layer (key‐dependent linear function) before the distinguisher. In this study, the authors extend his method and give a general method, which they can use to extend some rounds (non‐linear) before the distinguisher to attack more rounds with data complexity smaller than the whole space and little extra time consumption. The basic idea is that for different subkeys guessed in the forward rounds, they set different constant values for the input of the distinguisher. Finally, the selected data space is not full. For substitution permutation network (SPN) (Feistel with SPN round function) structures with 4 bit S‐box and bit permutation, they estimate the data complexity when adding one round before the distinguishers for all 4 bit S‐boxes. Using the method, they improve the integral attacks on PRESENT, RECTANGLE, TWINE and LBlock, and their results could cover one more round.
Zhihui Chu, Huaifeng Chen, Xiaoyun Wang 0001, Lu Li 0006, Xiaoyang Dong 0001, Yaoling Ding, Yonglin Hao
IET Inf. Secur.4
2018 Improved Integral Attacks on SIMON32 and SIMON48 with Dynamic Key-Guessing Techniques
abstract
Dynamic key-guessing techniques, which exploit the property of AND operation, could improve the differential and linear cryptanalytic results by reducing the number of guessed subkey bits and lead to good cryptanalytic results for SIMON. They have only been applied in differential and linear attacks as far as we know. In this paper, dynamic key-guessing techniques are first introduced in integral cryptanalysis. According to the features of integral cryptanalysis, we extend dynamic key-guessing techniques and get better integral cryptanalysis results than before. As a result, we present integral attacks on 24-round SIMON32, 24-round SIMON48/72, and 25-round SIMON48/96. In terms of the number of attacked rounds, our attack on SIMON32 is better than any previously known attacks, and our attacks on SIMON48 are the same as the best attacks.
Zhihui Chu, Huaifeng Chen, Xiaoyun Wang 0001, Xiaoyang Dong 0001, Lu Li 0006
Secur. Commun. Networks5