Lang Li 0002

dblp:l/LangLi2 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0002-4832-4499ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 9 since 2021Computer networks · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The ILLcipher family of low-latency block ciphers for industrial internet of things
Wei Sun 0020, Lang Li 0002
Comput. Networks2
2026 Cakr: a collision-aware cryptanalysis scheme for lightweight block ciphers
abstract
Abstract Partial neural distinguishers limit the available ciphertext bit combinations in differential neural cryptanalysis. When the training data size and the number of bits are not appropriately selected, label collisions can occur, which adversely affects key recovery efficiency. This paper conducts an analysis to investigate the correlation between the number of bits and the data size, aiming to address the aforementioned issue. It develops a strategy to control collisions and mitigate the impact of these collisions on model performance. A Collision-Aware Key Recovery (CAKR) framework is proposed tailored for high-collision data based on this strategy. This framework leverages the distribution characteristics of labels, eliminating the need for training neural distinguishers and significantly reducing both time and resource consumption. Experimental results show that the CAKR framework reduces the key recovery time by 96.8%, 95.5%, and 91.0% for the Speck32/64, Speck64/96, and Speck96/128, respectively. Additionally, a bit search algorithm is proposed that incorporates a differential evolution strategy and uses the non-uniformity of the ciphertext difference distribution among positive samples as the fitness criterion. Frequent calls to the neural distinguisher are avoided by our method, reducing the search time from 3.286 h to 7.464 s for 8-bit combinations in Speck32/64. The CAKR framework also offers a quantum version that theoretically further reduces time complexity.
Siqi Zhu, Lang Li 0002, Ruihan Xu 0006, Zhiwen Hu, Yemao Hu
Cybersecur.2
2026 Average Relative SNR: New Metric to Evaluate the Attack Performance of Non-Profiled Side-Channel Traces
Lang Li 0002
J. Electron. Test.2
2026 HDHL: A hybrid GSP lightweight block cipher with two-round high diffusion
Xingqi Yue, Lang Li 0002, Qingling Song
Integr.2
2026 NTSD: An Efficient Method to Enhance the Dataset for Differential-Neural Cryptanalysis
abstract
Differential-neural cryptanalysis has become a frontier method for evaluating the security of block ciphers. This optimizes the security boundaries of embedded devices in the Internet of Things (IoT) more effectively. However, it is challenging to construct an effective dataset for the differential neural distinguisher. The main challenge lies in that when expanding to deep rounds, the distribution of data features becomes scattered and a computational complexity disaster arises. Therefore, this paper proposes a dual coupling optimization model called the Normality Test Search Dataset Model (NTSD). The model achieves synergistic breakthroughs in search dataset enhancement and search efficiency. First, we propose a differential-key constraint mechanism. The mechanism enhances the anti-decay capability of the dataset by constraining the number of random keys and the form of ciphertexts in the dataset. This makes the probability of collision higher and also strengthens the overall features. Additionally, the distribution features of the enhanced dataset in high rounds are significantly different from those of random datasets. Second, we propose a more efficient dual evaluation strategy. This strategy employs a dual normality distribution detection method to reduce the complexity of searching input differences. Experiments show that the execution time of the NTSD model is approximately 80% less than that of Seok’s method. Finally, the NTSD model is also used in a variety of ciphers. For instance, GIFT-COFB has advanced from round 4 to round 6. The accuracy of ASCON-PERMUTATION in four rounds has increased from 50.69% to 58.23%.
Zhiwen Hu, Lang Li 0002, Yemao Hu
IEEE Internet Things J.2
2026 TmsNet: An Efficient Transformer Network Based on Multiscale Shift-Invariance Mechanism for Side-Channel Analysis
abstract
Transformer architectures have shown strong potential for profiling side-channel analysis (SCA), but their high computational complexity and substantial parameter overhead remain major obstacles to practical deployment in resource-constrained scenarios, especially under severe leakage-trace desynchronization. To address this challenge, we propose TmsNet, an efficient Transformer-based architecture tailored for desynchronized side-channel leakage analysis. TmsNet revisits the Transformer backbone by coupling multi-scale leakage representation with shift-aware modeling, thereby reducing model complexity while maintaining strong key-recovery capability under pronounced temporal misalignment. TmsNet is characterized by three key design components. First, it employs a multi-scale feature extraction module with temporal convolution kernels of different receptive fields to jointly capture fine-grained local leakage patterns and longer-range temporal dependencies, thereby reducing the sensitivity of vanilla Transformers to local misalignment. Second, it introduces a residual fusion mechanism to align and aggregate cross-scale features, yielding more robust feature representations under temporal perturbations. Third, it integrates relative positional encoding with Softmax-based normalization, replacing conventional LayerNorm to better preserve subtle leakage contrasts that may be weakened by mean-variance normalization, while mitigating the impact of countermeasure-induced random delays and clock jitter. Extensive evaluations on the desynchronized ASCADf, ASCADr, and TinyPower benchmark datasets show that TmsNet reduces parameter count by up to 78.5% relative to baseline models while maintaining competitive attack performance. Notably, it recovers the secret key using only 48, 42, and 28 attack traces on the three datasets, respectively, demonstrating an effective balance between architectural efficiency and practical profiling SCA capability.
Shengtao Tang, Lang Li 0002, Lianrui Deng, Xingqi Yue
IEEE Internet Things J.2
2026 ILAD: A hardware-efficient authenticated encryption scheme for VANET applications based on Ascon
Jiali Tang, Lang Li 0002, Xingqi Yue
J. Netw. Comput. Appl.2
2026 Low-Latency Implementation of Bitsliced SPN-Cipher on IoT Processors
abstract
Bitsliced cipher implementations demonstrate enhanced performance and security on high-end processors through specialized instruction utilization. However, IoT processors, particularly 32-bit architectures, present significant implementation challenges due to limited register sizes and instruction sets, hindering efficient parallelism in bitsliced SPN ciphers. This study presents optimization strategies for implementing bitsliced SPN ciphers on 32-bit processors using common instruction sets. The linear layer optimization employs a decomposition algorithm that transforms complex permutation operations into minimal instruction sequences. This approach recursively identifies optimal instruction combinations while maintaining computational efficiency. For the non-linear layer, operations are constrained to basic logic instructions (NOT, AND, OR, XOR). A novel encoding method for the Bit-slice Gate Complexity (BGC) model is proposed to optimize S-box transformations within these constraints using Boolean satisfiability solvers. Additionally, a comprehensive benchmarking framework facilitates standardized performance evaluation across implementations. Experimental evaluation of the optimized implementations on ARM Cortex-M and Xtensa LX processors demonstrates significant performance improvements. The proposed techniques achieve reductions of 9.7% and 67.6% in Cycles Per Byte for AES and QARMAv2 implementations, respectively.
Jiahao Xiang, Lang Li 0002
IEEE Trans. Computers2
2026 KD-SCA: Improving Lightweight CNN Model Profiling Side-Channel Analysis With Knowledge Distillation
Lianrui Deng, Lang Li 0002, Jiahao Xiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Construction of Lightweight S-Boxes With Low Boomerang Uniformity
Qingling Song, Lang Li 0002, Liuyan Yan
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 LLBC: A Novel Feistel-Based Low-Latency Block Cipher for IoT Applications
abstract
Low-latency has been an important criterion in the design of block ciphers, especially in lightweight cryptography for IoT (Internet of Things) constrained devices to ensure secure real-time data transmission with limited resources. However, the design of most low-latency block ciphers today employ the so-called SPN (Substitution Permutation Network) structure with a few exceptions such as the SCARF cipher. In this article, by adopting the ideas of parallel execution for bridging the gap in latency between the standard Feistel and SPN structures, we propose a new low-latency cipher that uses the standard Feistel structure named as LLBC (Low Latency Block Cipher). It has a 128-bit block length with 128-bit (or 256-bit) key length. For the purpose of minimizing the latency and implementation costs, we were able to specify two 4-bit S-boxes, which not only have good cryptographic properties but are also excellent in terms of their hardware performance. More specifically, these S-boxes require only 18.5 GEs (Gate Equivalents) and have depth 3, which to the best of our knowledge are currently the best performing low-latency 4-bit S-boxes. Using these S-boxes in the design of LLBC, we achieve a significant reduction in both latency and implementation cost compared to Midori and QARMA. For instance, implementation of LLBC on the NanGate 45 nm open cell library achieves delay of about 2.79 ns which can be compared to the delay of PRINCE, Midori and QARMA being 4.06 ns, 4.94 ns, and 4.02 ns, respectively. Moreover, a standard security analysis shows that LLBC has enough security margin against various known attacks.
Yongzhuang Wei, Enes Pasalic, Lang Li 0002, Ting Fan
IEEE Internet Things J.4
2025 QLW: a lightweight block cipher with high diffusion
Xingqi Yue, Lang Li 0002, Jiahao Xiang, Zhiwen Hu
J. Supercomput.2
2024 IoVCipher: A low-latency lightweight block cipher for internet of vehicles
Xiantong Huang, Lang Li 0002, Jinling Yang, Juanli Kuang
Ad Hoc Networks2
2024 INLEC: An involutive and low energy lightweight block cipher for internet of things
Lang Li 0002, Liuyan Yan, Chutian Deng
Pervasive Mob. Comput.2
2024 Side-channel analysis based on Siamese neural network
Lang Li 0002, Yu Ou
J. Supercomput.2
2023 An efficient differential analysis method based on deep learning
Lang Li 0002, Ying Guo 0006, Yu Ou, Xiantong Huang
Comput. Networks2
2023 DBST: a lightweight block cipher based on dynamic S-box
Liuyan Yan, Lang Li 0002, Ying Guo 0006
Frontiers Comput. Sci.2
2023 SAND-2: An optimized implementation of lightweight block cipher
Lang Li 0002, Ying Guo 0006
Integr.2
2022 SCENERY: a lightweight block cipher based on Feistel structure
Jingya Feng, Lang Li 0002
Frontiers Comput. Sci.2
2022 Side-channel analysis attacks based on deep learning network
Yu Ou, Lang Li 0002
Frontiers Comput. Sci.2
2022 DULBC: A dynamic ultra-lightweight block cipher with high-throughput
Jinling Yang, Lang Li 0002, Ying Guo 0006, Xiantong Huang
Integr.2
2022 A new S-box construction method meeting strict avalanche criterion
Lang Li 0002, Jinggen Liu, Ying Guo 0006
J. Inf. Secur. Appl.1
2021 Shadow: A Lightweight Block Cipher for IoT Nodes
abstract
The advancement of the Internet of Things (IoT) has promoted the rapid development of low-power and multifunctional sensors. However, it is seriously significant to ensure the security of data transmission of these nodes. Meanwhile, sensor nodes have the characteristics of converting analog signals into digital signals for operation processing in wireless sensor networks (WSNs). Given the particularity of Addition or AND, Rotation, and XOR (ARX) operations, its round function can only be based on the Feistel structure or generalized Feistel structure, otherwise, the process of decryption cannot be completed correctly. Furthermore, the existing ARX ciphers have the problems of only changing half of the plaintext block in one round and iterating for many rounds. In this article, a new logical combination method of generalized Feistel structure and ARX operations is proposed to improve the diffusion speed of ARX ciphers, called Shadow. Shadow overcomes the shortcomings of traditional ARX ciphers that only diffuse half of the block in one round. To ensure the efficiency of the encryption hardware circuit while ensuring the security of the physical-layer signal, we studied the round-based hardware architecture and the serial hardware architecture for Shadow cipher. Particularly, we conducted a series of performance tests on Shadow, including the avalanche effect, FPGA implementation, and ASIC implementation. Also, we conducted a security analysis of the Shadow. As shown by our experiments and comparisons, Shadow is compact in IoT nodes and is of high security against cryptanalysis.
Ying Guo 0006, Lang Li 0002
IEEE Internet Things J.2
2021 Implementation of PRINCE with resource-efficient structures based on FPGAs
abstract
In this era of pervasive computing, low-resource devices have been deployed in various fields. PRINCE is a lightweight block cipher designed for low latency, and is suitable for pervasive computing applications. In this paper, we propose new circuit structures for PRINCE components by sharing and simplifying logic circuits, to achieve the goal of using a smaller number of logic gates to obtain the same result. Based on the new circuit structures of components and the best sharing among components, we propose three new hardware architectures for PRINCE. The architectures are simulated and synthesized on different programmable gate array devices. The results on Virtex-6 show that compared with existing architectures, the resource consumption of the unrolled, low-cost, and two-cycle architectures is reduced by 73, 119, and 380 slices, respectively. The low-cost architecture costs only 137 slices. The unrolled architecture costs 409 slices and has a throughput of 5.34 Gb/s. To our knowledge, for the hardware implementation of PRINCE, the new low-cost architecture sets new area records, and the new unrolled architecture sets new throughput records. Therefore, the newly proposed architectures are more resource-efficient and suitable for lightweight, latency-critical applications.
Lang Li 0002, Jingya Feng, Ying Guo 0006
Frontiers Inf. Technol. Electron. Eng.1
2020 Research on a high-order AES mask anti-power attack
abstract
The cryptographic algorithm has been gradually improved in design, but its implementations are vulnerable to side‐channel analysis (SCA). Generally speaking, adding a mask to the primitive is the best way to counteract SCA. In the high‐order mask, the key to affecting performance and security lies in the multiplication design. Based on the research of the advanced encryption standard (AES) algorithm, internal round function structure, and zero‐knowledge proof, a high‐order AES mask scheme is designed to optimise the implementation. In this scheme, the substitution‐box protects sensitive variables in the algorithm with the use of secure multiplication and secure inversion by column. The scheme named as in columns higher‐order mask (ICHM), features low cost and high security. The result of the experiment proves the security and effectiveness of the ICHM.
Yu Ou, Lang Li 0002
IET Inf. Secur.2