VLDB 2026 Research / reviewers in the wild / expert
An Wang 0001
dblp:06/4924-1
· DBLP profile ↗
48ranked-venue papers
6as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 19 · 3 first-author · 10 since 2021Systems, architecture and hardware · 13 · 11 since 2021Computer networks · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperNTT: An Ultra-High Throughput Number Theoretic Transform Accelerator for FHE
Xiyan Dong, Leyan Zhang, An Wang 0001, Xinghua Wang 0005, Liehuang Zhu |
ISCAS | 5 |
| 2026 | More Practical and Robust: Enhancing Simple Power Analysis on Cryptosystems With Double ClusteringabstractThe widespread use of public key cryptographic algorithms in embedded devices has made them a primary target for side-channel analysis. Clustering-based Simple Power Analysis (SPA) poses a significant threat to public key implementations by inferring secret keys through the identification of distinguishable patterns in side-channel information. However, traditional clustering-based SPA methods are highly sensitive—even to non-key-dependent patterns—thereby limiting their robustness and practical applicability. To address these limitations, this paper proposes a double clustering method that enhances the flexibility, accuracy, and robustness of clustering-based SPA. By progressively adjusting the the number of clusters, the method adaptively identifies optimal clustering configurations, mitigating the need for fixed assumptions and improving resistance to noise and other interfering factors. Experiments covering multiple cryptographic algorithms, hardware platforms, and countermeasure settings demonstrate that the proposed method consistently outperforms traditional clustering-based SPA methods. Annyu Liu, Weijia Wang 0003, An Wang 0001 |
IEEE Internet Things J. | 4 |
| 2026 | AutoProfile: Automated profiling in deep learning-based side-channel analysis
Changshan Su, Yuxing Tang, An Wang 0001 |
Neural Networks | 7 |
| 2026 | Locality Does Matter: An Assessment Metric Adapted for Cluster-Based Side-Channel Analysis on Public Key CryptosystemsabstractCluster-based side-channel analysis (SCA) is a commonly used side-channel analysis method for public key cryptography systems. This paper focuses on assessment metrics to improve the efficiency and accuracy of key recovery in cluster-based SCA processes. We introduce a novel metric called adjacent distance coefficient (AD coefficient). Different from traditional metrics like silhouette coefficient, membership degree, and information entropy, the AD coefficient, by considering local density and avoiding the computation of cluster centers, is less influenced by the shape and size of data, thereby overcoming limitations of traditional metrics. Experiments on RSA, SM2, and ECC demonstrate that the AD coefficient exhibits higher accuracy compared to traditional metrics, especially under conditions with noise and random delays, where it shows significant robustness. Incorporating the AD coefficient, we propose an assessment method to detect whether cryptographic devices are resistant to cluster-based SCA attacks, offering a convenient quantitative measure for the security evaluation of cryptographic devices. Jinghong Ding, An Wang 0001, Congming Wei, Weiping Gong, Jingjie Wu, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | An Efficient Ensemble Framework to Assist Profiled Side-Channel Analysis by Machine LearningabstractThe application of machine learning techniques in side-channel analysis has recently received increased attention. Finding the best hyperparameters to achieve optimal performance for machine learning models in side-channel analysis is still a challenging endeavor. In order to solve the problem, we present an efficient ensemble framework designed to support profiled side-channel analysis for attacking cryptographic devices with countermeasures. Our proposed framework can partially mitigate the impact of traditional countermeasures employed in cryptographic devices. Additionally, we introduce a novel voting method called elite voting, which leverages candidate keys with higher probabilities to recover the secret key and adjusts the voting weights for better candidate keys. Experimental results illustrate that our proposed framework can effectively recover the right key from cryptographic devices with countermeasures through multiple experiments. It enhances the signal-to-noise ratio of traces and successfully recovers the right key across various datasets. Furthermore, when compared to traditional methods, our elite voting method further enhances the performance of ensemble learning by reducing the number of traces needed to recover the secret key. It exhibits superior performance compared to other ensemble methods, as it can reduce the minimum required number of traces significantly. Yaoling Ding, An Wang 0001, Shaofei Sun, Congming Wei, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | HTMMM: Novel Hybrid Truncated Montgomery Modular Multiplication Algorithm and Hardware ArchitectureabstractModular multiplication is one of the key operations in modern public-key cryptography. Montgomery Modular Multiplication (MMM) is a mainstream method to avoid modulo operations, which contains one variable multiplication and two constant multiplications. In this paper, for high performance, a novel Hybrid Truncated Montgomery Modular Multiplication (HTMMM) algorithm and its hardware architecture are proposed, which achieves state-of-the-art Area-Time-Product (ATP) and throughput. We propose an error-free high-part truncated multiplication for the first time, which solves the problem that conventional methods cannot be applied to MMM due to the introduced error, and reduces the complexity to the same level as low-part truncated multiplication. Besides, a hybrid multiplication based on Toom-Cook and Karatsuba is proposed to optimize variable multiplication, Non-Adjacent Form (NAF) encoding is adopted with truncated multiplication to optimize constant multiplications. The quantitative analysis of complexity for the integer multipliers with different schemes are illustrated to find the optimal multiplier under various cases. Based on these, we took the bit widthN= 1024 as an example to introduce the hardware architecture in detail and gave the implementation results ofN= 256 andN= 1024 in different processes. The experimental results demonstrate that compared with the best existing design, the throughput and ATP of our proposed design are improved by 1.25× and 2.22×, respectively. Zeying Li, Yue Hao 0007, Hongshuo Li, An Wang 0001, Zhiming Chen 0001, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | A UMAP-Based Clustering Side-Channel Analysis on Public-Key CryptosystemsabstractHorizontal analysis is a widely adopted method in side-channel analysis, particularly for public-key cryptosystems, where attackers aim to recover the key from a single trace. Current methods rely on trace segmentation, dimensionality reduction, and classification, but high noise and poor feature preservation hinder accuracy. Noise blurs cryptographic operation segment boundaries, and existing dimensionality reduction techniques fail to maintain the inherent distribution of trace points in a high-dimensional space. As a result, secret information recovery based on clustering remains inaccurate. This paper proposes an automated horizontal analysis framework named UMAP-HC to improve secret information recovery accuracy. The framework employs a sliding segmentation method to locate cryptographic operations in noisy traces with blurred segment boundaries. It leverages uniform manifold approximation and projection (UMAP) for feature preservation and hierarchical clustering for secret information recovery. Experimental results on four open access public-key algorithm power trace datasets, an SM2 power trace collected from a smart card, and an ECC power trace with dummy operation countermeasures demonstrate that UMAP-HC effectively classifies cryptographic operations, accurately locates operation segments, and recovers secret key. It achieves up to 100% recovery accuracy, surpassing previous methods by 40%-70%, with normalized mutual information reaching 1, an improvement of 0.02-0.99 over existing approaches. Yuhan Qian, Yaoling Ding, Shaofei Sun, Congming Wei, An Wang 0001, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Fama: An FPGA-Oriented Multiscalar Multiplication Accelerator Optimized via Algorithm-Hardware Co-DesignabstractMulti-scalar multiplication (MSM) is the primary computational bottleneck in zero-knowledge proof protocols. To address this, we introduce FAMA, an FPGA-oriented MSM accelerator developed through algorithm-hardware co-optimization. By integrating a 3D-Pippenger optimization algorithm, FAMA minimizes computational complexity, while its compact dual-mode point addition (PADD) unit significantly reduces hardware overhead. Compared to the best CPU-based design, FAMA achieves over 184.20× speedup. It also outperforms state-of-the-art FPGA-based MSM accelerators, reducing resource overhead by more than 64% and boosting area-time product (ATP) by up to 37.09×. Xiyan Dong, An Wang 0001, Xinghua Wang 0005, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Enhanced Template Attack Against Dilithium: Leveraging Dual-Loss Feature ExtractionabstractAs a post-quantum digital signature scheme, Dilithium was specifically designed to withstand known quantum algorithm attacks, and its side-channel resistance has garnered significant research attention. However, current side-channel attacks against Dilithium exhibit several limitations: (1) failure to leverage low-correlation characteristics in power traces, (2) loss functions limited to categorical information extraction from power traces, (3) dependency on specific coefficient recovery conditions while neglecting inter-coefficient statistical dependencies, (4) requirement for separate profiling models per intermediate value, resulting in substantial information loss. To address these limitations, we propose an enhanced template attack framework integrating deep learning with classical template attack methodology. Our approach employs a dual-loss similarity learning mechanism for feature extraction from high-dimensional power traces, enabling the construction of more discriminative templates while preserving weakly correlated features. Through assembly-level analysis of the y polynomial generation routine, we reveal inherent correlations among coefficientsyk0,yk1,yk2,yk3. Building on this discovery, our dual-loss similarity learning framework is designed to capture these inter-coefficient relationships, preserving their intrinsic dependencies while achieving effective inter-class separation and intra-class aggregation properties, which significantly enhances the effectiveness of subsequent template attacks. Experimental results on Cortex-M4 power traces demonstrate our method achieves 32.94% polynomial coefficient recovery accuracy for polynomial coefficients y, outperforming conventional SOD-based (83% improvement), T-Test-based (97%), and PCA-based template attacks (197% enhancement). Furthermore, complete private key recovery is achieved with merely 14 power traces under specific conditions. This DL-enhanced template attack framework demonstrates superior side-channel leakage exploitation, yielding substantial performance enhancements over conventional approaches. Haojin Zhang, Qingjun Yuan, Yaoling Ding, An Wang 0001, Hailong Zhang 0001, Haopeng Fan, Siqi Lu, Yongjuan Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Practical Differential Fault Attacks on the GPRS Standard CiphersabstractGEA-1 and GEA-2 are two standard stream ciphers used in GPRS (General Packet Radio Service) to protect against eavesdropping GPRS between the base station and the phone. Now, a range of current phones still support them. In this paper, a differential fault attack on the GEA-like stream ciphers under the random fault model is proposed for the first time. In this attack, an efficient dedicated algorithm for identifying the exact fault location is proposed. By using this dedicated algorithm, the attacker can succeed in determining the exact fault location. As applications, practical differential fault attacks on the GPRS standard ciphers (i.e., GEA-1 and GEA-2) are presented, which recover the 64-bit secret keys of GEA-1 and GEA-2 with time complexities of${2^{{\mathrm{{33}}}{\mathrm{{.807}}}}}$and${2^{{\mathrm{{33}}}{\mathrm{{.858}}}}}$, respectively. We validate the cryptanalytic results by simulating the whole attacks on the platform ChipWhisperer Lite. The experimental results show that both GEA-1 and GEA-2 can be broken within sixteen minutes on a common laptop. Finally, the possible countermeasures are presented to protect the processed data of massive GPRS devices. Zhengting Li, Lin Ding 0001, An Wang 0001, Haotong Xu, Zheng Liu 0029, Jiang Wan |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Bridging Lab and Industry: Practical SPA-GPT on Cryptosystems Boosted by LSTM and Simulated AnnealingabstractSimple Power Analysis is a commonly used method in Side-Channel Analysis on cryptosystems, which requires a significant amount of labor costs for segmentation. General Pulse Tailor for Simple Power Analysis (SPA-GPT) proposed in CHES 2024 utilizes reinforcement learning to achieve automated segmentation. However, its low efficiency and only targeting public-key algorithms limit its practical applications. In this paper, we propose a practical method, which utilize long short-term memory network and attention mechanism, coupled with a new deep Q-network policy using Simulated Annealing strategy, to solve the contradiction between reinforcement learning and high efficiency in trace segmentation. Moreover, the novel agent proposed in this paper also demonstrates transferability, enabling direct segmentation of a trace under varying lengths and signal-to-noise ratio conditions once the agent has been fully trained. In addition, our new approach is applicable for locating each execution of block ciphers in various encryption modes. Comparative experiments are conducted on 14 datasets, which are collected from software or hardware implementations of RSA, ECC, ML-KEM, AES, PRESENT, and SIMON, running on microcontrollers, FPGAs, or smart cards. Experimental results show that the new method enhances time efficiency by 50.34% to 94.24% while reducing network parameters by 87.84% compared to SPA-GPT. Yaoling Ding, An Wang 0001, Congming Wei, Liehuang Zhu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Programming Equation Systems of Arithmetization-Oriented Primitives with Constraints
Kexin Qiao, Mengyu Chang, Junjie Cheng, Changhai Ou, An Wang 0001, Liehuang Zhu |
Inscrypt (2) | 5 |
| 2025 | ProverNG: Efficient Verification of Compositional Masking for Cryptosystem's Side-Channel Security
Limin Fan, An Wang 0001 |
ICICS (3) | 6 |
| 2025 | PPCM-Fed: Privacy-Preserving Cross-Modal Federated Learning in IoT
Dongjue Wang, Keke Gai, Jing Yu 0007, An Wang 0001, Zhijing Cao, Liehuang Zhu |
SecureComm (5) | 4 |
| 2025 | Intelligent Re-encryption and Commitment-Driven Dynamic Data Sharing
An Wang 0001, Keke Gai, Jing Yu 0007, Liehuang Zhu |
WASA (1) | 2 |
| 2025 | Low-Latency and Area-Efficient Elliptic Curve Point Multiplication Architectures Over Koblitz CurvesabstractPoint multiplication is the core operation in elliptic curve cryptography. Koblitz curves are a special class of curves that can utilize the Frobenius mapping to accelerate the implementation of point multiplication operations. For point multiplication on Koblitz curves, this paper first proposes an optimized tNAF scalar conversion algorithm along with its corresponding hardware architecture. Additionally, for the computation of point multiplication, this paper proposes two optimal computational architectures: an area-efficient architecture and a low-latency architecture, both of which achieve the highest pipeline efficiency. The area-efficient architecture adopts a compact four-stage pipeline with a single multiplier, ensuring high circuit area utilization efficiency while achieving relatively low computation latency. The low-latency architecture implements two-stage and three-stage pipeline designs respectively in different binary fields, using two multipliers to reduce the clock cycles for point addition and further decrease the computation latency. The proposed architectures were implemented on Virtex-7 FPGA. For the GF(2163), GF(2283), and GF(2571) fields, the latency for the area-efficient architecture are 1.683μs, 3.455μs and 7.511μs, with slices usage of 3631, 7867 and 20612, and the point multiplication latency for the low-latency architecture are 1.347μs, 3.279μs and 7.071μs, with slices usage of 6026, 14246 and 38515. A comparison with state-of-the-art designs shows that the proposed point multiplication architectures offer significant advantages in terms of performance. In the GF(2163) field, the computation latency of the area-efficient architecture and the low-latency architecture is reduced by at least 21.158% and 43.952%, respectively. And in the GF(2283) field, the reduction in latency is 38.985% and 42.459%, while in the GF(2571) field, the reduction in latency is 58.111% and 59.720%. An Wang 0001, Yue Hao 0007, Zhiming Chen 0001, Liehuang Zhu |
IEEE Internet Things J. | 3 |
| 2025 | An Intelligent Framework for Cluster-Based Side-Channel Analysis on Public-Key CryptosystemsabstractClassical cluster-based side-channel analysis (SCA) uses clustering algorithms to analyze power traces and often, principal component analysis to reduce the dimension of data, resulting in that clustering may not deal well with high-dimensional traces, such as cryptographic algorithm implementations with countermeasures. In this article, we propose an intelligent framework for cluster-based SCA, which includes three steps of clustering, classification and correction, for processing large high-dimensional data. By combining unsupervised clustering and supervised deep learning techniques, the framework succeeds in mining the data for additional in-depth information. In addition, unlike traditional cluster-based SCA, our approach focuses on deep learning and deliberately avoids over-reliance on cluster labels during classification. And metrics for correction are adopted to achieve a high level of reliability in key recovery. Experiments on the RSA smart card based on Montgomery ladder implementation and FPGA-based ECC with random delay demonstrate that our framework can significantly improve the success rate with strong robustness. Congming Wei, Shulin He, An Wang 0001, Shaofei Sun, Yaoling Ding, Liehuang Zhu |
IEEE Internet Things J. | 3 |
| 2025 | Make It Easy! Timing Leakage Analysis on Cryptographic Chips Based on Horizontal LeakageabstractTiming analysis presents a significant threat to cryptographic modules. However, traditional timing leakage analysis has notable limitations, especially when precise execution times cannot be obtained. In this article, we propose a novel timing leakage analysis method that leverages horizontal leakage in the power/electromagnetic channel by detecting the trace length of encryption processes under varying inputs. To demonstrate the effectiveness of our approach, we conducted systematic experimental evaluations across a range of cryptographic devices. In comparison to timing leakage analysis based on plaintext-ciphertext correlation, our method offers higher accuracy at lower testing costs and exhibits improved resistance to vertical noise. Guangze Hong, An Wang 0001, Congming Wei, Yaoling Ding, Shaofei Sun, Liehuang Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | CL-SCA: A Contrastive Learning Approach for Profiled Side-Channel AnalysisabstractSide-channel analysis (SCA) based on machine learning, particularly neural networks, has gained considerable attention in recent years. However, previous works predominantly focus on establishing connections between labels and related profiled traces. These approaches primarily capture label-related features and often overlook the connections between traces of the same label, resulting in the loss of some valuable information. Besides, the attack traces also contain valuable information that can be used in the training process to assist model learning. In this paper, we propose a profiled SCA approach based on contrastive learning named CL-SCA to address these issues. This approach extracts features by emphasizing the similarities among traces, thereby improving the effectiveness of key recovery while maintaining the advantages of the original SCA approach. Through experiments of different datasets from different platforms, we demonstrate that CL-SCA significantly outperforms other approaches. Moreover, by incorporating attack traces into the training process using our approach, we can further enhance its performance. This extension can improve the effectiveness of key recovery, which is fully verified through experiments on different datasets. Annyu Liu, An Wang 0001, Shaofei Sun, Congming Wei, Yaoling Ding, Yongjuan Wang, Liehuang Zhu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Release the Power of Rejected Signatures: An Efficient Side-Channel Attack on the ML-DSA CryptosystemabstractThe module-lattice-based digital signature standard, formerly known as CRYSTALS-DILITHIUM, is a lattice-based post-quantum cryptographic scheme. In August 2024, the National Institute of Standards and Technology officially standardized ML-DSA under FIPS 204. ML-DSA generates one valid signature and multiple rejected signatures during a single signing process. Most side-channel attacks targeting ML-DSA have focused solely on the valid signature, while largely neglecting the hints contained in rejected signatures. Building on prior SASCA frameworks originally proposed for ML-DSA, in this paper we present an efficient and fully practical instantiation of a private-key recovery attack on ML-DSA that jointly exploits side-channel leakages from both valid and rejected signatures within a unified factor graph. This concrete instantiation maximizes the information extracted from a single signing attempt and minimizes the number of required traces for full key recovery. We conducted a proof-of-concept experiment with both reference and ASM-optimized implementations on a Cortex-M4 core chip, where the results demonstrate that incorporating rejected signatures reduces the required number of traces by at least 50.0% for full key recovery. Moreover, we show that using only rejected signatures suffices to recover the key with fewer than 30 traces under our setup. Our findings highlight that protecting rejected signatures is crucial, as their leakage provides valuable side-channel information. We strongly recommend implementing countermeasures for rejected signatures during the signing process to mitigate potential threats. Zheng Liu 0029, An Wang 0001, Congming Wei, Yaoling Ding, Annyu Liu, Liehuang Zhu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | High-Performance Elliptic Curve Scalar Multiplication Architecture Based on Interleaved MechanismabstractHigh-performance (HP) elliptic curve scalar multiplication (ECSM) hardware implementations hold significant importance in ensuring communication security in high-capacity and high-concurrence application scenarios. By analyzing the inherent priorities and parallelism in ECSMs, we proposed a novel HP ECSM algorithm and a partially parallel inversion algorithm based on the interleaved mechanism. With two dedicated multipliers and one interleaved multiplier, we introduced a compact hardware scheduling scheme to realize the consumption of four clock cycles within each loop of ECSM. The proposed HP ECSM architecture consists of two Karatsuba-Ofman multipliers (KOMs) and one classical multiplier (CM). The multiplexors and pipeline stages are meticulously designed to optimize the critical path (CP). The proposed architecture is implemented over Virtex-7 field-programmable gate array (FPGA), and the throughput reaches 158.03, 138.23, and 117.50 Mbps over$\text {GF}(2^{163})$,$\text {GF}(2^{283})$, and$\text {GF}(2^{571})$using 8762, 20451, and 41974 slices, respectively. The comparisons with recent existing works demonstrate that the performance and throughput of our design are among the top. Zhiming Chen 0001, Mingzhi Ma, Rongkun Jiang, An Wang 0001, Weijiang Wang, Hua Dang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Modular Inversion Architecture over GF(2m) Using Optimal Exponentiation Blocks for ECC CryptosystemsabstractThe inversion over GF(2m) is crucial for elliptic curve cryptography algorithms such as ECDSA and SM2. The Itoh-Tsujii’s Algorithm (ITA) can compute inversions in a sequential procedure by utilizing multiplications and exponentiations. This paper proposes a series of novel low-latency architectures with Cascaded Exponentiation Blocks (CEBs) and then derives the estimated clock cycle latency. The complexity of CEBs is evaluated by the matrix weight. We also employ a movable internal pipeline stage to optimize the critical path. Experiments on the Virtex-7 FPGA show the optimal exponentiation blocks for GF(2163), GF(2283) and GF(2571), respectively. Compared with existing works, both the performance and latency of our proposed architecture with OEBs are at the cutting edge. An Wang 0001 |
ISCAS | 3 |
| 2024 | A closer look at the belief propagation algorithm in side-channel attack on CCA-secure PQC KEM
Kexin Qiao, Heng Chang, Siwei Sun, Zehan Wu, Junjie Cheng, Changhai Ou, An Wang 0001, Liehuang Zhu |
Sci. China Inf. Sci. | 8 |
| 2024 | Bitwise Mixture Differential Cryptanalysis and Its Application to SIMONabstractWith the proliferation of IoT devices today, the need to strengthen the security of these devices is becoming increasingly urgent, particularly the need to review the security of lightweight block ciphers. SIMON is a lightweight block cipher proposed by the National Security Agency (NSA) of US to provide efficient and secure encryption for resource-constrained devices in IoT systems. This paper aims to evaluate the security of SIMON against mixture differential cryptanalysis, which was proposed in Eurocrypt 2017 to launch the best key-recovery attacks on the most widely used encryption standard AES. Though there have been intensive studies on this cryptanalysis method, its current targets are all aligned block ciphers. Whether the numerous bitwise block ciphers, including SIMON, have weaknesses regarding this method remains unknown. In this paper, we extend the mixture differential cryptanalysis to bit-wise ciphers and develop an SAT-based automatic tool to search for such distinguishers. We interpret the bit-wise mixture differential distinguisher as a variant of differential distinguisher in the multi-key setting with 2-3n as the boundary (n:block size), potentially boosting rounds or improving the signal-to-noise ratio of previous boomerang or classical differential distinguisher. Using SIMON as an example, we discover multi-key distinguishers for up to 17-round SIMON32, 18-round SIMON48, and 23-round SIMON64, which outperform previous results in terms of the number of rounds. This paper reconciles the disparity between mixture differential cryptanalysis applied to word-oriented target ciphers and its application to bit-oriented targets, thereby extending the mixture differential cryptanalysis to a broader range of block ciphers. Kexin Qiao, Zehan Wu, Junjie Cheng, Changhai Ou, An Wang 0001, Liehuang Zhu |
IEEE Internet Things J. | 5 |
| 2024 | An efficient heuristic power analysis framework based on hill-climbing algorithm
Shaofei Sun, Shijun Ding, An Wang 0001, Yaoling Ding, Congming Wei, Liehuang Zhu, Yongjuan Wang |
Inf. Sci. | 3 |
| 2024 | Efficient Multi-Byte Power Analysis Architecture Focusing on Bitwise Linear LeakageabstractAs the most commonly used side-channel analysis method, Correlation Power Analysis (CPA) usually uses the divide-and-conquer strategy to guess the single-byte key in the scenario of block cipher parallel implementation. However, this method cannot effectively use the power consumption information, resulting in a large number of power consumption traces. Therefore, genetic algorithm-based CPA is proposed, which can efficiently extract keys by multi-byte power analysis. However, genetic algorithm-based CPA tends to sacrifice computational cost to achieve a high key guessing success rate. To solve the above problems, this article focuses on bitwise linear leakage and proposes a multi-byte power analysis architecture based on the raindrop ripple algorithm. First, we propose to complete the key initialization by multiple linear regression. Second, we propose a novel swarm intelligence algorithm, the raindrop ripple algorithm, tailored for multi-byte power analysis based on the principles of “family planning” and “eugenics,” which greatly improves the probability of producing individuals with high fitness values. Third, we further enhance the possibility of the correct key being recovered by traversing the candidate key space in specific conditions. To verify the key guessing efficiency of the multi-byte power analysis architecture based on the raindrop ripple algorithm, comparative experiments are conducted on SAKURA-G with three power analysis methods based on genetic algorithms. Experimental results show that our proposal not only has the efficient power information utilization of multi-byte power analysis but also has a convergence speed comparable to or even faster than that of single-byte CPA. Its efficiency of key guessing is improved by 85.64% compared to EfficiencyGa-CPA, and its convergence speed is even faster than that of single-byte CPA at 725 power traces, and 83.87% faster than single-byte CPA at 1000 power traces, which is astonishing as a multi-byte power analysis. Zijing Jiang, Qun Ding, An Wang 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Time Is Not Enough: Timing Leakage Analysis on Cryptographic Chips via Plaintext-Ciphertext Correlation in Non-Timing ChannelabstractIn side-channel testing, the standard timing analysis works when the vendor can provide a measurement to indicate the execution time of cryptographic algorithms. In this paper, we find that there exists timing leakage in power/electromagnetic channels, which is often ignored in traditional timing analysis. Hence a new method of timing analysis is proposed to deal with the case where execution time is not available. Different execution time leads to different execution intervals, affecting the locations of plaintext and ciphertext transmission. Our method detects timing leakage by studying changes in plaintext-ciphertext correlation when traces are aligned forward and backward. Experiments are then carried out on different cryptographic devices. Furthermore, we propose an improved timing analysis framework which gives appropriate methods for different scenarios. Congming Wei, Guangze Hong, An Wang 0001, Jing Wang 0150, Shaofei Sun, Yaoling Ding, Liehuang Zhu, Wenrui Ma |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Autoencoder Assist: An Efficient Profiling Attack on High-Dimensional Datasets
Zijia Yang, Qin Wang 0008, Yaoling Ding, An Wang 0001 |
ICICS | 6 |
| 2022 | Correlation leakage analysis based on masking schemes
Yongchuan Niu, An Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2022 | SCARE and power attack on AES-like block ciphers with secret S-box
An Wang 0001, Liehuang Zhu, Yaoling Ding, Zeyuan Lyu, Zongyue Wang |
Frontiers Comput. Sci. | 2 |
| 2022 | Attacking the Edge-of-Things: A Physical Attack PerspectiveabstractThe concepts between Internet of Things (IoT) and edge computing are increasingly intertwined, as an edge-computing architecture generally comprises a (large) number of diverse IoT devices. This, however, increases the potential attack vectors since any one of these connected IoT devices can be targeted to facilitate other malicious cyber activities. Physical attacks are generally harder to mitigate and less studied, in comparison to their cyber counterparts. Thus, in this article we present an attack framework targeting true random number generators (TRNGs), which are a key component in cryptosystems for edge devices. We then demonstrate how such a framework can guide our investigation of a commercial ASIC chip that runs ring-oscillator-based TRNG. Specifically, we show that our template power attack, low voltage fault attack, and voltage glitch fault attack do not require prior knowledge of the TRNG implementation. Keke Gai, Yaoling Ding, An Wang 0001, Liehuang Zhu, Kim-Kwang Raymond Choo, Qi Zhang 0010, Zhuping Wang |
IEEE Internet Things J. | 3 |
| 2021 | Efficient Framework for Genetic Algorithm-Based Correlation Power AnalysisabstractVarious Artificial Intelligence (AI) techniques are combined with classic side-channel methods to improve the efficiency of attacks. Among them, Genetic-Algorithms-based Correlation Power Analysis (GA-CPA) is proposed to launch attacks on hardware cryptosystems to extract the secret key efficiently. However, the convergence efficiency of GA-CPA is unsatisfactory due to two problems: the randomly generated initial population generally have low fitness, and the mutation operation in each iteration hardly produces high-quality individuals because of the confusion and diffusion characteristics of S-boxes. In this paper, we propose an analysis framework of GA-CPA which focuses on solving these two problems. First, we explore the list of candidate key bytes which is the result of Correlation Power Analysis (CPA) on a limited number of power traces, so that the population can be initialized with high quality candidates. Second, we improve the mutation operation by guiding the candidate key to mutate in a higher-fitness direction instead of randomly. Third, we make full use of the fitness calculation method and combine it with key enumeration algorithms to further improve the efficiency of key recovery. Simulation experimental results show that our method reduces the number of traces by 33.3% and 43.9% compared to CPA with key enumeration and GA-CPA respectively when the success rate is fixed to 90%. Real experiments performed on SAKURA-G confirm that the number of traces required in our method is much less than the numbers of traces required in CPA and GA-CPA. Besides, we adjust our method to deal with DPA contest v1 dataset, and achieve a better result of 40.76 traces than the winning proposal of 42.42 traces. The computation cost of our proposal is nearly 16.7% of the winner. An Wang 0001, Yaoling Ding, Liehuang Zhu, Yongjuan Wang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | A Multiple Sieve Approach Based on Artificial Intelligent Techniques and Correlation Power AnalysisabstractSide-channel analysis achieves key recovery by analyzing physical signals generated during the operation of cryptographic devices. Power consumption is one kind of these signals and can be regarded as a multimedia form. In recent years, many artificial intelligence technologies have been combined with classical side-channel analysis methods to improve the efficiency and accuracy. A simple genetic algorithm was employed in Correlation Power Analysis (CPA) when apply to cryptographic algorithms implemented in parallel. However, premature convergence caused failure in recovering the whole key, especially when plenty of large S-boxes were employed in the target primitive, such as in the case of AES. In this article, we investigate the reason of premature convergence and propose a Multiple Sieve Method (MS-CPA), which overcomes this problem and reduces the number of traces required in correlation power analysis. Our method can be adjusted to combine with key enumeration algorithms and further improves the efficiency. Simulation experimental results depict that our method reduces the required number of traces by and , compared to classic CPA and the Simple-Genetic-Algorithm-based CPA (SGA-CPA), respectively, when the success rate is fixed to . Real experiments performed on SAKURA-G confirm that the number of traces required for recovering the correct key in our method is almost equal to the minimum number that makes the correlation coefficients of correct keys stand out from the wrong ones and is much less than the numbers of traces required in CPA and SGA-CPA. When combining with key enumeration algorithms, our method has better performance. For the traces number being 200 (noise standard deviation ), the attacks success rate of our method is , which is much higher than the classic CPA with key enumeration ( success rate). Moreover, we adjust our method to work on that DPA contest v1 dataset and achieve a better result (40.04 traces) than the winning proposal (42.42 traces). Yaoling Ding, Liehuang Zhu, An Wang 0001, Yongjuan Wang, Siu-Ming Yiu, Keke Gai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Block-oriented correlation power analysis with bitwise linear leakage: An artificial intelligence approach based on genetic algorithms
Yaoling Ding, An Wang 0001, Yongjuan Wang, Guoshuang Zhang |
Future Gener. Comput. Syst. | 3 |
| 2020 | x-only coordinate: with application to secp256k1 " >Chosen base-point side-channel attack on Montgomery ladder with x-only coordinate: with application to secp256k1abstractThis study revisits the side‐channel security of the elliptic curve cryptography (ECC) scalar multiplication implemented with Montgomery ladder. Focusing on a specific implementation that does not use the y ‐coordinate for point addition (ECADD) and point doubling (ECDBL), the authors show that Montgomery ladder on Weierstrass curves is vulnerable to a chosen base‐point attack. Unlike the normal implementation with y ‐coordinate, in the scenario of this study, the chosen base‐point strategy will not lead to operations with two same inputs during the ECADD and/or ECDBL. Instead, by choosing a suitable base‐point, one will find that there are operations that share a common operand; while it is not the case if the base‐point is not chosen correctly. This results in the recovery of the secret (fixed) scalar. They also experiment the methods of shared operand detection on a real‐world SoC, where a secp256k1 dedicated Montgomery ladder scalar multiplication with x ‐only coordinate is implemented, to show the efficiency of the scalar recovery attack. Naturally, the attack can be generalised to other Weierstrass curves when they contain special points. Congming Wei, Jiazhe Chen, An Wang 0001, Hongsong Shi, Xiaoyun Wang 0001 |
IET Inf. Secur. | 3 |
| 2020 | A machine learning based golden-free detection method for command-activated hardware Trojan
Ning Shang 0001, An Wang 0001, Yaoling Ding, Keke Gai, Liehuang Zhu, Guoshuang Zhang |
Inf. Sci. | 2 |
| 2019 | Power Analysis and Protection on SPECK and Its Application in IoT
An Wang 0001, Liehuang Zhu, Ning Shang 0001, Guoshuang Zhang |
SecureComm (2) | 2 |
| 2019 | New second-order threshold implementation of AESabstractIn this work, the authors propose some alternative hardware efficient masking schemes dedicated to protect the Advanced Encryption Standard (AES) against higher order differential power analysis (DPA). In general, the existing masking schemes all have in common an intrinsic trade‐off between the two main parameters of interest, namely the generation of fresh random masking values and the cost of hardware implementation. The design of efficient masking schemes which are non‐expensive in both aspects appears to be a difficult task. In this study, the authors propose a second‐order threshold implementation of AES, which is characterised by a beneficial trade‐off between the two parameters. More precisely, compared to the masking scheme of De Cnudde et al . at CHES 2016, which currently attains the best practical trade‐off, the proposed masking scheme requires 28.4% less random masking bits, whereas the implementation cost is slightly increased for about 13.7% (thus the chip area is 1.4 kGE larger). This masking scheme has been used to implement AES on an field‐programmable gate array (FPGA) platform and its resistance against the second‐order DPA in a simulated attack environment has been confirmed. Yongzhuang Wei, Fu Yao, Enes Pasalic, An Wang 0001 |
IET Inf. Secur. | 4 |
| 2018 | Right or wrong collision rate analysis without profiling: full-automatic collision fault attack
An Wang 0001, Weina Tian, Qian Wang 0022, Guoshuang Zhang, Liehuang Zhu |
Sci. China Inf. Sci. | 1 |
| 2018 | Analysis of Software Implemented Low Entropy Masking SchemesabstractLow Entropy Masking Schemes (LEMS) are countermeasure techniques to mitigate the high performance overhead of masked hardware and software implementations of symmetric block ciphers by reducing the entropy of the mask sets. The security of LEMS depends on the choice of the mask sets. Previous research mainly focused on searching balanced mask sets for hardware implementations. In this paper, we find that those balanced mask sets may have vulnerabilities in terms of absolute difference when applied in software implemented LEMS. The experiments verify that such vulnerabilities certainly make the software LEMS implementations insecure. To fix the vulnerabilities, we present a selection criterion to choose the mask sets. When some feasible mask sets are already picked out by certain searching algorithms, our selection criterion could be a reference factor to help decide on a more secure one for software LEMS. Jiazhe Chen, An Wang 0001, Xiaoyun Wang 0001 |
Secur. Commun. Networks | 3 |
| 2018 | Side-Channel Attacks and Countermeasures for Identity-Based Cryptographic Algorithm SM9abstractIdentity-based cryptographic algorithm SM9, which has become the main part of the ISO/IEC 14888-3/AMD1 standard in November 2017, employs the identities of users to generate public-private key pairs. Without the support of digital certificate, it has been applied for cloud computing, cyber-physical system, Internet of Things, and so on. In this paper, the implementation of SM9 algorithm and its Simple Power Attack (SPA) are discussed. Then, we present template attack and fault attack on SPA-resistant SM9. Our experiments have proved that if attackers try the template attack on an 8-bit microcontrol unit, the secret key can be revealed by enabling the device to execute one time. Fault attack even allows the attackers to obtain the 256-bit key of SM9 by performing the algorithm twice and analyzing the two different results. Accordingly, some countermeasures to resist the three kinds of attacks above are given. Qi Zhang 0010, An Wang 0001, Yongchuan Niu, Ning Shang 0001, Rixin Xu, Guoshuang Zhang, Liehuang Zhu |
Secur. Commun. Networks | 2 |
| 2017 | Practical two-dimensional correlation power analysis and its backward fault-tolerance
An Wang 0001, Wenjing Hu, Weina Tian, Guoshuang Zhang, Liehuang Zhu |
Sci. China Inf. Sci. | 1 |
| 2017 | RFA: R-Squared Fitting Analysis Model for Power AttackabstractCorrelation Power Analysis (CPA) introduced by Brier et al. in 2004 is an important method in the side-channel attack and it enables the attacker to use less cost to derive secret or private keys with efficiency over the last decade. In this paper, we propose R -squared fitting model analysis (RFA) which is more appropriate for nonlinear correlation analysis. This model can also be applied to other side-channel methods such as second-order CPA and collision-correlation power attack. Our experiments show that the RFA-based attacks bring significant advantages in both time complexity and success rate. An Wang 0001, Liehuang Zhu, Weina Tian, Rixin Xu, Guoshuang Zhang |
Secur. Commun. Networks | 1 |
| 2017 | Differential Fault Attack on ITUbee Block CipherabstractDifferential Fault Attack (DFA) is a powerful cryptanalytic technique to retrieve secret keys by exploiting the faulty ciphertexts generated during encryption procedure. This article proposes a novel DFA attack that is effective on ITUbee, a software-oriented block cipher for resource-constrained devices. Different from other DFA, our attack makes use of not only faulty values, but also differences between fault-free intermediate values corresponding to 2 plaintexts, which combine traditional differential analysis with DFA. The possible injection positions with different number of faults are discussed. The most efficient attack takes 2 25 round function operations with 4 faults, which is achieved in a few seconds on a PC. Shan Fu, Guoai Xu, Juan Pan, Zongyue Wang, An Wang 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2015 | Transient-Steady Effect Attack on Block Ciphers
Yanting Ren, An Wang 0001, Liji Wu |
CHES | 2 |
| 2015 | Efficient collision attacks on smart card implementations of masked AES
An Wang 0001, Zongyue Wang, Xuexin Zheng, Guoshuang Zhang, Liji Wu |
Sci. China Inf. Sci. | 1 |
| 2015 | A novel bit scalable leakage model based on genetic algorithmabstractAbstract With the growing popularity of smart integrated circuit (IC) cards, the chip security is attracting more and more attention. Researches on the attack and protection of smart IC cards have become increasingly hot. Side‐channel attack is the practical and effective method, which has brought enormous threat. The efficiency of attack depends on the extent of the leakage model, which characterizes the practical applications. In the power analysis attack, the classical leakage model usually exploits the power consumption of single S‐box, which is called divide and conquer. Taking data encryption standard (DES) algorithm, for example, the attack on each S‐box needs to search the key space of 2 6 in a brute‐force way. In this paper, we propose a novel leakage model, which is more flexible than the classical leakage model. The novel leakage model is based on the power consumption of multiple S‐boxes, and the implementation of this method is combined with genetic algorithm. We can establish leakage model based on the Hamming distance of round output generated by eight S‐boxes in DES algorithm. The experiment verifies the fact that the leakage model of eight S‐boxes can decrease the traces number up to 52% than the classical one based on single S‐box for DES algorithm. It also decreases the traces number up to 32% for SM4 algorithm. All the measurements of power data are acquired from a practical smart IC card. We also conclude that increasing noise, using variable clock, and limiting the lifetime of root key can be the choices of defensive strategy. Copyright © 2015 John Wiley & Sons, Ltd. Zhenbin Zhang, Liji Wu, An Wang 0001, Zhaoli Mu, Xiangmin Zhang |
Secur. Commun. Networks | 3 |
| 2012 | Overcoming Significant Noise: Correlation-Template-Induction Attack
An Wang 0001, Zongyue Wang, Yaoling Ding |
ISPEC | 1 |