Zhiming Chen 0001

dblp:25/2326-1 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-9195-1327ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HTMMM: Novel Hybrid Truncated Montgomery Modular Multiplication Algorithm and Hardware Architecture
abstract
Modular multiplication is one of the key operations in modern public-key cryptography. Montgomery Modular Multiplication (MMM) is a mainstream method to avoid modulo operations, which contains one variable multiplication and two constant multiplications. In this paper, for high performance, a novel Hybrid Truncated Montgomery Modular Multiplication (HTMMM) algorithm and its hardware architecture are proposed, which achieves state-of-the-art Area-Time-Product (ATP) and throughput. We propose an error-free high-part truncated multiplication for the first time, which solves the problem that conventional methods cannot be applied to MMM due to the introduced error, and reduces the complexity to the same level as low-part truncated multiplication. Besides, a hybrid multiplication based on Toom-Cook and Karatsuba is proposed to optimize variable multiplication, Non-Adjacent Form (NAF) encoding is adopted with truncated multiplication to optimize constant multiplications. The quantitative analysis of complexity for the integer multipliers with different schemes are illustrated to find the optimal multiplier under various cases. Based on these, we took the bit widthN= 1024 as an example to introduce the hardware architecture in detail and gave the implementation results ofN= 256 andN= 1024 in different processes. The experimental results demonstrate that compared with the best existing design, the throughput and ATP of our proposed design are improved by 1.25× and 2.22×, respectively.
Zeying Li, Yue Hao 0007, Hongshuo Li, An Wang 0001, Zhiming Chen 0001, Liehuang Zhu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 Low-Latency and Area-Efficient Elliptic Curve Point Multiplication Architectures Over Koblitz Curves
abstract
Point multiplication is the core operation in elliptic curve cryptography. Koblitz curves are a special class of curves that can utilize the Frobenius mapping to accelerate the implementation of point multiplication operations. For point multiplication on Koblitz curves, this paper first proposes an optimized tNAF scalar conversion algorithm along with its corresponding hardware architecture. Additionally, for the computation of point multiplication, this paper proposes two optimal computational architectures: an area-efficient architecture and a low-latency architecture, both of which achieve the highest pipeline efficiency. The area-efficient architecture adopts a compact four-stage pipeline with a single multiplier, ensuring high circuit area utilization efficiency while achieving relatively low computation latency. The low-latency architecture implements two-stage and three-stage pipeline designs respectively in different binary fields, using two multipliers to reduce the clock cycles for point addition and further decrease the computation latency. The proposed architectures were implemented on Virtex-7 FPGA. For the GF(2163), GF(2283), and GF(2571) fields, the latency for the area-efficient architecture are 1.683μs, 3.455μs and 7.511μs, with slices usage of 3631, 7867 and 20612, and the point multiplication latency for the low-latency architecture are 1.347μs, 3.279μs and 7.071μs, with slices usage of 6026, 14246 and 38515. A comparison with state-of-the-art designs shows that the proposed point multiplication architectures offer significant advantages in terms of performance. In the GF(2163) field, the computation latency of the area-efficient architecture and the low-latency architecture is reduced by at least 21.158% and 43.952%, respectively. And in the GF(2283) field, the reduction in latency is 38.985% and 42.459%, while in the GF(2571) field, the reduction in latency is 58.111% and 59.720%.
An Wang 0001, Yue Hao 0007, Zhiming Chen 0001, Liehuang Zhu
IEEE Internet Things J.6
2025 High-Performance Elliptic Curve Scalar Multiplication Architecture Based on Interleaved Mechanism
abstract
High-performance (HP) elliptic curve scalar multiplication (ECSM) hardware implementations hold significant importance in ensuring communication security in high-capacity and high-concurrence application scenarios. By analyzing the inherent priorities and parallelism in ECSMs, we proposed a novel HP ECSM algorithm and a partially parallel inversion algorithm based on the interleaved mechanism. With two dedicated multipliers and one interleaved multiplier, we introduced a compact hardware scheduling scheme to realize the consumption of four clock cycles within each loop of ECSM. The proposed HP ECSM architecture consists of two Karatsuba-Ofman multipliers (KOMs) and one classical multiplier (CM). The multiplexors and pipeline stages are meticulously designed to optimize the critical path (CP). The proposed architecture is implemented over Virtex-7 field-programmable gate array (FPGA), and the throughput reaches 158.03, 138.23, and 117.50 Mbps over$\text {GF}(2^{163})$,$\text {GF}(2^{283})$, and$\text {GF}(2^{571})$using 8762, 20451, and 41974 slices, respectively. The comparisons with recent existing works demonstrate that the performance and throughput of our design are among the top.
Zhiming Chen 0001, Mingzhi Ma, Rongkun Jiang, An Wang 0001, Weijiang Wang, Hua Dang
IEEE Trans. Very Large Scale Integr. Syst.2
2024 High-Performance ECC Scalar Multiplication Architecture Based on Comb Method and Low-Latency Window Recoding Algorithm
abstract
Elliptic curve scalar multiplication (ECSM) is the essential operation in elliptic curve cryptography (ECC) for achieving high performance and security. We introduce a novel high-performance ECSM architecture over binary fields to meet the growing demand for performance and security. A low-latency window (LLW) recoding algorithm for hardware implementation is proposed to enhance the resistance toward side-channel attacks (SCAs). Based on the LLW algorithm, we propose an enhanced comb method for ECSM with a unified point addition (PA) and point doubling (PD) pattern. The theoretical analysis demonstrates that the enhanced comb method with$w=4$strikes the balance of computation burden for both extreme cases. To achieve short clock cycle latency and high frequency, the data dependency of ECSM is thoroughly analyzed, and we explore a timing schedule with one two-stage pipelined Karatsuba multiplier accumulator (MAC). The datapath of the proposed architecture is well-designed, ensuring that the critical path (CP) only contains minimal logic primitives apart from the MAC. Besides, the ideal placement of pipeline stages for MAC is illustrated. The proposed architecture has been implemented on Xilinx Virtex-7 series field-programmable gate arrays (FPGAs) and performs ECSM in 2.51, 4.93, and$10.85 ~\mu \text { s}$with 3422, 7983, and 20158 slices over$\text {GF}(2^{163})$,$\text {GF}(2^{283})$, and$\text {GF}(2^{571})$, respectively. Implementation results reveal that our design shows 53.60%, 39.36%, and 32.64% performance improvement over the existing state-of-the-art works, respectively.
Zhiming Chen 0001, Mingzhi Ma, Rongkun Jiang, Hongshuo Li, Weijiang Wang
IEEE Trans. Very Large Scale Integr. Syst.2
2023 Deep-Reinforcement-Learning-Based NOMA-Aided Slotted ALOHA for LEO Satellite IoT Networks
abstract
The low earth orbit (LEO) satellites have received extensive attention as an essential supplement to the terrestrial network for supporting global Internet of Things (IoT) services. Considering the rapid growth of IoT devices and the significant satellite-to-ground latency, proposing low-latency, low-overhead access protocols for LEO satellite IoT systems is challenging. In this article, we propose a multibeam random access (RA) framework and deploy the deep reinforcement learning (DRL) algorithm to control the nonorthogonal multiple access (NOMA) aided RA strategy. First, we divide the satellite coverage region into multiple beams and assume that the adjacent beams share parts of regions. Hence, the devices in the sharing region are allowed to transmit packets in two periods allocated for the two beams. Then, packets in multiple beams can be decoded jointly by an interslot successive interference cancelation (SIC) decoder. In addition, we consider the heterogeneity among devices and assign different power levels for heterogeneous types of devices, which enables power-domain NOMA and the intraslot SIC decoder in this system to mitigate the collision resolution. To maximize the average throughput, the deep deterministic policy gradient (DDPG) algorithm is adopted to achieve an online decision to optimize the RA protocol where the packet repetition strategies of devices are adjusted dynamically. The simulation results show that the proposed scheme outperforms the traditional benchmark schemes with significant throughput gain.
Hanxiao Yu, Zesong Fei, Jing Wang 0037, Zhiming Chen 0001, Yuping Gong
IEEE Internet Things J.5
2022 A Ka-band calibratable phased-array front-end chip with high element-consistency
Shiyan Sun, An'an Li, Yingtao Ding, Sijia Jiang, Zhiming Chen 0001, Baoyong Chi
Sci. China Inf. Sci.7