Cheng Chen 0076

dblp:10/217-76 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-5733-4528ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 FlexMSM: A Flexible FPGA-Based Accelerator for Multi-Scalar Multiplication with Reconfigurable Modular Arithmetic and Optimized Pippenger Scheduling
abstract
Zero-Knowledge Proofs (ZKPs), especially zk-SNARKs, rely heavily on Multi-Scalar Multiplication (MSM), a compute-intensive elliptic curve operation. While prior work targets curves with optimized operations like BLS12-377, MSM on general-purpose curves such as BLS12-381 remains challenging due to imbalanced resource usage, performance gaps between curve operations, and low utilization of point addition. This paper proposes FlexMSM to support scalable MSM cores on a single FPGA for BLS12-381 curve, delivering significant gains over existing works for input sizes from 218 to 226 .
Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Zhiguo Wan
FPGA1
2026 High-Performance Accelerator for Constant-Time Cross-Domain Integer and Montgomery Inversion on FPGA
abstract
Modular Inversion (MI) is one of the fundamental arithmetic operations in the finite field, which plays an essential role in various cryptographic applications and requires high performance and security. Unfortunately, the simple MI algorithm is vulnerable to side-channel attacks, such as the timing attack, which can compromise the cryptographic system by analyzing the time taken to execute cryptographic algorithms. Attackers may recover the initial data since the time can differ based on the input. Besides, the low complexity and low resource consumption of hardware implementations in MI are also challenging. In this article, we propose two novel modular inversion algorithms, named Constant-Time Integer Modular Inversion (CT-IMI) and Constant-Time Complementary Montgomery Modular Inversion (CT-CMMI). They both consist of constant iteration rounds to resist the timing attack. CT-IMI processes the data in the integer field, which is designed for common scenarios. CT-CMMI is suitable for the cross-domain case, which can directly use data in the Montgomery domain and avoid the conversion steps for some specific applications, e.g., scalar multiplication in Elliptic Curve Cryptography (ECC). In software simulations, we measure the average clock cycles for a single inversion and illustrate the relationship between various bit lengths and the latency. The significant differences between constant and non-constant algorithms demonstrate the vulnerability of modular inversion to timing attacks. In addition, we design two efficient hardware architectures on FPGA. Experimental results show that our CT-IMI can finish a single inversion in 2.56 \(\mu\) s with 4.2k LUTs, 1.8k FFs, and our CT-CMMI requires 2.45 \(\mu\) s with 2.7k LUTs, 1.6k FFs. The product of area and latency of our CT-IMI and CT-CMMI can reach 10.50 and 6.62, respectively, which shows optimal performance compared with all the results in the existing literature.
Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Xianye Ben, Zhiguo Wan
ACM Trans. Embed. Comput. Syst.2
2025 Customized FPGA Implementation of Authenticated Lightweight Cipher Fountain for IoT Systems
abstract
Authenticated Encryption with Associated-Data (AEAD) can ensure both confidentiality and integrity of information in encrypted communication. Distinctive variants are customized from AEAD to satisfy various requirements. In this paper, we take a 128-bit lightweight AEAD stream cipher Fountain as an example. We provide a general cryptographic solution with three Fountain variants. These three variants are for encryption, message authentication code (MAC) generation, and authenticated encryption with associated data, respectively. Besides, we propose area-saved and throughput-improved strategies for the FPGA implementation of Fountain. The conventional paralleled hardware implementation leads to much resource-consuming with higher parallel width. We propose a hybrid architecture with parallel and serial update modes simultaneously. We also analyze the trade-off between area occupation and authentication latency for those two architectures. According to our discussion, hybrid architectures can perform efficiently with higher throughput than most ciphers, including Grain-128 x32. Our Fountain keystream generator occupies 46 slices on Spartan-3 FPGAs, smaller than most ciphers with the same security level, and even smaller than the 80-bit security level cipher Trivium. In summary, the customized Fountain with optimized implementations on FPGA is suitable for various applications in the field of IoT.
Zhengyuan Shi, Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Zhiguo Wan
ACM Trans. Embed. Comput. Syst.2
2023 Design Space Exploration of Galois and Fibonacci Configuration Based on Espresso Stream Cipher
abstract
Fibonacci and Galois are two different kinds of configurations in stream ciphers. Although many transformations between two configurations have been proposed, there is no sufficient analysis of their FPGA performance. Espresso stream cipher provides an ideal sample to explore such a problem. The 128-bit secret key Espresso is designed in Galois configuration, and there is a Fibonacci-configured Espresso variant proved with the equivalent security level. To fully leverage the efficiency of two configurations, we explore the hardware optimization approaches toward area and throughput, respectively. In short, the FPGA-implemented Fibonacci cipher is more suitable for extremely resource-constrained or high-throughput applications, while the Galois cipher compromises both area and speed. To the best of our knowledge, this is the first work to systematically compare the FPGA performance of cipher configurations under relatively fair cryptographic security. We hope this work can serve as a reference for the cryptography hardware architecture research community.
Zhengyuan Shi, Cheng Chen 0076, Gangqiang Yang, Hailiang Xiong, Fudong Li 0002, Honggang Hu, Zhiguo Wan
ACM Trans. Reconfigurable Technol. Syst.2
2023 Hardware Optimizations of Fruit-80 Stream Cipher: Smaller than Grain
abstract
Fruit-80, which emerged as an ultra-lightweight stream cipher with 80-bit secret key, is oriented toward resource-constrained devices in the Internet of Things. In this article, we propose area and speed optimization architectures of Fruit-80 on FPGAs. Our implementations include both serial and parallel structure and optimize area, power, speed, and throughput, respectively. The area optimization architecture aims to achieve the most suitable ratio of look-up-tables and flip-flops to fully utilize the reconfigurable unit. It also reuses NFSR and LFSR feedback functions to save resources for high throughput. The speed optimization architecture adopts a hybrid approach for parallelization and reduces the latency of long data paths by pre-generating primary feedback and inserting flip-flops. Besides, we recommend using the round key function to optimize serial or parallel implementations for Fruit-80 and using indexing and shifting methods for different throughput. In conclusion, our results show that the area optimization architecture occupies up to 35 slices on Xilinx Spartan-3 FPGA and 18 slices on Xilinx 7 series FPGA, smaller than that of Grain and other common stream ciphers. The optimal throughput/area ratio of the speed optimization architecture is 7.74 Mbps/slice, better than that of Grain v1, which is 5.98 Mbps/slice. The serial implementation of Fruit-80 with round key function occupies only 75 slices on Spartan-3 FPGA. To the best of our knowledge, the result sets a new record of the minimum area in lightweight cipher implementation on FPGA.
Gangqiang Yang, Zhengyuan Shi, Cheng Chen 0076, Hailiang Xiong, Fudong Li 0002, Honggang Hu, Zhiguo Wan
ACM Trans. Reconfigurable Technol. Syst.3
2022 Work-in-Progress: Towards a Smaller than Grain Stream Cipher: Optimized FPGA Implementations of Fruit-80
abstract
Fruit-80, an ultra-lightweight stream cipher with 80-bit secret key, is oriented toward resource constrained devices in the Internet of Things. In this paper, we propose area and speed optimization architectures of Fruit-80 on FPGAs. The area optimization architecture reuses NFSR&LFSR feedback functions and achieves the most suitable ratio of look-up-tables and flip-flops. The speed optimization architecture adopts a hybrid approach for parallelization and reduces the latency of long data paths by pre-generating primary feedback and inserting flip-flops. In conclusion, the optimal throughput-to-area ratio of the speed optimization architecture is better than that of Grain v1. The area optimization architecture occupies only 35 slices on Xilinx Spartan-3 FPGA, smaller than that of Grain and other common stream ciphers. To the best of our knowledge, this result sets a new record of the minimum area in lightweight cipher implementations on FPGA.
Gangqiang Yang, Zhengyuan Shi, Cheng Chen 0076, Hailiang Xiong, Honggang Hu, Zhiguo Wan, Keke Gai, Meikang Qiu
CASES3