EDBT 2026 Demo / reviewers in the wild / expert
Zeming Cheng
dblp:298/7833
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-4742-9600ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flex-NTT: Design of a Flexible and Compact Number Theoretic Transform Architecture for Homomorphic Encryption ApplicationsabstractThis article presents Flex-NTT, a flexible (configurable) and area-efficient number theoretic transform (NTT) architecture featuring novel unified memory access patterns. The proposed design offers compile-time configurability (CTC) to support various parallel butterfly units (BUs) and run-time configurability (RTC) to accommodate diverse NTT operation sizes without recompilation. A single hardware instance is reconfigurable for NTT, inverse NTT (INTT), and point-wise multiplication, improving hardware utilization. The proposed coefficient and twiddle-factor memory access schemes achieve the theoretical minimum memory sizes, eliminate intermediate buffers, and enable natural-order outputs without additional reordering. FPGA evaluations demonstrate that Flex-NTT achieves up to$4.25 \times $BRAM savings and$2.82 \times $performance improvements over prior NTT/INTT designs. When applied to polynomial multiplication, it delivers up to$2.35 \times $performance and$2.93 \times $BRAM utilization improvements under the same metric. Zeming Cheng, Xiao Tuo, Mingye Li, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | MCCMS: Achieve fine-grained phase distribution design in cement microstructure using diffusion models
Lin Wang 0004, Shuangrong Liu, Haozhong Gao, Zeming Cheng, Chaoran Pang, Bo Yang 0001 |
Comput. Aided Des. | 6 |
| 2024 | A High-Performance, Conflict-Free Memory-Access Architecture for Modular Polynomial MultiplicationabstractIn this article, we present the HiCoP architecture, a high-performance, conflict-free memory access, modular polynomial multiplication design that accelerates the number-theoretic transform (NTT), inverse NTT (INTT), and modular polynomial multiplications. To optimize hardware costs, the HiCoP architecture utilizes a high-radix reconfigurable butterfly unit (RBU) that can be dynamically configured to perform NTT, INTT, and point-wise multiplications, alongside an area-efficient Montgomery modular multiplier (MMM) tailored for NTT-friendly modulus. Moreover, by integrating pre-processing, post-processing, and Montgomery domain transformations into NTT and INTT operations, we effectively minimize the cycle count for modular polynomial multiplication. Additionally, we propose a novel conflict-free memory access algorithm that simplifies the control logic and eliminates the need for ping-pong memory in the HiCoP architecture. Experimental results of modular polynomial multiplications demonstrate significant performance gains for the HiCoP architecture implemented on the Xilinx Virtex-7 field-programmable gate array (FPGA) platform, with up to$8.75\times $,$4.15\times $,$10.57\times $, and$8.50\times $improvements in throughput-to-hardware-cost ratio for LUT count, FF count, BRAM count, and DSP count, respectively. Zeming Cheng, Bo Zhang 0098, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Design of a High-Performance Iterative Barrett Modular Multiplier for Crypto SystemsabstractModular multiplication (MM) is a fundamental operation in many cryptographic and arithmetic applications. In this article, we present an improved Barrett modular multiplication (BMM) algorithm and its hardware-efficient implementation. The proposed algorithm leverages parallel computation of quotient and intermediate results, enhancing overall efficiency. To further optimize the algorithm, two optimizations are introduced, replacing expensive multiplications and additions with more efficient compression and encoding operations at each iteration. We first introduce a novel data model that enables the use of a 2-bit adder to handle potential overflow in signed addition. Moreover, by employing a 3-bit addition on intermediate results, we eliminate the need for complete round operations while ensuring the desired result range. The experimental results demonstrate significant improvements in terms of area and computation time compared to existing classic BMM and Montgomery modular multiplication (MMM) designs. Our improved BMM outperforms these designs, particularly in high-radix scenarios. This work provides a valuable contribution to the field of MM, offering a hardware-efficient solution for achieving improved performance in cryptographic and arithmetic systems. Bo Zhang 0098, Zeming Cheng, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | An Iterative Montgomery Modular Multiplication Algorithm With Low Area-Time ProductabstractThis paper presents a highly efficient iterative Montgomery modular multiplication algorithm, wherein the computations of quotient and intermediate result in each iteration are done in parallel. This parallelism breaks the data dependency and thus reduces the computation latency. Moreover, this paper replaces required multiplications and additions in each iteration with compressions and encoding, thereby achieving a computation latency of order$d+6$where$d=\left\lceil N/m \right\rceil +2$is the number of iterations,$N$denotes the bitwidth of modulus$M$, and$m$is the number of bits of the multiplier that are processed in each iteration of the algorithm. Hardware realization of the proposed Montgomery modular multiplication on a Xilinx Virtex-7 FPGA device shows$> 41\%$computation latency saving and$>31\%$area saving when$N=1,024$and$m=8$, compared with the best of previous state-of-art references. These savings amount to more than 63% reduction in terms of the area-latency product metric. Bo Zhang 0098, Zeming Cheng, Massoud Pedram |
IEEE Trans. Computers | 2 |
| 2022 | High-Radix Design of a Scalable Montgomery Modular Multiplier With Low LatencyabstractThe proposed herein is a scalable high-radix (i.e.,$2^m$2m) Montgomery Modular (MM) Multiplication circuit replacing the integer multiplications in each iteration of the Montgomery MM algorithm (related to the product of$m$mbits of the multiplier and the multiplicand) with carry-save compressions and completely eliminating costly multiplications. Furthermore, the proposed Montgomery MM decomposes the multiplicand itself using a radix of$2^w$2wwith$w\geq 2m$w≥2m, thereby achieving a scalable design, which can deliver an issue latency of one cycle and a cycle (count) latency of$O(N^2/(wmp))$O(N2/(wmp))where$p$pdenotes the number of available processing elements, each of which is designed to complete the above iteration by computing in part the product of$w$wbits of the multiplicand and$m$mbits of the multiplier. The area complexity of the proposed Montgomery MM is$O(wmp)$O(wmp), and thus, the Area-Latency-Product complexity is$O(N^{2})$O(N2). Bo Zhang 0098, Zeming Cheng, Massoud Pedram |
IEEE Trans. Computers | 2 |
| 2021 | A High-Performance Low-Power Barrett Modular Multiplier for CryptosystemsabstractThis paper presents a fast architecture for Barrett modular multiplication. By replacing the integer multiplications in each iteration with carry-save compressions and using Booth coding plus operation rescheduling to increase parallelism, we eliminate costly multiplications while concurrently avoiding large-bitwidth additions. Our detailed error analysis proves that intermediate results are always less than twice the modulus. Experimental results show that the removal of multiplication eliminates the need for any DSPs. Even not accounting for this key benefit, compared to the best of prior art results, the proposed design results in 46.8% latency reduction with a similar area. Bo Zhang 0098, Zeming Cheng, Massoud Pedram |
ISLPED | 2 |