Hao Cheng 0009

dblp:09/5158-9 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0002-4539-3034ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hardware masking with buffer chain
abstract
Abstract Side-channel attacks pose a major threat to cryptographic implementations, as they can exploit physical leakages to recover secret information. Masking is one of the most widely adopted countermeasures, aiming to protect sensitive intermediate values by randomization. However, when deployed in hardware, masking faces the challenges from the glitch leakage, which can easily undermine the security integrity of masking techniques and affect the foundational independent assumptions upon which they are based. To address the intricacies of hardware implementation, circuit separation using registers has emerged as a straightforward method. In this paper, we investigate low-latency hardware masking by exploring the use of buffers (rather than registers) to prevent glitch propagation. Rather than directly inserting buffers into the circuit path, our approach involves employing a chain of buffers to generate signals that serve as controls, thereby synchronizing blocks that require sequential computation. This significantly reduces the power consumption of the shielding circuit while also decreasing latency within the circuit.
Guofeng Qin, Chun Guo 0002, Hao Cheng 0009, Weijia Wang 0003
Cybersecur.5
2025 High-Throughput EdDSA Verification on Intel Processors with Advanced Vector Extensions
Hao Cheng 0009, Johann Großschädl, Peter Y. A. Ryan
SAC2
2024 RISC-V Instruction Set Extensions for Multi-Precision Integer Arithmetic: A Case Study on Post-Quantum Key Exchange Using CSIDH-512
abstract
Multi-Precision Integer (MPI) arithmetic is a performance-critical component of many public-key cryptosystems, including besides classical ones (e.g., RSA, ECC) also isogeny-based post-quantum schemes. In this paper, we analyze and compare two widely-used MPI representations, namely full-radix and reduced-radix, for the efficient implementation of modular arithmetic operations on the 64-bit RISC-V (RV64GC) architecture. We also evaluate how the execution times of both can be further improved with Instruction Set Extensions (ISEs). The ISEs we propose are able to accelerate a CSIDH-512 class group action by a factor of 1.71 compared to a standard software implementation on a 64-bit Rocket core. This speed-up comes at the cost of a hardware overhead of about 10%.
Hao Cheng 0009, Georgios Fotiadis, Johann Großschädl, Dan Page, Thinh Hung Pham, Peter Y. A. Ryan
DAC1
2021 AVRNTRU: Lightweight NTRU-based Post-Quantum Cryptography for 8-bit AVR Microcontrollers
abstract
Introduced in 1996, NTRUEncrypt is not only one of the earliest but also one of the most scrutinized lattice-based cryptosystems and expected to remain secure in the upcoming era of quantum computing. Furthermore, NTRUEncrypt offers some efficiency benefits over “pre-quantum” cryptosystems like RSA or ECC since the low-level arithmetic operations are less computation-intensive and, thus, more suitable for constrained devices. In this paper we present Avrntru, a highly-optimized implementation of NTRUEncrypt for 8-bit AVR microcontrollers that we developed from scratch to reach high performance and resistance to timing attacks. Avrntru complies with the EESS #1 v3.1 specification and supports product-form parameter sets such as ees443ep1, ees587ep1, and ees743ep1. An entire encryption (including mask generation and blinding-polynomial generation) using the ees443ep1 parameters requires 847973 clock cycles on an ATmega1281 microcontroller; the decryption is more costly and has an execution time of 1051871 cycles. We achieved these results with the help of a novel hybrid technique for multiplication in a truncated polynomial ring, whereby one of the operands is a sparse ternary polynomial in product form and the other an arbitrary element of the ring. A constant-time multiplication in the ring given by the ees443ep1 parameters takes only 192577 cycles, which sets a new speed record for the arithmetic part of a lattice-based cryptosystem on AVR.
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
DATE1
2020 Lightweight Post-quantum Key Encapsulation for 8-bit AVR Microcontrollers
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
CARDIS1
2020 High-Throughput Elliptic Curve Cryptography Using AVX2 Vector Instructions
Hao Cheng 0009, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
SAC1
2019 A Lightweight Implementation of NTRU Prime for the Post-quantum Internet of Things
Hao Cheng 0009, Daniel Dinu, Johann Großschädl, Peter B. Rønne, Peter Y. A. Ryan
WISTP1