EDBT 2026 Demo / reviewers in the wild / expert
Yifan Zhao 0007
dblp:13/7050-7
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-7304-8206ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DNA-HHE: Dual-mode Near-network Accelerator for Hybrid Homomorphic Encryption on the Edge
Yifan Zhao 0007, Xinglong Yu, Honglin Kuang, Jun Han 0003 |
ISCAS | 1 |
| 2026 | RAEnc: A Stall-Less Metadata Compression Framework for Return Address Integrity on High-Performance Embedded ProcessorsabstractProtecting return address integrity (RAI) on high-performance embedded processors (HPEPs) is more challenging than on traditional microcontroller units (MCUs) or server processors. Isolation-based schemes such as shadow stacks fail against hardware memory threats, while crypto-based schemes remain vulnerable to forgery and replay attacks. To address the challenges of RAI protection in HPEPs, this article presentsRAEnc, a novel crypto-based scheme, which provides comprehensive RAI against both forgery and replay attacks with negligible performance overhead via hardware–software co-design. We first introduce the metadata compression during return address encryption (MCRAE), a cryptographic primitive compressing the call stack’s integrity state into a single, register-resident value using custom RISC-V instructions, thereby eliminating the memory attack surface. We then detail the stall-less encryption/decryption unit (EDU) that is tightly coupled with the processor pipeline. By employing speculative scheduling, a low-latency 128-bit QARMA unit, encryption/decryption cache (EDCache), and multichain parallelism, the EDU eliminates pipeline stalls common in decoupled accelerators. Implemented on the open-source BOOM processor,RAEncincurs a negligible 0.2% performance overhead on embedded and SPEC CPU2017 workloads with modest hardware cost, outperforming existing crypto-based RAI schemes in both security and performance. Zikang Zhou, Kanheng Jiang, Yifan Zhao 0007, Jun Han 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Sliding-Window Scheduling to Exploit Hybrid-Bonding-Based Accelerators for Fully Homomorphic Encryption
Xinhua Chen, Xinglong Yu, Yifan Zhao 0007, Honglin Kuang, Jun Han 0003 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | ARV-Q: An Adaptive RISC-V Vector Processor for Unified Support of Post-Quantum Standards and Side-Channel Protection on the EdgeabstractUnder the threat of quantum computers, the public-key cryptosystems need to transition to the Post-Quantum Cryptography (PQC) standards. However, this migration process is hindered by the diverse mathematical structures of PQC standards as well as the side-channel attacks, especially for the resource-constrained and physically accessible edge devices. To address this issue, we present ARV-Q, an adaptive RISC-V vector processor that efficiently offers unified support of PQC standards and side-channel security enhancement. Firstly, we propose an adaptive RISC-V-based computing paradigm to adapt to PQC algorithms across diverse mathematical bases, the core of which is a crypto extension supporting all PQC standards plus schemes in the fourth round in NIST standardization process. This crypto extension can well cooperate with the RISC-V Vector Extension (RVV) and is capable of adaptive operator-type support. Second, we design a highly resource-efficient vector crypto engine featuring versatile Butterfly Units and multi-Selected-Element-Width-adaptive modular arithmetic constructs, achieving high hardware utilization with configurable parameter settings. The crypto engine is integrated into the RISC-V core via an agile extension interface capable of bridging all types of register transfers. Besides, due to the capability of cooperative hybrid vector computing of RVV and crypto extension, ARV-Q can flexibly adapt to side-channel attacks with no hardware overhead while maintaining high performance. ARV-Q is implemented in 22nm process, and post-layout simulations are conducted. Results outperform the state-of-the-art counterparts in primary PQC standards with 1.21-5.14× better latency and more than 4.41× better area efficiency. Yifan Zhao 0007, Honglin Kuang, Xinglong Yu, Ziyi Hao, Jian-Yi Meng, Jun Han 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | RVCE-FAL: A RISC-V Scalar-Vector Custom Extension for Faster FALCON Digital SignatureabstractThe National Institute of Standards and Technology (NIST) has selected FALCON as one of the standardized digital signature algorithms against quantum attacks in 2022. Compared with the other post-quantum cryptography (PQC) schemes, lattice-based FALCON is more appropriate for future Internet of Things (loT) applications due to the fastest signature verification process and the lowest transmission overhead. In this paper, we propose a custom extension based on the RISC-V scalar-vector framework for efficient implementation of FALCON. To our best knowledge, this work is the first hardware-software co-design for complete FALCON signature generation and verification routines. Besides, we design the first FALCON Gaussian sampling hardware and a RISC-V vector extension (RVV) based domain-specific core. The proposed architecture accelerates kernel operations in FALCON, such as discrete Gaussian sampling, number theoretic transform (NTT), inverse NTT, and polynomial operations. Compared with the reference implementation, results on the gem5-RTL simulation platform present a speedup for signature generation and verification of up to 18 x and 6.9 x. Xinglong Yu, Yifan Zhao 0007, Honglin Kuang, Jun Han 0003 |
DATE | 3 |
| 2023 | Enhancing RISC-V Vector Extension for Efficient Application of Post-Quantum CryptographyabstractWe present a cryptography extension built on RISC-V Vector Extension for efficient application of lattice-based post-quantum cryptography, offering custom instructions that can perform vectorized operations on polynomials of variable length and data width. We use micro-operation architecture to simplify the execution of variable-latency vector instructions and propose fracturable modular arithmetic units to support operations on variable coefficient width. On this basis, a vector unit is designed, achieving significant speed-up compared to the state-of-the-art counterparts for number-theoretic-transform-based polynomial multiplication. This cryptography extension is further integrated into the gem5 simulator to evaluate CRYSTALS-Kyber and CRYSTALS-Dilithium; results outperform the state-of-the-art implementations with more than 2.3 × improvement in cycle count. Yifan Zhao 0007, Honglin Kuang, Chen Chen 0058, Jian-Yi Meng, Jun Han 0003 |
ASAP | 1 |
| 2022 | A High-Performance Domain-Specific Processor With Matrix Extension of RISC-V for Module-LWE ApplicationsabstractThe 5G edge computing infrastructure should be empowered with quantum attack resistance by implementing post-quantum cryptography (PQC). Among various PQC schemes, lattice-based cryptography (LBC) based on learning with error (LWE) has attracted much attention because of its performance efficiency and security guarantee. In LWE-based LBCs, the Module-LWE-based schemes gain advantage over the others benefiting from the unique polynomial matrix and vector structure. To provide a high-performance implementation of Module-LWE applications for the edge computing paradigm, we propose a domain-specific processor based on a matrix extension of RISC-V architecture. This custom extension encapsulates the matrix-based ring operations with a high-level functional abstraction. A 2-D systolic array with configurable functionality is proposed to perform matrix-based number theoretic transform (NTT) and other arithmetic operations, achieving high data-level parallelism with support for the variable-sized polynomial matrix and vector structure. As this structure of Module-LWE involves no data dependency between different inner elements, an out-of-order mechanism is further developed to exploit the instruction-level parallelism. We implement the proposed architecture under TSMC 28nm technology. The evaluation results show that our implementation can achieve up to$3.5\times $and$3.3\times $improvement in cycle count respectively in Kyber and Dilithium, compared to the state-of-the-art crypto-processor counterparts. Yifan Zhao 0007, Ruiqi Xie, Guozhu Xin, Jun Han 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | A Multi-Layer Parallel Hardware Architecture for Homomorphic Computation in Machine LearningabstractHomomorphic Encryption (HE) allows untrusted parties to process encrypted data without revealing its content. People could encrypt the data locally and send it to the cloud to conduct neural network training or inferencing, which achieves data privacy in AI. However, the combined AI and HE computation could be extremely slow. To deal with it, we propose a multi-level parallel hardware accelerator for homomorphic computations in machine learning. The vectorized Number Theoretic Transform (NTT) unit is designed to form the low-level parallelism, and we apply a Residue Number System (RNS) to form the mid-level parallelism in one polynomial. Finally, a fully pipelined and parallel accelerator for two ciphertext operands is proposed to form the high-level parallelism. To address the core computation (matrix-vector multiplication) in neural networks, our work is designed to support Multiply-Accumulate (MAC) operations natively between ciphertexts. We have analyzed our design on FPGA ZCU102, and experimental results show that it outperforms previous works and achieves over an order of magnitude acceleration than software implementations. Guozhu Xin, Yifan Zhao 0007, Jun Han 0003 |
ISCAS | 2 |