EDBT 2026 Demo / reviewers in the wild / expert
Yi Bian 0001
dblp:133/2197-1
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0008-4308-3084ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ML-Cube: Accelerating Module-Lattice-Based Cryptography using Machine Learning Accelerators with a Memory-Less DesignabstractThe rapid advancement of AI technologies has led to a dramatic surge in computational demands, driving significant breakthroughs in ML accelerators. The powerful performance of these accelerators has attracted the attention of cryptography researchers, and recent studies have begun to explore their use in accelerating cryptographic operations. However, treating these accelerators as black boxes leads to high latency, and strict concurrency requirements, which hinder their practical deployment. In this paper, we go beyond the black-box treatment of ML accelerators and introduce ML-Cube (ML3), a novel memory-less framework that leverages ML accelerators to implement module-lattice-based PQC, FIPS 203 ML-KEM, and FIPS 204 ML-DSA. The performance benefits of ML-Cube arise from our thorough analysis of ML accelerator internals. Rather than treating the accelerators as black boxes, we dissect their operating mechanisms and design tailored mathematical transformations for cryptographic acceleration. This enables memory-less (I)NTT and polynomial multiplication that minimizes external memory dependencies and reduces latency. We further address the high latency and excessive parallelism demands of traditional SIMT-based implementations by fully parallelizing both ML-KEM and ML-DSA schemes. Our experiments show that our Tensor Core-based (I)NTT achieves a 2.03x--3.56x speedup over a highly-optimized CUDA-core implementation. Moreover, our memory-less polynomial multiplication attains a 10x speedup, and the full ML-KEM reaches up to a 3.58x speedup with only less than one-tenth of the latency compared with SOTA approach (CHES '24). Additionally, our enhanced ML-DSA implementation offers a 30% to 55% throughput improvement over the previous SOTA methods (TDSC '24) under the server-oriented model. Importantly, by confining core computations within registers, our approach inherently mitigates memory disclosure and cache-based side-channel attacks, thereby enhancing overall security. Fangyu Zheng, Zhuoyu Xie, Wenxu Tang, Guang Fan 0001, Yijing Ning, Yi Bian 0001, Jingqiang Lin 0001, Jiwu Jing |
CCS | 7 |
| 2025 | AsyncGBP${}^{+}$+: Bridging SSL/TLS and Heterogeneous Computing Power With GPU-Based ProvidersabstractThe rapid evolution of GPUs has emerged as a promising solution for accelerating the worldwide used SSL/TLS, which faces performance bottlenecks due to its underlying heavy cryptographic computations. Nevertheless, substantial structural adjustments from the parallel mode of GPUs to the serial mode of the SSL/TLS stack are imperative, potentially constraining the practical deployment of GPUs. In this paper, we propose AsyncGBP${}^{+}$, a three-level framework that facilitates the seamless conversion of cryptographic requests from synchronous to asynchronous mode. We conduct an in-depth analysis of the OpenSSL provider and cryptographic primitive features relevant to GPU implementations, aiming to fully exploit the potential of GPUs. Notably, AsyncGBP${}^{+}$supports three working settings (offline/online/hybrid), finely tailored for various public key cryptographic primitives, including traditional ones like X25519, Ed25519, ECDSA, and the quantum-safe CRYSTALS-Kyber. A comprehensive evaluation demonstrates that AsyncGBP${}^{+}$can efficiently achieve an improvement of up to 137.8$\times$compared to the default OpenSSL provider (for X25519, Ed25519, ECDSA) and 113.30$\times$compared to OpenSSL-compatibleliboqs(for CRYSTALS-Kyber) in a single-process setting. Furthermore, AsyncGBP${}^{+}$surpasses the current fastest commercial-off-the-shelf OpenSSL-compatible TLS accelerator with a 5.3$\times$to 7.0$\times$performance improvement. Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Guang Fan 0001, Jiwu Jing |
IEEE Trans. Computers | 1 |
| 2024 | TensorPolyMul: Accelerating Polynomial Multiplication in NTT-unfriendly Lattice-based Cryptography Using Tensor CoresabstractThe urgent demand for computing power in Artificial intelligence (AI) technology has driven the rapid development of dedicated accelerators. Meanwhile, the threat posed by quantum computing to traditional public-key cryptography has prompted the emergence of post-quantum algorithms, such as lattice-based cryptography. However, performance issues with these algorithms have raised concerns within the industry about the transition to quantum-safe solutions. In this paper, we propose a novel universal framework for NTT-unfriendly lattice-based post-quantum algorithms, leveraging NVIDIA’s AI accelerator Tensor Core to address this challenge. By employing techniques such as polynomial matrixization and multi-precision representation, we effectively transform the primary workload (i.e., polynomial multiplication) into a series of small-coefficient matrix multiplications that can be directly accelerated by Tensor Cores. This approach effectively bridges the gap between typical Tensor Core workloads and the core workloads of lattice-based post-quantum cryptography. As a case study, we implemented a prototype called TensorPolyMul to provide an implementation of Saber, a quantum-safe Key Encapsulation Mechanism (KEM). The experiments showcase that TensorPolyMul surpasses the state-of-the-art Tensor Core-based work, achieving remarkable speed-ups of $1.53 \times 1.33 \times, 1.62 \times$, and $1.22 \times$ for Inner Product, MatrixVecMul, Encaps, and Decaps, respectively. Yi Bian 0001, Fangyu Zheng, Jiwu Jing |
ICPADS | 1 |
| 2023 | AsyncGBP: Unleashing the Potential of Heterogeneous Computing for SSL/TLS with GPU-based ProviderabstractThe proliferation of IoT and 5G technologies has led to an explosion of data traffic that data centers must handle while ensuring secure transmission via SSL/TLS. The high volume of cryptographic operations required imposes performance bottlenecks. The GPU-based cryptographic accelerator is one of the competitive solutions. However, significant structural differences with practical applications confine their capacities to specific domains, such as offline cryptanalysis, undermining their potential for real-world cryptographic acceleration. Yi Bian 0001, Fangyu Zheng, Yuewu Wang, Lingguang Lei, Jiankuo Dong, Jiwu Jing |
ICPP | 1 |