VLDB 2026 Research / reviewers in the wild / expert
Jiangbin Dong
dblp:298/8815
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CROPHE: Cross-Operator Dataflow Optimization for Fully Homomorphic Encryption AcceleratorsabstractFully homomorphic encryption (FHE) enables the protection of data privacy at the cost of significantly higher computational demands. To alleviate its memory-bound bottlenecks, dataflow optimizations that maximize on-chip data reuse and minimize off-chip accesses could be leveraged. In this work, we exploit the opportunities of cross-operator dataflow optimizations in FHE accelerators, and propose a hardware-software co-design called CROPHE. On the hardware level, instead of overly-specialized functional units, CROPHE provisions a homogeneous and unified architecture that allows for flexible resource allocation and operator mapping. On the software level, the scheduling framework of CROPHE takes a comprehensive and systematic approach to explore various spatial and temporal data pipelining and sharing schemes across multiple operators, resulting in more efficient dataflow than prior work. We also propose novel cross-operator dataflow optimizations for the unique operators in FHE including number theoretic transforms and homomorphic rotations. The evaluation shows CROPHE significantly outperforms state-of-the-art designs by$1.77 \times$to$4.86 \times$. Xinhua Chen, Jiangbin Dong, Hongren Zheng, Tian Tang 0001, Mingyu Gao 0001 |
HPCA | 2 |
| 2026 | GenZA: A General and Efficient Accelerator for Diverse Zero-Knowledge Proof Protocols
Jiangbin Dong, Mingyu Gao 0001 |
ISCA | 2 |
| 2025 | A Unified Vector Processing Unit for Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) algorithms enable privacy-preserving computing directly on encrypted data without leaking sensitive contents, while their excessive computational overheads could be alleviated by specialized hardware accelerators. The vector architecture has been prominently used for FHE accelerators to match the underlying polynomial data structures. While most FHE operations can be efficiently supported by vector processing units, the number theoretic transform (NTT) and automorphism operators involve complex and irregular data permutations among vector elements, and thus are handled with separate dedicated hardware units in existing FHE accelerators. In this paper, we present an efficient inter-lane network design and the corresponding dataflow control scheme, in order to realize NTT and automorphism operations among the multiple lanes of a vector unit. An arbitrarily large operator is first decomposed to fit in the fixed width of the vector unit, and the required data permutation and transposition are conducted on the specialized inter-lane network. Compared to previous designs, our solution reduces the hardware resources needed, with up to 9.4x area and 6.0x power savings for only the inter-lane network, and up to 1.2 x area and 1.1 x power savings for the whole vector unit. Jiangbin Dong, Xinhua Chen, Mingyu Gao 0001 |
DATE | 1 |
| 2021 | PipeZK: Accelerating Zero-Knowledge Proof with a Pipelined ArchitectureabstractZero-knowledge proof (ZKP) is a promising cryptographic protocol for both computation integrity and privacy. It can be used in many privacy-preserving applications including verifiable cloud outsourcing and blockchains. The major obstacle of using ZKP in practice is its time-consuming step for proof generation, which consists of large-size polynomial computations and multi-scalar multiplications on elliptic curves. To efficiently and practically support ZKP in real-world applications, we propose PipeZK, a pipelined accelerator with two subsystems to handle the aforementioned two intensive compute tasks, respectively. The first subsystem uses a novel dataflow to decompose large kernels into smaller ones that execute on bandwidth-efficient hardware modules, with optimized off-chip memory accesses and on-chip compute resources. The second subsystem adopts a lightweight dynamic work dispatch mechanism to share the heavy processing units, with minimized resource underutilization and load imbalance. When evaluated in 28 nm, PipeZK can achieve 10x speedup on standard cryptographic benchmarks, and 5x on a widely-used cryptocurrency application, Zcash. Ye Zhang 0042, Shuo Wang 0009, Xian Zhang 0001, Jiangbin Dong, Xingzhong Mao, Fan Long, Dong Zhou 0006, Mingyu Gao 0001, Guangyu Sun 0003 |
ISCA | 4 |