Zhenxiang Zhang

dblp:75/5544 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Acceleration of Multi-Body Molecular Dynamics With Customized Parallel Dataflow
abstract
FPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case.
Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.8
2023 Topgun: An ECC Accelerator for Private Set Intersection
abstract
Elliptic Curve Cryptography (ECC), one of the most widely used asymmetric cryptographic algorithms, has been deployed in Transport Layer Security (TLS) protocol, blockchain, secure multiparty computation, and so on. As one of the most secure ECC curves, Curve25519 is employed by some secure protocols, such as TLS 1.3 and Diffie-Hellman Private Set Intersection (DH-PSI) protocol. High-performance implementation of ECC is required, especially for the DH-PSI protocol used in privacy-preserving platform. Point multiplication, the chief cryptographic primitive in ECC, is computationally expensive. To improve the performance of DH-PSI protocol, we propose Topgun, a novel and high-performance hardware architecture for point multiplication over Curve25519. The proposed architecture features a pipelined Finite-field Arithmetic Unit and a simple and highly efficient instruction set architecture. Compared to the best existing work on Xilinx Zynq 7000 series FPGA, our implementation with one Processing Element can achieve 3.14× speedup on the same device. To the best of our knowledge, our implementation appears to be the fastest among the state-of-the-art works. We also have implemented our architecture consisting of 4 Compute Groups, each with 16 PEs, on an Intel Agilex AGF027 FPGA. The measured performance of 4.48 Mops/s is achieved at the cost of 86 Watts power, which is the record-setting performance for point multiplication over Curve25519 on FPGAs.
Guiming Wu, Qianwen He, Jiali Jiang, Zhenxiang Zhang, Yuan Zhao 0015, Yinchao Zou, Jie Zhang 0144, Changzheng Wei, Ying Yan 0002, Hui Zhang 0002
ACM Trans. Reconfigurable Technol. Syst.4
2022 A Reliability-constrained Association Rule Mining Method for Explaining Machine Learning Predictions on Continuity of Asthma Care
abstract
Continuity of care has been shown to possess numerous health benefits for asthma. However, the use of methodology to predetermine the level of care continuity has been underinvestigated. We recently built a machine learning model to predict the continuity of care for diagnosed asthma patients at the University of Washington Medicine. As there is a consensus on the un-explainable featured nowadays by black-box machine learning, which impedes the clinical deployment of our model. To tackle this issue, we proposed a reliability-constrained association rule mining method, called RC-ARM, to automatically explain the predictions of any machine learning model. First, we introduced the belief function to build a reliability-constrained rule framework. Then, the two-step reliable association rule mining algorithms were developed to generate reliable explanation rules for the machine learning model’s predictions. At last, clinical intervention suggestions regarding each mined rule were embedded for further understanding. The results showed that the proposed method could explain all (110/110) predictions of our machine learning model for asthma patients with low levels of continuity of care. This semantic-fused method sheds light on black-box models and encourages clinical experts to embrace the benefits of machine learning without any prior concern about its lack of explainability.
Qibin Zhang, Zhenxiang Zhang, Gang Chen 0037
BIBM4
2022 A High-Performance Hardware Architecture for ECC Point Multiplication over Curve25519
abstract
As one of the most secure ECC curves, Curve25519 is employed by some secure protocols, such as TLS 1.3, IRTF’s RFC7748, Diffie-Hellman Private Set Intersection (DH-PSI) protocol, etc. High performance implementation of ECC is required, especially for the DH-PSI protocol. Point multiplication, the chief cryptographic primitive in ECC, is computationally expensive. To improve the performance of DH-PSI protocol, we propose a novel and high-performance hardware architecture for point multiplication over Curve25519. The proposed architecture features a pipelined Finite-field Arithmetic Unit (FAU) and a simple and highly efficient instruction set architecture (ISA). Compared to the best existing work on Xilinx Zynq 7000 series FPGA, our implementation with one Processing Element (PE) can achieve 3.14x speedup on the same device. To the best of our knowledge, our implementation appears to be the fastest among the state-of-the-art works. We also have implemented our proposed architecture consisting of 4 Compute Groups (CGs), each with 16 PEs, on an Intel Agilex AGF027 FPGA. The experimental results show the peak performance of 4.52 Mops/s (million point multiplication operations per seconds) can be achieved. Moreover, the measured performance of 4.48 Mops/s is achieved, with the PE utilization of 99% and at the cost of 86 Watts power, which is the record-setting performance for point multiplication over Curve25519 on FPGAs.
Guiming Wu, Qianwen He, Jiali Jiang, Zhenxiang Zhang, Yuan Zhao 0015, Yinchao Zou
FCCM4
2020 HaoCL: Harnessing Large-scale Heterogeneous Processors Made Easy
abstract
The pervasive adoption of Deep Learning (DL) and Graph Processing (GP) makes it a de facto requirement to build large-scale clusters of heterogeneous accelerators including GPUs and FPGAs. The OpenCL programming framework can be used on the individual nodes of such clusters but is not intended for deployment in a distributed manner. Fortunately, the original OpenCL semantics naturally fit into the programming environment of heterogeneous clusters. In this paper, we propose a heterogeneity-aware OpenCL-like (HaoCL) programming framework to facilitate the programming of a wide range of scientific applications including DL and GP workloads on large-scale heterogeneous clusters. With HaoCL, existing applications can be directly deployed on heterogeneous clusters without any modifications to the original OpenCL source code and without awareness of the underlying hardware topologies and configurations. Our experiments show that HaoCL imposes a negligible overhead in a distributed environment, and provides near-liner speedups on standard benchmarks when computation or data size exceeds the capacity of a single node. The system design and the evaluations are presented in this demo paper.
Yao Chen 0008, Jiong He, Hongshi Tan, Zhenxiang Zhang, Marianne Winslett, Deming Chen
ICDCS6
2010 On the effectiveness of a generalization of Miller's primality theorem
Zhenxiang Zhang
J. Complex.1