Bei Zhou 0004

dblp:14/153-4 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-1515-0602ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021
YearPublicationVenuePosition
2025 GATUNER:Genetic Algorithm Applied to Floating-Point Precision Tuning
abstract
Floating-point arithmetic is widely used in high-precision computing fields such as national defense, aerospace, and finance. Different applications have varying requirements for floating-point computation precision, and how to maximize program performance while meeting precision requirements is a crucial challenge in high-performance program design. Mixed-precision technology is a common optimization strategy to address this issue, which involves using multiple precision types within the same program. However, most existing research on mixed precision tends to fall into local optima and fails to directly provide users with usable mixed-precision programs. Since genetic algorithms are optimization methods with global search capabilities, we propose a genetic algorithm-based mixed-precision tuning method and develop a tool called GATUNER. Specifically, GATUNER first analyzes the variables and statements of the program to construct a dependency graph of the program. Then, it uses the Leiden algorithm to detect community structures in the dependency graph, employs an improved genetic algorithm to search for mixed-precision configurations, and finally automatically generates the corresponding mixed-precision program. We evaluated GATUNER on 14 benchmark tests and application programs. The experimental results show that GATUNER can find valid mixed-precision configurations for 76.79% of the test programs, with an average performance improvement of 31.65% for the found mixed-precision configurations. Compared with HIFPTUNER, ADAPT, and CHEF-FP, the average performance improvement of the mixed-precision configurations found by GATUNER increased by 14.1%, 24.09%, and 15.18% respectively. In addition, compared with PESA, the average performance improvement of the mixed-precision configurations obtained by GATUNER increased by 18.67%.
Jinchen Xu, Hongru Yang, Bei Zhou 0004
APSEC6
2025 Optimization of General Matrix Multiplication for RISC-V Processors
abstract
General Matrix Multiplication (GEMM) is a fundamental operation in high-performance computing and artificial intelligence, with its performance directly impacting the overall efficiency of higher-level applications. This paper optimizes GEMM by leveraging the characteristics of the RISC-V architecture, achieving significant performance improvements on the XuanTie C908 and C910 processors. It is also observed that traditional optimization method become a bottleneck for small-scale GEMM performance on RISC-V. To address this, we propose an analytical model-based approach that effectively enhances small-scale GEMM performance. In performance evaluation, we tested square matrices and irregular matrices (derived from convolutional neural networks) to validate the optimized GEMM across various scenarios and assess its performance in realworld applications. Compared to the GEMM routines in popular open-source BLAS (Basic Linear Algebra Subprograms) libraries (OpenBLAS and BLIS), the optimized GEMM achieves average speedups of$1.29 \times$and$2.20 \times$, with maximum speedups of$6.51 \times$and$27.85 \times$on the C908 platform. On the C910 platform, the average speedups are$2.38 \times$and$5.15 \times$, with maximum speedups of$17.15 \times$and$84.16 \times$.
Bei Zhou 0004, Jianmin Pang, Fei Li 0045, Hongru Yang, Mengyao Duan, Jinchen Xu
HPCC2
2025 PESA: error sensitivity analysis tool for floating-point computational programs
Mengqi Cui, Jinchen Xu, Yuchang Zhou, Hongru Yang, Liguang Ji, Bei Zhou 0004
J. Supercomput.6
2024 Arfa: An Agile Regime-Based Floating-Point Optimization Approach for Rounding Errors
abstract
We introduce a floating-point (FP) error optimization approach called Arfa that partitions the domain D of an FP expression fe into regimes and rewrites fe in each regime where fe shows larger errors. First, Arfa seeks a rewrite substitution fo with lower errors across D, whose error distribution is plotted for effective regime inference. Next, Arfa generates an incomplete set of ordered rewrite candidates within each regime of interest, so that searching for the best rewrite substitutions is performed efficiently. Finally, Arfa selects the best rewrite substitution by inspecting the errors of top ranked rewrite candidates, with enhancing precision also considered. Experiments on 56 FPbench examples and four real-life programs show that Arfa not only reduces the maximum and average errors of fe by 4.73 and 2.08 bits on average (and up to 33 and 16 bits), but also exhibits lower errors, sometimes to a significant degree, than Herbie and NumOpt.
Jinchen Xu, Mengqi Cui, Fei Li 0045, Zuoyan Zhang, Hongru Yang, Bei Zhou 0004, Jie Zhao 0002
ISSTA6
2024 A Holistic Approach to Automatic Mixed-Precision Code Generation and Tuning for Affine Programs
abstract
Reducing floating-point (FP) precision is used to trade the quality degradation of a numerical program's output for performance, but this optimization coincides with type casting, whose overhead is undisclosed until a mixed-precision code version is generated. This uncertainty enforces the decoupled implementation of mixed-precision code generation and autotuning in prior work. In this paper, we present a holistic approach called PrecTuner that consolidates the mixed-precision code generator and the autotuner by defining one parameter. This parameter is first initialized by some automatically sampled values and used to generate several code variants, with various loop transformations also taken into account. The generated code variants are next profiled to solve a performance model formulated using the aforementioned parameter, possibly under a pre-defined quality degradation budget. The best-performing value of the defined parameter is finally predicted without evaluating all code variants. Experimental results of the PolyBench benchmarks on CPU demonstrate that PrecTuner outperforms LuIs by 3.28× while achieving smaller errors, and we also validate its effectiveness in optimizing a real-life large-scale application. In addition, PrecTuner also obtains a mean speedup of 1.81× and 1.52×-1.73× over Pluto on single- and multi-core CPU, respectively, and 1.71× over PPCG on GPU.
Jinchen Xu, Guanghui Song, Bei Zhou 0004, Fei Li 0045, Jiangwei Hao, Jie Zhao 0002
PPoPP3
2024 SCR-LIBM: A Correctly Rounded Elementary Function Library in Double-Precision
abstract
The MPFR and CR-LIBM math libraries are frequently utilized due to their ability to generate correctly rounded results for all double-precision inputs. However, it is worth noting that MPFR has a slower average performance, while CR-LIBM achieves correct rounding over two iterations, rendering it less stable. In addition, CR-LIBM has a poor performance in handling the worst-case of correct rounding. This paper implements a correctly rounded elementary function library called SCR-LIBM in double-precision, which is stable and efficient. Our key idea is to divide subdomains and use the low-degree Taylor polynomial to approximate the elementary function in each subdomain. We simulate the high-precision representation based on the double–double data format, and use the error-free transformation and Double-double algorithm to control the error in the process of polynomial approximation and output compensation. Our approach ensures that the elementary function is correctly rounded, without the need for redundant iterations. The experimental evaluation shows that the average performance of elementary functions implemented in SCR-LIBM is 8.534 times faster than that of MPFR, and 2.492 times faster than that of CR-LIBM when dealing with the worst-case of correct rounding. What’s more, our SCR-LIBM is more stable than CR-LIBM.
Jinchen Xu, Bei Zhou 0004, Jiangwei Hao, Fei Li 0045, Zuoyan Zhang
Int. J. Softw. Eng. Knowl. Eng.3
2024 Hierarchical search algorithm for error detection in floating-point arithmetic expressions
Zuoyan Zhang, Jinchen Xu, Jiangwei Hao, Haotian He, Bei Zhou 0004
J. Supercomput.6
2023 AMPT: Automatic Mixed-Precision Tuning Based on Accuracy Gain
abstract
Mixed-precision tuning is a valuable approach for striking a balance between the accuracy and performance of floating-point computations. However, selecting suitable configurations from numerous mixed-precision configurations can be challenging. The discrete and finite nature of floating-point numbers introduces unavoidable rounding errors that are difficult to predict. To address these challenges, we propose an accuracy gain evaluation function for expressions. This evaluation function quantifies the extent of accuracy improvement achieved by adjusting the precision of the operations within the expression. This is achieved through a thorough examination of the impact of each operation's condition number and implementation error on the final error. Based on this evaluation function, we present AMPT, a mixed-precision tuning method specifically for expressions. AMPT not only identifies the optimal mixed-precision configuration but also automatically generates the corresponding code. Experimental results demonstrate that AMPT accurately assesses the accuracy gains of different mixed-precision configurations, and the mixed-precision code generated by AMPT can achieve low-error computation of expressions with low performance overhead.
Jiangwei Hao, Yuanyuan Xia, Fei Li 0045, Hongru Yang, Zongjiang Yi, Bei Zhou 0004, Jianmin Pang
APSEC6
2023 Eiffel: Inferring Input Ranges of Significant Floating-point Errors via Polynomial Extrapolation
abstract
Existing search heuristics used to find input values that result in significant floating-point (FP) errors or small ranges that cover them are accompanied by severe constraints, complicating their implementation and restricting their general applicability. This paper introduces an error analysis tool called Eiffel to infer error-inducing input ranges instead of searching them. Given an FP expression with its domain$\mathcal{D}$, Eiffel first constructs an error data set by sampling values across a smaller domain$\mathcal{R}$and assembles these data into clusters. If more than two clusters are formed, Eiffel derives polynomial curves that best fit the bound coordinates of the error-inducing ranges in$\mathcal{R}$, extrapolating them to infer all target ranges of$\mathcal{D}$and reporting the maximal error. Otherwise, Eiffel simply returns the largest error across$\mathcal{R}$. Experimental results show that Eiffel exhibits a broader applicability than Atomu and$\mathbf{S}^{3}$FP by successfully detecting the errors of all 70 considered benchmarks while the two baselines only report errors for part of them. By taking as input the inferred ranges of Eiffel, Herbie obtains an average accuracy improvement of 3.35 bits and up to 53.3 bits.
Zuoyan Zhang, Bei Zhou 0004, Jiangwei Hao, Hongru Yang, Mengqi Cui, Yuchang Zhou, Guanghui Song, Fei Li 0045, Jinchen Xu, Jie Zhao 0002
ASE2
2019 Memory latency optimizations for the elementary functions on the Sunway architecture
Bei Zhou 0004, Yongzhong Huang, Jinchen Xu, Shaozhong Guo, Hongyuan Qi
J. Supercomput.1