VLDB 2026 Research / reviewers in the wild / expert
Fei Li 0045
dblp:87/3534-45
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0002-1706-3625ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimization of General Matrix Multiplication for RISC-V ProcessorsabstractGeneral Matrix Multiplication (GEMM) is a fundamental operation in high-performance computing and artificial intelligence, with its performance directly impacting the overall efficiency of higher-level applications. This paper optimizes GEMM by leveraging the characteristics of the RISC-V architecture, achieving significant performance improvements on the XuanTie C908 and C910 processors. It is also observed that traditional optimization method become a bottleneck for small-scale GEMM performance on RISC-V. To address this, we propose an analytical model-based approach that effectively enhances small-scale GEMM performance. In performance evaluation, we tested square matrices and irregular matrices (derived from convolutional neural networks) to validate the optimized GEMM across various scenarios and assess its performance in realworld applications. Compared to the GEMM routines in popular open-source BLAS (Basic Linear Algebra Subprograms) libraries (OpenBLAS and BLIS), the optimized GEMM achieves average speedups of$1.29 \times$and$2.20 \times$, with maximum speedups of$6.51 \times$and$27.85 \times$on the C908 platform. On the C910 platform, the average speedups are$2.38 \times$and$5.15 \times$, with maximum speedups of$17.15 \times$and$84.16 \times$. Bei Zhou 0004, Jianmin Pang, Fei Li 0045, Hongru Yang, Mengyao Duan, Jinchen Xu |
HPCC | 4 |
| 2024 | Arfa: An Agile Regime-Based Floating-Point Optimization Approach for Rounding ErrorsabstractWe introduce a floating-point (FP) error optimization approach called Arfa that partitions the domain D of an FP expression fe into regimes and rewrites fe in each regime where fe shows larger errors. First, Arfa seeks a rewrite substitution fo with lower errors across D, whose error distribution is plotted for effective regime inference. Next, Arfa generates an incomplete set of ordered rewrite candidates within each regime of interest, so that searching for the best rewrite substitutions is performed efficiently. Finally, Arfa selects the best rewrite substitution by inspecting the errors of top ranked rewrite candidates, with enhancing precision also considered. Experiments on 56 FPbench examples and four real-life programs show that Arfa not only reduces the maximum and average errors of fe by 4.73 and 2.08 bits on average (and up to 33 and 16 bits), but also exhibits lower errors, sometimes to a significant degree, than Herbie and NumOpt. Jinchen Xu, Mengqi Cui, Fei Li 0045, Zuoyan Zhang, Hongru Yang, Bei Zhou 0004, Jie Zhao 0002 |
ISSTA | 3 |
| 2024 | A Holistic Approach to Automatic Mixed-Precision Code Generation and Tuning for Affine ProgramsabstractReducing floating-point (FP) precision is used to trade the quality degradation of a numerical program's output for performance, but this optimization coincides with type casting, whose overhead is undisclosed until a mixed-precision code version is generated. This uncertainty enforces the decoupled implementation of mixed-precision code generation and autotuning in prior work. In this paper, we present a holistic approach called PrecTuner that consolidates the mixed-precision code generator and the autotuner by defining one parameter. This parameter is first initialized by some automatically sampled values and used to generate several code variants, with various loop transformations also taken into account. The generated code variants are next profiled to solve a performance model formulated using the aforementioned parameter, possibly under a pre-defined quality degradation budget. The best-performing value of the defined parameter is finally predicted without evaluating all code variants. Experimental results of the PolyBench benchmarks on CPU demonstrate that PrecTuner outperforms LuIs by 3.28× while achieving smaller errors, and we also validate its effectiveness in optimizing a real-life large-scale application. In addition, PrecTuner also obtains a mean speedup of 1.81× and 1.52×-1.73× over Pluto on single- and multi-core CPU, respectively, and 1.71× over PPCG on GPU. Jinchen Xu, Guanghui Song, Bei Zhou 0004, Fei Li 0045, Jiangwei Hao, Jie Zhao 0002 |
PPoPP | 4 |
| 2024 | SCR-LIBM: A Correctly Rounded Elementary Function Library in Double-PrecisionabstractThe MPFR and CR-LIBM math libraries are frequently utilized due to their ability to generate correctly rounded results for all double-precision inputs. However, it is worth noting that MPFR has a slower average performance, while CR-LIBM achieves correct rounding over two iterations, rendering it less stable. In addition, CR-LIBM has a poor performance in handling the worst-case of correct rounding. This paper implements a correctly rounded elementary function library called SCR-LIBM in double-precision, which is stable and efficient. Our key idea is to divide subdomains and use the low-degree Taylor polynomial to approximate the elementary function in each subdomain. We simulate the high-precision representation based on the double–double data format, and use the error-free transformation and Double-double algorithm to control the error in the process of polynomial approximation and output compensation. Our approach ensures that the elementary function is correctly rounded, without the need for redundant iterations. The experimental evaluation shows that the average performance of elementary functions implemented in SCR-LIBM is 8.534 times faster than that of MPFR, and 2.492 times faster than that of CR-LIBM when dealing with the worst-case of correct rounding. What’s more, our SCR-LIBM is more stable than CR-LIBM. Jinchen Xu, Bei Zhou 0004, Jiangwei Hao, Fei Li 0045, Zuoyan Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2023 | AMPT: Automatic Mixed-Precision Tuning Based on Accuracy GainabstractMixed-precision tuning is a valuable approach for striking a balance between the accuracy and performance of floating-point computations. However, selecting suitable configurations from numerous mixed-precision configurations can be challenging. The discrete and finite nature of floating-point numbers introduces unavoidable rounding errors that are difficult to predict. To address these challenges, we propose an accuracy gain evaluation function for expressions. This evaluation function quantifies the extent of accuracy improvement achieved by adjusting the precision of the operations within the expression. This is achieved through a thorough examination of the impact of each operation's condition number and implementation error on the final error. Based on this evaluation function, we present AMPT, a mixed-precision tuning method specifically for expressions. AMPT not only identifies the optimal mixed-precision configuration but also automatically generates the corresponding code. Experimental results demonstrate that AMPT accurately assesses the accuracy gains of different mixed-precision configurations, and the mixed-precision code generated by AMPT can achieve low-error computation of expressions with low performance overhead. Jiangwei Hao, Yuanyuan Xia, Fei Li 0045, Hongru Yang, Zongjiang Yi, Bei Zhou 0004, Jianmin Pang |
APSEC | 3 |
| 2023 | Eiffel: Inferring Input Ranges of Significant Floating-point Errors via Polynomial ExtrapolationabstractExisting search heuristics used to find input values that result in significant floating-point (FP) errors or small ranges that cover them are accompanied by severe constraints, complicating their implementation and restricting their general applicability. This paper introduces an error analysis tool called Eiffel to infer error-inducing input ranges instead of searching them. Given an FP expression with its domain$\mathcal{D}$, Eiffel first constructs an error data set by sampling values across a smaller domain$\mathcal{R}$and assembles these data into clusters. If more than two clusters are formed, Eiffel derives polynomial curves that best fit the bound coordinates of the error-inducing ranges in$\mathcal{R}$, extrapolating them to infer all target ranges of$\mathcal{D}$and reporting the maximal error. Otherwise, Eiffel simply returns the largest error across$\mathcal{R}$. Experimental results show that Eiffel exhibits a broader applicability than Atomu and$\mathbf{S}^{3}$FP by successfully detecting the errors of all 70 considered benchmarks while the two baselines only report errors for part of them. By taking as input the inferred ranges of Eiffel, Herbie obtains an average accuracy improvement of 3.35 bits and up to 53.3 bits. Zuoyan Zhang, Bei Zhou 0004, Jiangwei Hao, Hongru Yang, Mengqi Cui, Yuchang Zhou, Guanghui Song, Fei Li 0045, Jinchen Xu, Jie Zhao 0002 |
ASE | 8 |