EDBT 2026 Demo / reviewers in the wild / expert
Jinchen Xu
dblp:120/1768
· DBLP profile ↗
19ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-6275-2617ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GATUNER:Genetic Algorithm Applied to Floating-Point Precision TuningabstractFloating-point arithmetic is widely used in high-precision computing fields such as national defense, aerospace, and finance. Different applications have varying requirements for floating-point computation precision, and how to maximize program performance while meeting precision requirements is a crucial challenge in high-performance program design. Mixed-precision technology is a common optimization strategy to address this issue, which involves using multiple precision types within the same program. However, most existing research on mixed precision tends to fall into local optima and fails to directly provide users with usable mixed-precision programs. Since genetic algorithms are optimization methods with global search capabilities, we propose a genetic algorithm-based mixed-precision tuning method and develop a tool called GATUNER. Specifically, GATUNER first analyzes the variables and statements of the program to construct a dependency graph of the program. Then, it uses the Leiden algorithm to detect community structures in the dependency graph, employs an improved genetic algorithm to search for mixed-precision configurations, and finally automatically generates the corresponding mixed-precision program. We evaluated GATUNER on 14 benchmark tests and application programs. The experimental results show that GATUNER can find valid mixed-precision configurations for 76.79% of the test programs, with an average performance improvement of 31.65% for the found mixed-precision configurations. Compared with HIFPTUNER, ADAPT, and CHEF-FP, the average performance improvement of the mixed-precision configurations found by GATUNER increased by 14.1%, 24.09%, and 15.18% respectively. In addition, compared with PESA, the average performance improvement of the mixed-precision configurations obtained by GATUNER increased by 18.67%. Jinchen Xu, Hongru Yang, Bei Zhou 0004 |
APSEC | 2 |
| 2025 | Optimization of General Matrix Multiplication for RISC-V ProcessorsabstractGeneral Matrix Multiplication (GEMM) is a fundamental operation in high-performance computing and artificial intelligence, with its performance directly impacting the overall efficiency of higher-level applications. This paper optimizes GEMM by leveraging the characteristics of the RISC-V architecture, achieving significant performance improvements on the XuanTie C908 and C910 processors. It is also observed that traditional optimization method become a bottleneck for small-scale GEMM performance on RISC-V. To address this, we propose an analytical model-based approach that effectively enhances small-scale GEMM performance. In performance evaluation, we tested square matrices and irregular matrices (derived from convolutional neural networks) to validate the optimized GEMM across various scenarios and assess its performance in realworld applications. Compared to the GEMM routines in popular open-source BLAS (Basic Linear Algebra Subprograms) libraries (OpenBLAS and BLIS), the optimized GEMM achieves average speedups of$1.29 \times$and$2.20 \times$, with maximum speedups of$6.51 \times$and$27.85 \times$on the C908 platform. On the C910 platform, the average speedups are$2.38 \times$and$5.15 \times$, with maximum speedups of$17.15 \times$and$84.16 \times$. Bei Zhou 0004, Jianmin Pang, Fei Li 0045, Hongru Yang, Mengyao Duan, Jinchen Xu |
HPCC | 9 |
| 2025 | Scalable Detection of Floating-Point Errors via Adaptive Parallel Subdomain SearchabstractFloating-point error detection is crucial in numerical computing, particularly for multi-parameter functions where even minor errors can propagate and significantly impact results. The sparse distribution of floating-point errors poses a significant detection challenge, as significant deviations are triggered by only rare inputs. Existing search algorithms face two major limitations: poor scalability for multi-parameter functions and insufficient utilization of floating-point representation characteristics. To address these challenges, we propose SDPS (Scalable Detection via Parallel Subdomain Search), a novel algorithm that combines adaptive domain partitioning with floating-point-specific heuristics. SDPS employs a multi-level error classification system and specialized point generation strategies, supported by efficient parallel processing through dynamic task allocation. Our comprehensive evaluation demonstrates that SDPS significantly outperforms state-of-the-art methods in both detection accuracy and computational efficiency, especially for multi-parameter functions where it effectively addresses the exponential growth of search space that limits existing approaches. Zuoyan Zhang, Shihan Yuan, Hongru Yang, Jie Zhao 0002, Jinchen Xu |
QRS | 5 |
| 2025 | PESA: error sensitivity analysis tool for floating-point computational programs
Mengqi Cui, Jinchen Xu, Yuchang Zhou, Hongru Yang, Liguang Ji, Bei Zhou 0004 |
J. Supercomput. | 2 |
| 2024 | LCOC: A Low-cost Optimizing Compiler for Large Scale Quantum Programs Towards Realistic HardwareabstractQuantum compilers play an essential role in the translation of quantum algorithms on quantum hardware. As the size of quantum programs increases, the problem of extreme compilation overhead becomes apparent. To address this problem, compilation frameworks with high-level intermediate representations such as QCOR have been introduced. However, these frameworks still face significant overhead from qubit mapping and optimizing sequence redundancy, resulting in challenges when compiling circuits to realistic hardware. In our paper, a low-cost optimizing compiler (LCOC) is proposed, which addresses the overhead in qubit mapping and enhanced the optimization capabilities. A qubit mapping optimizer based on the LLVM Polly is introduced in LCOC, which implements an efficient mapping of quantum programs containing specific loop types. While achieving low-cost mapping, it considerably reduces the depth of the circuit. We also implement an optimization sequence generator that leverages the modularity of LLVM and iterative compilation to provide LCOC with a more flexible and efficient optimization sequence. We conduct experiments with 100+ qubit circuits. Experiments show that LCOC is up to 1540× faster compared to Qiskit and 74× faster compared to QCOR in compile time. The average speedup relative to Qiskit and QCOR are 194× and 8.64×, respectively. In addition, it can effectively reduce the circuit depth by 50.32% compared to Qiskit and 83.5% compared to QCOR on average. Jinchen Xu, Qing Mu, Hongru Yang, Xiaodong Ding, Qiming Du, Zheng Shan |
ISPA | 2 |
| 2024 | Arfa: An Agile Regime-Based Floating-Point Optimization Approach for Rounding ErrorsabstractWe introduce a floating-point (FP) error optimization approach called Arfa that partitions the domain D of an FP expression fe into regimes and rewrites fe in each regime where fe shows larger errors. First, Arfa seeks a rewrite substitution fo with lower errors across D, whose error distribution is plotted for effective regime inference. Next, Arfa generates an incomplete set of ordered rewrite candidates within each regime of interest, so that searching for the best rewrite substitutions is performed efficiently. Finally, Arfa selects the best rewrite substitution by inspecting the errors of top ranked rewrite candidates, with enhancing precision also considered. Experiments on 56 FPbench examples and four real-life programs show that Arfa not only reduces the maximum and average errors of fe by 4.73 and 2.08 bits on average (and up to 33 and 16 bits), but also exhibits lower errors, sometimes to a significant degree, than Herbie and NumOpt. Jinchen Xu, Mengqi Cui, Fei Li 0045, Zuoyan Zhang, Hongru Yang, Bei Zhou 0004, Jie Zhao 0002 |
ISSTA | 1 |
| 2024 | A Holistic Approach to Automatic Mixed-Precision Code Generation and Tuning for Affine ProgramsabstractReducing floating-point (FP) precision is used to trade the quality degradation of a numerical program's output for performance, but this optimization coincides with type casting, whose overhead is undisclosed until a mixed-precision code version is generated. This uncertainty enforces the decoupled implementation of mixed-precision code generation and autotuning in prior work. In this paper, we present a holistic approach called PrecTuner that consolidates the mixed-precision code generator and the autotuner by defining one parameter. This parameter is first initialized by some automatically sampled values and used to generate several code variants, with various loop transformations also taken into account. The generated code variants are next profiled to solve a performance model formulated using the aforementioned parameter, possibly under a pre-defined quality degradation budget. The best-performing value of the defined parameter is finally predicted without evaluating all code variants. Experimental results of the PolyBench benchmarks on CPU demonstrate that PrecTuner outperforms LuIs by 3.28× while achieving smaller errors, and we also validate its effectiveness in optimizing a real-life large-scale application. In addition, PrecTuner also obtains a mean speedup of 1.81× and 1.52×-1.73× over Pluto on single- and multi-core CPU, respectively, and 1.71× over PPCG on GPU. Jinchen Xu, Guanghui Song, Bei Zhou 0004, Fei Li 0045, Jiangwei Hao, Jie Zhao 0002 |
PPoPP | 1 |
| 2024 | SCR-LIBM: A Correctly Rounded Elementary Function Library in Double-PrecisionabstractThe MPFR and CR-LIBM math libraries are frequently utilized due to their ability to generate correctly rounded results for all double-precision inputs. However, it is worth noting that MPFR has a slower average performance, while CR-LIBM achieves correct rounding over two iterations, rendering it less stable. In addition, CR-LIBM has a poor performance in handling the worst-case of correct rounding. This paper implements a correctly rounded elementary function library called SCR-LIBM in double-precision, which is stable and efficient. Our key idea is to divide subdomains and use the low-degree Taylor polynomial to approximate the elementary function in each subdomain. We simulate the high-precision representation based on the double–double data format, and use the error-free transformation and Double-double algorithm to control the error in the process of polynomial approximation and output compensation. Our approach ensures that the elementary function is correctly rounded, without the need for redundant iterations. The experimental evaluation shows that the average performance of elementary functions implemented in SCR-LIBM is 8.534 times faster than that of MPFR, and 2.492 times faster than that of CR-LIBM when dealing with the worst-case of correct rounding. What’s more, our SCR-LIBM is more stable than CR-LIBM. Jinchen Xu, Bei Zhou 0004, Jiangwei Hao, Fei Li 0045, Zuoyan Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2024 | Hierarchical search algorithm for error detection in floating-point arithmetic expressions
Zuoyan Zhang, Jinchen Xu, Jiangwei Hao, Haotian He, Bei Zhou 0004 |
J. Supercomput. | 2 |
| 2023 | Eiffel: Inferring Input Ranges of Significant Floating-point Errors via Polynomial ExtrapolationabstractExisting search heuristics used to find input values that result in significant floating-point (FP) errors or small ranges that cover them are accompanied by severe constraints, complicating their implementation and restricting their general applicability. This paper introduces an error analysis tool called Eiffel to infer error-inducing input ranges instead of searching them. Given an FP expression with its domain$\mathcal{D}$, Eiffel first constructs an error data set by sampling values across a smaller domain$\mathcal{R}$and assembles these data into clusters. If more than two clusters are formed, Eiffel derives polynomial curves that best fit the bound coordinates of the error-inducing ranges in$\mathcal{R}$, extrapolating them to infer all target ranges of$\mathcal{D}$and reporting the maximal error. Otherwise, Eiffel simply returns the largest error across$\mathcal{R}$. Experimental results show that Eiffel exhibits a broader applicability than Atomu and$\mathbf{S}^{3}$FP by successfully detecting the errors of all 70 considered benchmarks while the two baselines only report errors for part of them. By taking as input the inferred ranges of Eiffel, Herbie obtains an average accuracy improvement of 3.35 bits and up to 53.3 bits. Zuoyan Zhang, Bei Zhou 0004, Jiangwei Hao, Hongru Yang, Mengqi Cui, Yuchang Zhou, Guanghui Song, Fei Li 0045, Jinchen Xu, Jie Zhao 0002 |
ASE | 9 |
| 2023 | Amplitude transformed quantum convolutional neural network
Shiqin Di, Jinchen Xu, Guoqiang Shu, Congcong Feng, Xiaodong Ding, Zheng Shan |
Appl. Intell. | 2 |
| 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine RelationsabstractLoop tiling and fusion are two essential transformations in optimizing compilers to enhance the data locality of programs. Existing heuristics either perform loop tiling and fusion in a particular order, missing some of their profitable compositions, or execute ad-hoc implementations for domain-specific applications, calling for a generalized and systematic solution in optimizing compilers. In this article, we present a so-called basteln (an abbreviation for backward slicing of tiled loop nests) strategy in polyhedral compilation to better model the interplay between loop tiling and fusion. The basteln strategy first groups loop nests by preserving their parallelism/tilability and next performs rectangular/parallelogram tiling to the output groups that produce data consumed outside the considered program fragment. The memory footprints required by each tile are then computed, from which the upward exposed data are extracted to determine the tile shapes of the remaining fusion groups. Such a tiling mechanism can construct complex tile shapes imposed by the dependences between these groups, which are further merged by a post-tiling fusion algorithm for enhancing data locality without losing the parallelism/tilability of the output groups. The basteln strategy also takes into account the amount of redundant computations and the fusion of independent groups, exhibiting a general applicability. We integrate the basteln strategy into two optimizing compilers, with one a general-purpose optimizer and the other a domain-specific compiler for deploying deep learning models. The experiments are conducted on CPU, GPU, and a deep learning accelerator to demonstrate the effectiveness of the approach for a wide class of application domains, including deep learning, image processing, sparse matrix computation, and linear algebra. In particular, the basteln strategy achieves a mean speedup of 1.8× over cuBLAS/cuDNN and 1.1× over TVM on GPU when used to optimize deep learning models; it also outperforms PPCG and TVM by 11% and 20%, respectively, when generating code for the deep learning accelerator. Jie Zhao 0002, Jinchen Xu, Peng Di, Wang Nie, Yanzhi Yi, Zhen Geng, Renwei Zhang, Bojie Li, Zhiliang Gan, Xuefeng Jin 0004 |
ACM Trans. Comput. Syst. | 2 |
| 2022 | Superblock-based performance optimization for Sunway Math Library on SW26010 many-core processor
Shaozhong Guo, Jiangwei Hao, Yuanyuan Xia, Jinchen Xu |
J. Supercomput. | 5 |
| 2022 | Design of variable precision transcendental function automatic generator
Jiangwei Hao, Jinchen Xu, Shaozhong Guo, Yuanyuan Xia |
J. Supercomput. | 2 |
| 2021 | Error detection of arithmetic expressions
Yuanyuan Xia, Shaozhong Guo, Jiangwei Hao, Jinchen Xu |
J. Supercomput. | 5 |
| 2019 | Open-set human activity recognition based on micro-Doppler signatures
Yang Yang 0045, Chunping Hou, Yue Lang, Dai Guan, Danyang Huang, Jinchen Xu |
Pattern Recognit. | 6 |
| 2019 | Memory latency optimizations for the elementary functions on the Sunway architecture
Bei Zhou 0004, Yongzhong Huang, Jinchen Xu, Shaozhong Guo, Hongyuan Qi |
J. Supercomput. | 3 |
| 2018 | Detection of the maximum error of mathematical functions
Hongyuan Qi, Jinchen Xu, Shaozhong Guo |
J. Supercomput. | 2 |
| 2016 | Code Generation for Distributed-Memory ArchitecturesabstractCompiling for distributed-memory architectures comprise two main phases. The first phase is to determine computation and data composition. In the 1990s, a great deal of work addressed this problem. The second phase is code generation. However, there is still no effective solution to this problem. Existing methods try to generate codes on the basis of computation and data composition. To enhance the performance of generated codes, various communication optimizations are introduced since communication is one of the main factors degrading the performance. These approaches would bring redundant communication data, as they did not optimize communications jointly with code generation. In this paper, we propose a novel code generation technique for distributed-memory architectures. First, we determine the communication sender and receiver by traversing a loop-based tree structure. To support message aggregation, we find the most appropriate point to insert a message. Secondly, we construct the communication set by proposing some code generation rules, and prove their correctness and accuracy. Redundant communication is thus eliminated. Also, we have evaluated some programs ranging from micro-kernels to applications in NAS parallel benchmarks, and have compared the performance with their message passing interface (MPI), High Performance Fortran (HPF) and Unified Parallel C (UPC) versions. Compared with these versions, our compiler can generate fewer communication points. The generated codes of outperform the HPF and UPC versions and the state-of-the-art, and the average performance can reach 70% of the hand-coded MPI programs. Jie Zhao 0002, Rongcai Zhao, Jinchen Xu |
Comput. J. | 3 |