VLDB 2026 Research / reviewers in the wild / expert
Haipeng Jia
dblp:94/964
· DBLP profile ↗
25ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue SolversabstractKrylov subspace methods are widely used in scientific computing to solve large sparse linear systems and eigenvalue problems. Their performance bottleneck is often dominated by high-order matrix-power kernels (MPK), especially in polynomial preconditioners that must scale to millions or billions of variables. We present Diagonal Block MPK (DBMPK), a lightweight and parallel-friendly optimization that partitions the input matrix into diagonal blocks and off-diagonal regions. This design enables efficient intra-block data reuse and eliminates inter-block dependencies. It improves cache locality, parallelism, and reduces preprocessing overheads, compared to existing techniques. Our evaluation on x86 and Arm HPC platforms shows that DBMPK improves MPK performance by 26.6%-38.4%. When applied to polynomial preconditioners for linear systems and eigenvalue problems, it achieves consistent end-to-end speedups of 18.6%-34.0%, including in weak scaling tests on 128 nodes, demonstrating strong scalability and practical impact. Xiaojian Yang, Yuhui Ni, Shengguo Li, Dezun Dong, Chuanfu Xu, Haipeng Jia, Jie Liu 0002 |
PPoPP | 7 |
| 2026 | CGA: Accelerating BFS Through an Sparsity-Aware Adaptive Framework on Heterogeneous PlatformsabstractDirection optimization determines whether to use Sparse Matrix-Sparse Vector Multiplication (SpMSpV) or Sparse Matrix-Dense Vector Multiplication (SpMV) based on the input vector's sparsity at each iteration of Breadth-First Search (BFS), aiming to achieve the fastest graph traversal. Although prior work on direction optimization has achieved state-of-the-art performance on either CPUs or GPUs, it has not fully leveraged the capabilities of modern heterogeneous platforms. This is because SpMSpV/SpMV execution times on GPUs do not consistently outperform those on CPUs, particularly for SpMSpV. In response, this paper introducesCGA, a machine learning-based adaptive framework for BFS that optimally selects betweenCPU andGPU kernels, effectivelyAdapting to diverse real-world graphs, vectors, and computing platforms. Our contributions include a novel set of bucket-based SpMSpV algorithms that significantly enhance kernel performance inhigh-sparsity scenarios, along with a low-overhead decision tree model and reduced CPU-GPU data transfers. Experimental results show that our framework outperforms previous state-of-the-art methods, achieving up to a 4.91x speedup over CPU-only baseline and 3.27x speedup over GPU-only baseline. Lei Xu 0023, Haipeng Jia, Yunquan Zhang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | HAM-SpMSpV: an Optimized Parallel Algorithm for Masked Sparse Matrix-Sparse Vector Multiplications on multi-core CPUsabstractThe efficiency of Sparse Matrix-Sparse Vector Multiplication (SpM-SpV) is critically important in fields such as machine learning and graph analytics. In certain algorithms, masked SpMSpV computes only a subset of the result entries. Despite its significance, this selective computation poses unique challenges, and existing algorithms often struggle to exploit the sparsity of the input and the mask vectors concurrently. To boost the efficiency of masked SpMSpV on shared memory architectures, we introduce a hybrid adaptive masked SpMSpV algorithm (HAM-SpMSpV) designed to select the efficient kernel automatically based on input features. This approach builds upon the foundation of a conventional algorithm, incorporating two novel masked SpMSpVs: the pre-bucketing masked SPA-based algorithm and the pre-masking bucketed hash-based algorithm. The newly proposed algorithms significantly expedite computation, especially in scenarios with high sparsity in input vectors and masks. Our evaluation involved extensive testing across a diverse range of real-world graphs, utilizing various sparsity of input vectors and masks. This rigorous testing confirmed that our approach notably outperforms existing solutions. Specifically, it achieves a speedup of up to 1.96 times compared to SuiteSparse:GraphBLAS and a remarkable 6.28 times relative to MKL Graph, demonstrating significant advancements in SpMSpV efficiency. Lei Xu 0023, Haipeng Jia, Yunquan Zhang, Xianmeng Jiang |
HPDC | 2 |
| 2024 | VNEC: A Vectorized Non-Empty Column Format for SpMV on CPUsabstractSparse matrix-vector multiplication (SpMV) is a widely used computational kernel for many applications. The performance of existing vectorization-oriented and locality-optimized SpMV works is limited by increasing additional memory accesses to the output vector or using expensive gather operations. To address these issues, we present the Vectorized Non-Empty Column (VNEC), a novel SpMV storage format aiming to optimize locality and vectorization while alleviating the existing limitations. The VNEC chunks the sparse matrix by rows and removes the empty columns from each row block to improve input vector locality and reduce extra output vector memory access. It can also relieve the cost of expensive gather operations by padding zeros and employing less costly vector load instruction. Specifically, we design two variants of VNEC for different non-zero distributions and propose an effective heuristic selection model by introducing the Intra-Row Density (IRD) to evaluate which variant is suitable for optimizing a given matrix. Experimental results show that in a multicore environment, VNEC achieves up to 6.94× speedup (2.10× on average) against the standard MKL SpMV routine on the x86 CPU and up to 5.92× speedup (1.73× on average) over ArmPL on the ARM CPU. We emphasize that the VNEC format is practical for real-world iterative applications because of its low preprocessing overhead for format conversion. Haipeng Jia, Lei Xu 0023, Cunyang Wei, Kun Li 0016, Xianmeng Jiang, Yunquan Zhang |
IPDPS | 2 |
| 2024 | OpenFFT-SME: An Efficient Outer Product Pattern FFT Library on ARM SME CPUsabstractFast Fourier transform (FFT) is widely used in scientific and engineering computation. Recently developed matrix computation units for AI and high-performance computing provide new optimization opportunities for the FFT algorithm. Compared to dedicated matrix multiplication architectures like Intel AMXs, ARM’s Scalable Matrix Extension (SME) provides more flexible outer product instructions to construct matrix multiplications in software. To leverage ARM SME’s matrix multiplication capabilities, this paper proposes a novel optimized outer product pattern for the Cooley-Tukey FFT algorithm and presents OpenFFT-SME, the first FFT library for ARM SME based on this pattern. This pattern can reduce the number of outer product operations and memory accesses to the DFT matrix by exploiting the symmetric and periodic properties of twiddle factors in the DFT matrix. To further boost performance, OpenFFT-SME incorporates software pipelining to enhance execution pipelines in the assembly code kernels. Meanwhile, a butterfly network more suitable for this pattern is designed and integrated. Experiments demonstrate that OpenFFT-SME outperforms vectorization methods on ARM SME CPUs, achieving 3.60x (power of two) and 4.14x speedups (non-power of two) in double-precision compared to FFTW, 2.47x (power of two) and 3.21x (non-power of two) speedups in double-precision compared to FFTE, and speedups of 4.38x (power of two) and 7.02x (non-power of two) in single-precision compared to FFTW. Furthermore, we compare the advantages and disadvantages of our implementation against vectorization methods and analyze its performance characteristics through additional experiments. Ruge Zhang, Haipeng Jia, Yunquan Zhang, Baicheng Yan, Penghao Ma |
IPDPS | 2 |
| 2024 | Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation SystemsabstractTransportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease. Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | IrGEMM: An Input-Aware Tuning Framework for Irregular GEMM on ARM and X86 CPUsabstractThe matrix multiplication algorithm is a fundamental numerical technique in linear algebra and plays a crucial role in many scientific computing applications. Despite the high performance of mainstream basic linear algebra libraries for large-scale dense matrix multiplications, they exhibit poor performance when applied to matrix multiplication with irregular input. This paper proposes an input-aware tuning framework that accounts for application scenarios and computer architectures to provide high-performance irregular matrix multiplication on ARMv8 and X86 CPUs. The framework comprises two stages: the install-time stage and the run-time stage. The install-time stage utilizes our proposed computational template to generate high-performance kernels for general data layout and SIMD-friendly data layout. The run-time stage utilizes a tiling algorithm suitable for irregular GEMM to select the optimal kernel and link as an execution plan. Additionally, load-balanced multi-threaded optimization algorithms are defined to exploit the multi-threading capability of modern processors. Experiments demonstrate that the proposed IrGEMM framework can achieve significant performance improvements for irregular GEMM on both ARMv8 and X86 CPUs compared to other mainstream BLAS libraries. Cunyang Wei, Haipeng Jia, Yunquan Zhang, Jianyu Yao, Chendi Li, Wenxuan Cao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | SA_TRSM: A Shape-Aware Auto-Tuning Framework for Small-Scale Irregular-Shaped TRSMabstractTRSM (Triangular Solve with Matrix) is an algorithm in the BLAS library for efficiently solving systems of linear equations, which is widely used in scientific computing, engineering computing, and machine learning. The traditional TRSM algorithm performs well in solving large-scale converging squareshaped matrices but is inefficient in solving small-scale irregularshaped matrices. In this paper, we propose SATRSM, a Shape-Aware auto-tuning framework that is aware of scale size and irregularity, aiming to improve performance on small-scale irregular-shaped TRSM computations. SA TRSM consists of the install-time stage and the run-time stage. In the install-time stage, we designed five components for generating high-performance kernels. In the run-time stage, we designed the Shape-Aware tiling algorithm and Plan Generator for generating an efficient execution plan. The experimental results show that the average performance of SA TRSM in this paper improves by 29.4,16.1,24.6 times, and 7.8 times on double-precision real, single-precision real, doubleprecision complex, and single-precision complex in turn, relative to the algorithms in MKL. Rongyuan Guo, Haipeng Jia, Yunquan Zhang, Mingsen Deng, Cunyang Wei, Wenbin Chang |
ICPADS | 2 |
| 2023 | OpenFFT: An Adaptive Tuning Framework for 3D FFT on ARM Multicore CPUsabstractThe sophisticated hierarchy and shared characteristics of cache in multicore CPU architectures bring challenges to the performance improvement of fundamental algorithms, especially in implementing and optimizing 3D FFT. 3D FFT is a memory-bounded algorithm that contains many highly discretized memory accesses. With the working set scaling, the data locality becomes poor, which is prone to cause serious memory access overhead, especially for high-dimensional data transposition. This paper proposes a 3D FFT optimization framework named OpenFFT. This framework optimizes the memory access of 3D FFT by the following methods, including 1) A novel tiling algorithm, Z-OpenFFT, based on the column-order algorithm for high-dimensional vectorization to improve data locality and eliminate transposition; 2) An efficient search algorithm Section-cache-aware algorithm to optimize the memory access of butterfly network of 1D FFT; 3) A multi-thread allocation model by analyzing the characteristics of cache hierarchy and task size to allocate threads adaptively. Experiments demonstrate that OpenFFT could obtain a more competitive performance than the best configuration of FFTW and ARMPL on ARM CPUs. Tun Chen, Haipeng Jia, Yunquan Zhang, Kun Li 0016, Zhihao Li 0001, Jianyu Yao, Chendi Li |
ICS | 2 |
| 2023 | Generating Fast FFT Kernels on CPUs via FFT-Specific IntrinsicsabstractThis paper proposes an algorithm-specific instruction (ASI)-based fast Fourier transform (FFT) code generation framework, named FFTASI, to generate unified architecture independent butterfly kernels that can be transformed into architecture-dependent kernels by establishing the mapping between ASIs and architecture-specific instructions for various hardware platforms. FFTASI strikes a good balance between performance and productivity on CPUs. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Yuyan Sun, Yiwei Zhang 0009, Tun Chen |
PPoPP | 2 |
| 2022 | IATF: An Input-Aware Tuning Framework for Compact BLAS Based on ARMv8 CPUsabstractRecently the mainstream basic linear algebra libraries have delivered high performance on large scale General Matrix Multiplication(GEMM) and Triangular System Solve(TRSM). However, these libraries are still insufficient to provide sustained performance for batch operations on large groups of fixed-size small matrices on specific architectures, which are extensively used in various scientific computing applications. In this paper, we propose IATF, an input-aware tuning framework for optimizing large group of fixed-size small GEMM and TRSM to boost near-optimal performance on ARMv8 architecture. The IATF contains two stages: install-time stage and run-time stage. In the install-time stage, based on SIMD-friendly data layout, we propose computing kernel templates for high-performance GEMM and TRSM, analyze optimal kernel sizes to increase computational instruction ratio, and design kernel optimization strategies to improve kernel execution efficiency. Furthermore, an optimized data packing strategy is also presented for computing kernels to minimize the cost of memory accessing overhead. In the run-time stage, we present an input-aware tuning method to generate an efficient execution plan for large group of fixed-size small GEMM and TRSM, according to the input matrix properties. The experimental results show that IATF could achieve significant performance improvements in GEMM and TRSM compared with other mainstream BLAS libraries. Cunyang Wei, Haipeng Jia, Yunquan Zhang, Liusha Xu |
ICPP | 2 |
| 2022 | Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 4 |
| 2022 | Publisher Correction: Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | IAAT: A Input-Aware Adaptive Tuning framework for Small GEMMabstractGEMM with the small size of input matrices is becoming widely used in many fields like HPC and machine learning. Although many famous BLAS libraries already supported small GEMM, they cannot achieve near-optimal performance. This is because the costs of pack operations are high and frequent boundary processing cannot be neglected. This paper proposes an input-aware adaptive tuning framework(IAAT) for small GEMM to overcome the performance bottlenecks in state-of-the-art implementations. IAAT consists of two stages, the install-time stage and the run-time stage. In the run-time stage, IAAT tiles matrices into blocks to alleviate boundary processing. This stage utilizes an input-aware adaptive tile algorithm and plays the role of runtime tuning. In the install-time stage, IAAT auto-generates hundreds of kernels of different sizes to remove pack operations. Finally, IAAT finishes the computation of small GEMM by invoking different kernels, which corresponds to the size of blocks. The experimental results show that IAAT gains better performance than other BLAS libraries on ARMv8 platform. Jianyu Yao, Boqian Shi, Chunyang Xiang, Haipeng Jia, Chendi Li, Yunquan Zhang |
ICPADS | 4 |
| 2020 | Automatic Generation of High-Performance FFT Kernels on Arm and X86 CPUsabstractThis article presents AutoFFT, a template-based code generation framework that can automatically generate high-performance FFT kernels for all natural-number radices. AutoFFT is based on the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix, as the outer parallelization framework. Because butterflies are the core operations of the Cooley-Tukey algorithm, we explore additional symmetric and periodic properties of the DFT matrix and formulate multiple optimized calculation templates to further reduce the number of floating-point operations for butterflies of arbitrary natural numbers. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these calculation templates and converts them into efficient assembly kernels using the template optimizer. Through a series of experiments on Arm, Intel, and AMD processors, we show that AutoFFT-generated kernels can outperform those in Fastest Fourier Transform in the West (FFTW), the Arm Performance Libraries (ARMPL), and the Intel Math Kernel Library (MKL). Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Richard W. Vuduc |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | AutoFFT: a template-based FFT codes auto-generation framework for ARM and X86 CPUsabstractThe discrete Fourier transform (DFT) is widely used in scientific and engineering computation. This paper proposes a template-based code generation framework named AutoFFT that can automatically generate high-performance fast Fourier transform (FFT) codes. AutoFFT employs the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix as the outer parallelization framework. To further reduce the number of floating-point operations of butterflies, we explore more symmetric and periodic properties of the DFT matrix and formulate two optimized calculation templates for prime and power-of-two radices. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these two templates and transfers them to efficient assembly codes using the template optimizer. Experiments show that AutoFFT outperforms FFTW, ARMPL, and Intel MKL on average across all FFT types on ARMv8 and Intel x86-64 processors. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Luning Cao |
SC | 2 |
| 2019 | Efficient parallel optimizations of a high-performance SIFT on GPUs
Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Shice Liu, Shigang Li 0002 |
J. Parallel Distributed Comput. | 2 |
| 2018 | Implementation and Optimization of Multi-dimensional Real FFT on ARMv8 Platform
Haipeng Jia, Zhihao Li 0001, Yunquan Zhang |
ICA3PP (2) | 2 |
| 2017 | HartSift: A High-Accuracy and Real-Time SIFT Based on GPUabstractScale Invariant Feature Transform (SIFT) is one of the most popular and robust feature extraction algorithms for its invariance to scale, rotation and illumination. It has been widely adopted in many fields, such as video tracking, image stitching, simultaneous localization and mapping (SLAM), structure from motion (SFM) and so on. However, high computational complexity constrains its further application in real-time systems. These systems have to make a tradeoff between accuracy and performance to achieve real-time feature extraction. They adopt other faster algorithms but with less accuracy, like SURF and PCA-SIFT. In order to address this problem, this paper proposes a GPU-accelerated SIFT using CUDA, named HartSift, which realizes high-accuracy and real-time feature extraction by making full use of computing resources of CPU and GPU within a single machine. Experiments show that, on the NIVDIA GTX TITAN Black GPU, HartSift can process an image within 3.14?10.57ms (94.61?318.47fps) according to the size of images. In addition, HartSift is 59.34?75.96 times and 4.01?6.49 times faster than OpenCV-SIFT (a CPU version) and SiftGPU (a GPU version), respectively. In the mean time, HartSift's performance and CudaSIFT's (the fastest GPU version so far) are almost the same, while HartSift's accuracy is much higher than CudaSIFT's. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang |
ICPADS | 2 |
| 2016 | Parallel Processing Systems for Big Data: A SurveyabstractThe volume, variety, and velocity properties of big data and the valuable information it contains have motivated the investigation of many new parallel data processing systems in addition to the approaches using traditional database management systems (DBMSs). MapReduce pioneered this paradigm change and rapidly became the primary big data processing system for its simplicity, scalability, and fine-grain fault tolerance. However, compared with DBMSs, MapReduce also arouses controversy in processing efficiency, low-level abstraction, and rigid dataflow. Inspired by MapReduce, nowadays the big data systems are blooming. Some of them follow MapReduce's idea, but with more flexible models for general-purpose usage. Some absorb the advantages of DBMSs with higher abstraction. There are also specific systems for certain applications, such as machine learning and stream data processing. To explore new research opportunities and assist users in selecting suitable processing systems for specific applications, this survey paper will give a high-level overview of the existing parallel data processing systems categorized by the data input as batch processing, stream processing, graph processing, and machine learning processing and introduce representative projects in each category. As the pioneer, the original MapReduce system, as well as its active variants and extensions on dataflow, data access, parameter tuning, communication, and energy optimizations will be discussed at first. System benchmarks and open issues for big data processing will also be studied in this survey. Yunquan Zhang, Shigang Li 0002, Xinhui Tian, Haipeng Jia, Athanasios V. Vasilakos |
Proc. IEEE | 6 |
| 2015 | Optimizing Image Sharpening Algorithm on GPUabstractSharpness is an algorithm used to sharpen images. As the increase of image size, resolution, and the requirements for real-time processing, the performance of sharpness needs to get improved greatly. The independent pixel calculation of sharpness makes a good opportunity to use GPU to largely accelerate the performance. However, to transplant it to GPU, one challenge is that sharpness involves several stages to execute. Each stage has its own characteristics, either with or without data dependency to other stages. Based on those characteristics, this paper proposes a complete solution to implement and optimize sharpness on GPU. Our solution includes five major and effective techniques: Data Transfer Optimization, Kernel Fusion, Vectorization for Data Locality, Border and Reduction Optimization. Experiments show that, compared to a well-optimized CPU version, our GPU solution can reach 10.7~ 69.3 times speedup for different image sizes on an AMD Fire Pro W8000 GPU. Mengran Fan, Haipeng Jia, Yunquan Zhang, Xiaojing An |
ICPP | 2 |
| 2013 | MPFFT: An Auto-Tuning FFT Library for OpenCL GPUs
Yan Li 0005, Yunquan Zhang, Yiqung Liu 0005, Guoping Long, Haipeng Jia |
J. Comput. Sci. Technol. | 5 |
| 2012 | GPURoofline: A Model for Guiding Performance Optimizations on GPUs
Haipeng Jia, Yunquan Zhang, Guoping Long, Jianliang Xu, Shengen Yan, Yan Li 0005 |
Euro-Par | 1 |
| 2012 | An Insightful Program Performance Tuning Chain for GPU Computing
Haipeng Jia, Yunquan Zhang, Guoping Long, Shengen Yan |
ICA3PP (1) | 1 |
| 2011 | Automatic FFT Performance Tuning on OpenCL GPUsabstractMany fields of science and engineering, such as astronomy, medical imaging, seismology and spectroscopy, have been revolutionized by Fourier methods. The fast Fourier transform (FFT) is an efficient algorithm to compute the discrete Fourier transform (DFT) and its inverse. The emerging class of high performance computing architectures, such as GPU, seeks to achieve much higher performance and efficiency by exposing a hierarchy of distinct memories to programmers. However, the complexity of GPU programming poses a significant challenge for programmers. In this paper, based on the Kronecker product form multi-dimensional FFTs, we propose an automatic performance tuning framework for various OpenCL GPUs. Several key techniques of GPU programming on AMD and NVIDIA GPUs are also identified. Our OpenCL FFT library achieves up to 1.5 to 4 times, 1.5 to 40 times and 1.4 times the performance of clAmdFft 1.0 for 1D, 2D and 3D FFT respectively on an AMD GPU, and the overall performance is within 90% of CUFFT 4.0 on two NVIDIA GPUs. Yan Li 0005, Yunquan Zhang, Haipeng Jia, Guoping Long |
ICPADS | 3 |