EDBT 2026 Demo / reviewers in the wild / expert
Jianbin Fang
dblp:49/8377
· DBLP profile ↗
71ranked-venue papers
12as first author
42since 2021 · last 2026
0000-0003-3542-4869ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 57 · 8 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRASP: Optimizing VLIW Instruction Scheduling via Graph Reinforcement LearningabstractVery Long Instruction Word (VLIW) processors can expose substantial instruction-level parallelism (ILP), yet compiler-generated schedules often fail to fully utilize available functional units, leaving performance on the table and forcing costly manual tuning. We present GRASP (Graph Reinforcement Assembly Scheduling Platform), an assembly-level post-pass optimizer that improves VLIW instruction scheduling and packing via graph reinforcement learning. GRASP represents each basic block as an Instruction Dependency Graph (IDG) that explicitly encodes dependence, latency, and machine resource constraints, and trains a GNN-based reinforcement learning agent to iteratively select legal scheduling decisions and construct high-throughput issue packets. By operating after compilation, GRASP requires no source-code changes and does not modify compiler internals, making it easy to integrate into existing toolchains. Across a suite of AI and HPC kernels, GRASP accelerates execution by up to 1.38 × (1.22 × on average), narrowing the gap between compiler output and expert-tuned VLIW code. Weiyuan Tong, Jianbin Fang, Wei Wang 0056, Jie Ren 0007, Zhanyong Tang |
ICS | 3 |
| 2026 | Demystifying ARM SME to Optimize General Matrix Multiplications
Chencheng Deng, Weiling Yang, Jianbin Fang, Dezun Dong |
IPDPS | 3 |
| 2026 | Optimizing small matrix multiplications via batch grouping on multi-core DSPs
Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 3 |
| 2026 | High-performance matrix multiplication micro-kernel generation for GPDSPs via MLIR progressive lowering
Kainan Yu, Peng Zhang 0061, Chun Huang 0006, Jianbin Fang |
CCF Trans. High Perform. Comput. | 5 |
| 2026 | Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core ProcessorabstractTransformer-based large language models (LLM) are increasingly deployed in high-performance computing environments, where the attention mechanism often becomes a key bottleneck during inference. Although state-of-the-art attention algorithms (e.g., FlashAttention) achieve high efficiency on GPUs, they are ill-suited to emerging heterogeneous many-core processors. In this work, we focus on MT-3000, a representative architecture deployed in the new-generation Tianhe supercomputer, and identify three principal challenges in realizing high-performance attention: complex multi-tier memory requiring manual data movement, excessive reduction overhead caused by sub-tile softmax operations, and static execution pipelines that fail to adapt to inference phases and sequence lengths. To overcome these challenges, we propose DeferAttention , a high-performance attention implementation designed for the MT-3000 many-core processor. DeferAttention introduces a novel deferred-reduction attention strategy to decouple reduction from the fused compute pipeline, enabling more efficient aggregation over large tiles. Moreover, DeferAttention adopts a memory-centric operator design, including data tiling, multi-level software pipelining, and modular micro-kernels, to maximize data reuse and execution throughput. Finally, to support runtime-adaptive execution, DeferAttention integrates a lightweight kernel selection strategy guided by an analytical cost model. Experimental results show that DeferAttention achieves up to 98% of the theoretical peak at the micro-kernel level and 85% at the operator level, outperforming baseline implementations and significantly accelerating end-to-end inference. Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che |
ACM Trans. Archit. Code Optim. | 2 |
| 2026 | Towards Efficient Symmetric Sparse Matrix-Vector Multiplication on Multi-CoresabstractExploiting matrix symmetry to halve memory footprint offers a substantial opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently, i.e., either are non-scalable or yield poor performance for large high-bandwidth irregular matrices. This article extends DCS-SpMV , a D ivide-and- C onquer (DC) based shared-memory implementation of S ymmetric SpMV. The key idea of DCS-SpMV is to recursively divide and reorder the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. The DC approach naturally transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution DCH-SpMV that executes these two parts using DCS-SpMV and the standard SpMV, respectively. We also develop a machine learning model for DCH-SpMV to predict the optimal number of DC recursions on a given matrix and architecture. In this work, we further optimize the hybrid DC implementation by reducing data conflicts before the DC preprocessing. First, we present a conflict-pruning strategy to decouple certain highly dense columns or rows from the conflict graph of a symmetric matrix. Second, we implement a heuristic to adaptively select the lower or upper triangular part of a symmetric matrix, leading to fewer data conflicts. Our optimizations not only facilitate the DC preprocessing, but also improve the performance of DCH-SpMV. We evaluate our work on both x86 and ARM multi-core CPUs using 298 symmetric sparse matrices from the SuiteSparse Matrix Collection. Our new optimizations improve the performance of previous version [ 42 ] by up to 4.89×, demonstrating significant speedup over the state-of-the-art approaches including the vendor-tuned Intel oneMKL library. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | mtGEMM: An Efficient GEMM Library for Modern Multi-Core DSPsabstractThe General Matrix Multiplication (GEMM) is a crucial subprogram in high-performance computing (HPC). With the increasing importance of power and energy consumption, modern Digital Signal Processors (DSPs) are being integrated into general-purpose HPC systems. However, due to architecture disparities, traditional optimizations for CPUs and GPUs are not easily applicable to modern DSPs. This paper shares our experience of optimizing the GEMM operation using a CPU-DSP platform as a case study. Our work employs a set of strategies to improve the performance and scalability of GEMM. These strategies focus on developing micro-kernels based on heterogeneous on-chip memory, addressing the memory access bottleneck in multi-core parallelism, and facilitating efficient transpose-GEMM. These approaches, collectively referred to as an efficient and practical library (a.k.a.mtGEMM), maximize computational capabilities and bandwidth utilization of multi-core DSPs, while achieving high performance for variously-shaped GEMMs. Our experimental results demonstrate thatmtGEMMcan attain between 92% and 96% of the hardware peak, with the multi-core scalability being almost linear. Jianbin Fang, Kainan Yu, Peng Zhang 0061, Dezun Dong, Xinxin Qi, Xingyu Hou, Ruibo Wang, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Me-MPK: Accelerating Krylov Subspace Solvers via Memory-efficient Matrix-Power KernelabstractThis paper focuses on optimizing the Matrix-Power Kernel (MPK), which relies on a series of Sparse Matrix-Vector multiplications (SpMVs) using the same sparse matrix. MPK is a crucial component of Krylov subspace methods for solving large sparse linear systems in various fields, including circuit simulations. MPK offers a potential for matrix reuse in cache, which can accelerate memory-bound sparse solvers. Additionally, many sparse matrices encountered in applications are symmetric, allowing us to reduce the memory footprint for SpMVs by half. However, reusing the matrix introduces data dependencies between subsequent SpMVs, and symmetric SpMVs can result in data conflicts during shared-memory parallelization. Previous research has often focused on either matrix reuse or symmetry, failing to leverage both aspects effectively. This paper proposes a unified, memory-efficient approach called Me-MPK that takes advantage of both cache reuse and matrix symmetry for MPK on shared-memory multi-core systems. We first introduce a unified dependency graph for a sparse matrix, which represents all potential data dependencies and conflicts. Next, we perform architecture-aware recursive partitioning on this graph to create subgraphs and formulate a separating subgraph that decouples all dependencies and conflicts among the subgraphs. These independent subgraphs are then scheduled for parallel execution of SpMV or symmetric SpMV in a specified order to optimize cache reuse. We apply Me-MPK in two s-Step Krylov subspace solvers, and our evaluations show that Me-MPK significantly outperforms the current state-of-the-art solutions, delivering an average speedup of up to 2.00X and 1.86X on X86 and ARM CPUs, respectively. As a result, we achieve overall speedup in the sparse solvers of up to $\mathbf{1. 6 5 X}$ and $\mathbf{1. 5 8 X}$. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Shengguo Li, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002 |
DAC | 3 |
| 2025 | Selection of Supervised Learning-Based Sparse Matrix Reordering Algorithms
Tao Tang 0001, Youfu Jiang, Yingbo Cui 0001, Jianbin Fang, Peng Zhang 0061, Lin Peng 0001, Chun Huang 0006 |
HiPC | 4 |
| 2025 | Optimizing Direct Convolutions on High-Performance Multi-Core DSPsabstractConvolution operations form the computational backbone of deep learning inference but often become performance bottlenecks on conventional architectures. While multi-core Digital Signal Processors (DSPs) offer energy-efficient alternatives through long vector units and software-managed memory hierarchies, existing convolution optimizations designed for CPUs/GPUs underperform due to architectural mismatches in memory systems and execution pipelines. We present mtConv, an optimizing convolution method for multi-core DSPs. mtConv achieves high performance by exploiting data reuse, managing on-chip memory, and designing efficient micro-kernels. It maximizes the overlap between computation and communication to hide data transfer latency, leveraging DSPs’ long vector units and hierarchical scratchpad memories. We evaluate mtConv against state-of-the-art convolution optimizations on DSPs. Experimental results show that mtConv delivers the best overall performance across various convolution layers, achieving up to 93.25% of the hardware’s peak performance on a single DSP core and 92.31% when using all 8 cores of a single DSP cluster. Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Chun Huang 0006, Jie Ren 0007 |
ICPP | 3 |
| 2025 | Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core ProcessorabstractOptimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries. Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Jie Ren 0007 |
SC | 2 |
| 2025 | An empirical performance evaluation of SYCL on ARM multi-core processors
Hanzheng Liang, Chencheng Deng, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 3 |
| 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph ProcessingabstractGraph attention networks (GATs) have advanced performance in various application domains by introducing the attention mechanism into the graph neural networks (GNNs). The inefficiency of running GATs on CPUs or GPUs necessitates specialized hardware designs. Unfortunately, previous specialized architecture designs have focused on either the GNN architecture or the attention mechanism, resulting in limited performance and leaving ample room for improvement. This article presents Gator , a joint optimization approach with software–hardware co-designs for GAT inference. On the software level, Gator leverages degree-weighted graph partitioning and parameter-adaptive feature selection techniques to preprocess the input graph data, mining subgraph-level parallelism and mitigating the computation bottleneck of the dedicated dataflow. On the hardware level, Gator designs a unified processing engine to support various kernels by extracting a common computation pattern and a dimension-aware microarchitecture for efficient partial sum reduction. Extensive experiments show that our approach can achieve 11.5× more efficiency compared to NVIDIA RTX 4090 and provide a speedup of 3× to 9.4×, along with a 2.6× to 4.7× reduction in memory traffic, when compared to six state-of-the-art methods, with minimal accuracy loss. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zixiao Yu |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | DCSolver: Accelerating Sparse Iterative Solvers via Divide-and-Conquer on GPUsabstractSparse iterative solvers are commonly used in various fields. However, certain essential kernels of these solvers, such as sparse triangular solves (SpTRSV), present significant challenges for efficient parallelization due to data dependencies . Previous methods, like level-scheduling or multi-coloring, typically involve creating a Task Dependency Graph (TDG) to represent data dependencies and identify independent sets from the TDG for parallel execution. However, these approaches often result in limited parallelism with substantial synchronization overheads or negatively impact the solver convergence rate. This article introduces DCSolver , a Divide-and-Conquer (DC) framework designed to efficiently parallelize sparse solvers with data dependencies on GPUs. To achieve this, we break down the solver TDG into independent subgraphs, allowing us to exploit both coarse-grained and fine-grained parallelism. To efficiently allocate GPU threads for subgraphs with varying degrees of parallelism, we have developed an adaptive in-warp scheduling strategy. Additionally, we propose a hybrid parallelization scheme in DCSolver, which involves employing different parallel approaches for different DC recursions to achieve a more optimal balance between parallelism and convergence for solvers. To evaluate the effectiveness of DCSolver, we apply it to two preconditioned Krylov subspace solvers and an unstructured mesh Computational Fluid Dynamics (CFD) solver. Our results show that when compared with the state-of-the-art methods, DCSolver accelerates the time-to-solution of solvers by an average speedup of up to 26.19X. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUsabstractConvolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007 |
IEEE Trans. Computers | 3 |
| 2024 | Optimizing SpMV on Heterogeneous Multi-Core DSPs through Improved Locality and VectorizationabstractThe sparse matrix-vector multiplication (SpMV) is widely used in large-scale scientific computing and engineering. However, optimizing SpMV for high-performance digital signal processors (DSPs) has received limited attention. We present HaLAV, a method to accelerate SpMV on CPU-DSP heterogeneous platforms, using the FT-M7032 DSP platform as a case study. HaLAV partitions the input matrix into ‘dense’ and ‘sparse’ parts through column reordering. For the dense part, HaLAV automatically selects storage formats optimized for vectorization to run on the DSP. At the same time, it offloads the sparse component to be processed by the CPU using the standard CSR algorithm. We evaluate our approach on the FT-M7032 platform and an Intel Xeon CPU. Experimental results show that our techniques achieve average speedups of 2.09 × and 1.66 × over the competing baselines on the FT-M7032 and the Xeon platform, respectively. Deshun Bi, Shengguo Li, Dezun Dong, Peng Zhang 0061, Jianbin Fang |
ICPP | 5 |
| 2024 | Optimizing Stencil Computation on Multi-core DSPsabstractStencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart. Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007 |
ICPP | 4 |
| 2024 | Optimizing General Matrix Multiplications on Modern Multi-core DSPsabstractGeneral Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware. Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001 |
IPDPS | 4 |
| 2024 | GraphCube: Interconnection Hierarchy-aware Graph ProcessingabstractProcessing large-scale graphs with billions to trillions of edges requires efficiently utilizing parallel systems. However, current graph processing engines do not scale well beyond a few tens of computing nodes because they are oblivious to the communication cost variations across the interconnection hierarchy. We introduce GraphCube, a better approach to optimizing graph processing on large-scale parallel systems with complex interconnections. GraphCube features a new graph partitioning approach to achieve better load balancing and minimize communication overhead across multiple levels of the interconnection hierarchy. We evaluate GraphCube by applying it to fundamental graph operations performed on synthetic and real-world graph datasets. Our evaluation used up to 79,024 computing nodes and 1.2+ million processor cores. Our large-scale experiments show that GraphCube outperforms state-of-the-art parallel graph processing methods in throughput and scalability. Furthermore, GraphCube outperformed the top-ranked systems on the Graph 500 list. Xinbiao Gan, Shenghao Qiu, Jiaqi Si, Jianbin Fang, Dezun Dong, Chunye Gong, Zheng Wang 0001 |
PPoPP | 6 |
| 2024 | Towards Scalable Unstructured Mesh Computations on Shared Memory Many-CoresabstractDue to data conflicts or data dependences, exploiting shared memory parallelism on unstructured mesh applications is highly challenging. The prior approaches are neither general nor scalable on emerging many-core processors. This paper presents a general and scalable shared memory approach for unstructured mesh computations. We recursively divide and reorder an unstructured mesh to construct a task dependency tree (TDT), where massive parallelism is exposed and data conflicts as well as data dependences are respected. We propose two recursion strategies to support popular programming models on both CPUs and GPUs for TDT. We evaluate our approach by applying it to an industrial unstructured Computational Fluid Dynamics (CFD) software. Experimental results show that our approach significantly outperforms the prior shared memory approaches, delivering up to 8.1× performance improvement over the engineer-tuned implementations. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Yonggang Che, Shizhao Chen, Jie Liu 0002 |
PPoPP | 3 |
| 2024 | A Conflict-aware Divide-and-Conquer Algorithm for Symmetric Sparse Matrix-Vector MultiplicationabstractExploiting matrix symmetry to halve memory footprint offers an opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently. This paper proposes DCS-SpMV, a Divide-and-Conquer (DC) algorithm for efficient Symmetric SpMV. The key idea is to recursively divide the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. Our DC algorithm transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution that executes these two parts using DCS-SpMV and traditional SpMV respectively. We develop a machine learning model to predict an optimal hybrid implementation for a given matrix and architecture. We evaluate our work on both X86 and ARM CPUs, demonstrating significant performance improvement over the state-of-the-art. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Shizhao Chen, Yonggang Che, Jie Liu 0002 |
SC | 3 |
| 2024 | Enhancing Compiler Optimization with Reinforcement Learning and Monte Carlo Tree SearchabstractReinforcement learning (RL) has shown promising performance in compiler optimization, demonstrating significant capabilities across various frameworks such as AutoPhase and CompilerGym.Despite its great success in compiler optimization tasks, challenges remain in training efficiency and optimization performance.This paper presents a novel approach that integrates reinforcement learning with Monte Carlo Tree Search (MCTS) to improve the optimization performance of the traditional Proximal Policy Optimization (PPO) algorithm.We also employ a lightweight search algorithm to reduce training time while maintaining comparable optimization performance.Experimental results show that our PPO-MCTS-Centor model achieves an average performance improvement of 4% over the traditional PPO algorithm.In terms of training efficiency, our PPO-Light model maintains 98% of the performance while reducing training costs by around 16%.Our approach offers a promising solution to enhance compiler optimization through improved efficiency and effectiveness. Jingxiang Ren, Jing Xie 0023, Jianbin Fang, Ting Wang 0009 |
SEKE | 4 |
| 2024 | Editorial for the special issue on programming models and system software for High-Performance Computing (HPC) environments
Jianbin Fang, Jidong Zhai |
CCF Trans. High Perform. Comput. | 1 |
| 2024 | Efficient compiler optimization by modeling passes dependenceabstractAbstract Selecting the optimal combination of compiler passes is a significant challenge to enhance performance and reduce the code size of compiled binaries. While a well-selected sequence of compiler passes can yield considerable benefits, the large number of potential combinations and the scarcity of effective ones make this task prohibitively complex. To tackle this problem, we propose a novel approach to group compiler passes into a small set of sub-sequences. This approach translates the task of identifying the right compiler passes combination into determining the appropriate combination of these sub-sequences. We apply our approach to CBench and PolyBench, demonstrating remarkable performance improvements. Our approach enhances runtime performance by 22% compared to the default LLVM ‘O3’ option, and achieves a code size reduction of 24% compared to the ‘Oz’ option. Our approach also outperforms state-of-the-art across various optimization tasks and hardware platforms. Jianbin Fang, Ting Wang 0009, Jing Xie 0023, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | thSORT: an efficient parallel sorting algorithm on multi-core DSPs
Mouzhi Yang, Peng Zhang 0061, Jianbin Fang, Weifeng Liu 0002, Chun Huang 0006 |
CCF Trans. High Perform. Comput. | 3 |
| 2024 | Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise ProductabstractSparse-dense matrix multiplication (SpMM) is the performance bottleneck of many high-performance and deep-learning applications, making it attractive to design specialized SpMM hardware accelerators. Unfortunately, existing hardware solutions do not take full advantage of data reuse opportunities of the input and output matrices or suffer from irregular memory access patterns. Their strategies increase the off-chip memory traffic and bandwidth pressure, leaving much room for improvement. We present Mentor , a new approach to designing SpMM accelerators. Our key insight is that column-wise dataflow, while rarely exploited in prior works, can address these issues in SpMM computations. Mentor is a software-hardware co-design approach for leveraging column-wise dataflow to improve data reuse and regular memory accesses of SpMM. On the software level, Mentor incorporates a novel streaming construction scheme to preprocess the input matrix for enabling a streaming access pattern. On the hardware level, it employs a fully pipelined design to unlock the potential of column-wise dataflow further. The design of Mentor is underpinned by a carefully designed analytical model to find the tradeoff between performance and hardware resources. We have implemented an FPGA prototype of Mentor . Experimental results show that Mentor achieves speedup by geomean 2.05× (up to 3.98×), reduces the memory traffic by geomean 2.92× (up to 4.93×), and improves bandwidth utilization by geomean 1.38× (up to 2.89×), compared with the state-of-the-art hardware solutions. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zidong Du, Yongwei Zhao 0001, Zheng Wang 0079 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 5 |
| 2024 | Optimizing Full-Spectrum Matrix Multiplications on ARMv8 Multi-Core CPUsabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. While the mainstream Basic Linear Algebra Subprograms (BLAS) libraries can deliver good performance on large and regular-shaped GEMMs, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in emerging HPC applications. Recent research has focused on improving GEMM performance on GPUs, but there is still significant room for improvement on emerging HPC hardware based on multi-core CPUs. We presentLibShalom2, an open-source library to optimize full-spectrum GEMMs, taking small, irregular-shaped, and large-scale regular-shaped matrices.LibShalom2explicitly targets the ARMv8 architecture, which is becoming common in HPC systems.LibShalom2is designed to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluateLibShalom2by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show thatLibShalom2consistently outperforms existing solutions across full-spectrum GEMM workloads and hardware architectures. We also show thatLibShalom2delivers an average speedup of 2.2x for real-life neural network workloads. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0079 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Optimizing HPC I/O Performance with Regression Analysis and Ensemble LearningabstractTo improve parallel I/O performance, it is imperative to optimize the adjustable parameters across the different layers of the I/O software stack. Finding an optimal configuration for different scenarios is hampered by the complex interaction dynamics between these parameters and the large parameter space. Previous research efforts have focused on tuning these parameters using independent algorithms; however, these approaches exhibit certain shortcomings such as unstable performance results and delayed convergence rates.This paper introduces OPRAEL, an auto-tuning approach on parallel I/O tasks by ensembles and performance modeling using regression analysis. To test its effectiveness, we applied this approach on the Tianhe-II supercomputer using one well-known I/O benchmark(IOR) and two I/O kernels(S3D-I/O, BT-I/O). Leveraging our experience in predictive modeling, we optimized the tuning of the I/O stack parameters. Our experimental results show a remarkable 10.2X improvement in write performance speedup for the optimization task with BT-I/O and a 500x500x500 input. We also compared the potential of using a single search algorithm versus using reinforcement learning search in the I/O parameter auto-optimization task. Our results show that OPRAEL outperforms the traditional approach, resulting in a maximum 8.4X improvement in write performance for the 128-process IOR optimization. Zhangyu Liu, Cheng Zhang 0007, Jianbin Fang, Lin Peng 0001, Guixin Ye, Zhanyong Tang |
CLUSTER | 4 |
| 2023 | Optimizing MPI Collectives on Shared Memory Multi-CoresabstractMessage Passing Interface (MPI) programs often experience performance slowdowns due to collective communication operations, like broadcasting and reductions. As modern CPUs integrate more processor cores, running multiple MPI processes on shared-memory machines to take advantage of hardware parallelism is becoming increasingly common. In this context, it is crucial to optimize MPI collective communications for shared-memory execution. However, existing MPI collective implementations on shared-memory systems have two primary drawbacks. The first is extensive redundant data movements when performing reduction collectives, and the second is the ineffective use of non-temporal instructions to optimize streamed data processing. To address these limitations, this paper proposes two optimization techniques that minimize data movements and enhance the use of non-temporal instructions. We evaluated our techniques by integrating them into the OpenMPI library and tested their performance using micro-benchmarks and real-world applications running on two multi-core clusters. Experimental results show that our approach significantly outperforms existing techniques, yielding a 1.2--6.4x performance improvement. Jintao Peng, Jianbin Fang, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Zheng Wang 0001 |
SC | 2 |
| 2023 | Optimizing Direct Convolutions on ARM Multi-CoresabstractConvolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms. Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001 |
SC | 3 |
| 2023 | wrBench: Comparing Cache Architectures and Coherency Protocols on ARMv8 Many-Core Systems
Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000abstractAs the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system. Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | PipeFB: An Optimized Pipeline Parallelism Scheme to Reduce the Peak Memory Usage
Sheng Ma, Xiang Hou, Libo Huang 0002, Jianbin Fang |
ICA3PP | 7 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 7 |
| 2022 | FlowDNN: a physics-informed deep neural network for fast and accurate flow predictionabstractfor flow-related design optimization problems, e.g., aircraft and automobile aerodynamic design, computational fluid dynamics (CFD) simulations are commonly used to predict flow fields and analyze performance. While important, CFD simulations are a resource-demanding and time-consuming iterative process. The expensive simulation overhead limits the opportunities for large design space exploration and prevents interactive design. In this paper, we propose FlowDNN, a novel deep neural network (DNN) to efficiently learn flow representations from CFD results. FlowDNN saves computational time by directly predicting the expected flow fields based on given flow conditions and geometry shapes. FlowDNN is the first DNN that incorporates the underlying physical conservation laws of fluid dynamics with a carefully designed attention mechanism for steady flow prediction. This approach not only improves the prediction accuracy, but also preserves the physical consistency of the predicted flow fields, which is essential for CFD. Various metrics are derived to evaluate FlowDNN with respect to the whole flow fields or regions of interest (RoIs) (e.g., boundary layers where flow quantities change rapidly). Experiments show that FlowDNN significantly outperforms alternative methods with faster inference and more accurate results. It speeds up a graphics processing unit (GPU) accelerated CFD solver by more than 14 000×, while keeping the prediction error under 5%. Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Siqi Wang 0001, Shizhao Chen, Jianbin Fang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2021 | Optimizing Barrier Synchronization on ARMv8 Many-Core ArchitecturesabstractSynchronization operations are commonly seen in OpenMP programs where a parallel construct often works with an explicit or implicit barrier operation. While OpenMP synchronization has been extensively studied on the traditional x86 CPU architectures, there is little work on understanding OpenMP barrier synchronization operations on ARMv8 high-performance many-cores. This paper presents the first comprehensive performance study on OpenMP barrier implementations on emerging ARMvS-based many-cores. We evaluate seven representative barrier algorithms on three distinct ARMv8 architectures: Phytium 2000+, ThunderX2, and Kunpeng920. We empirically show that the existing synchronization implementations exhibit poor scalability on ARMv8 architectures compared to the x86 counterpart. We then propose various optimization strategies for improving these widely used synchronization algorithms on each platform. We showcase that our optimizations yield 12.6x performance improvement over the GCC implementation and 4.7x improvement over the LLVM implementation, translating to 1.6x improvement over the state-of-the-art best-performing algorithm. We share our experience and practical insights on optimizing OpenMP synchronization operations on emerging ARMv8 multi-core CPU architectures. Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001 |
CLUSTER | 2 |
| 2021 | Characterizing Small-Scale Matrix Multiplications on ARMv8-based Many-Core ArchitecturesabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. There is a large body of work on evaluating and optimizing large-scale matrix multiplication, but how well the small-scale matrix multiplication (SMM) performs is largely unknown, especially for the ARMv8-based many-core architectures. In this work, we evaluate and characterize the performance of SMM subroutines on Phytium 2000 +, an ARMv8-based 64-core architecture. The evaluation work is extensively performed with the mainstream open-source libraries including OpenBLAS, BLIS, BALSFEO, and Eigen. Given various experimental settings, we observe how well the small-scale GEMM routines perform on Phytium 2000 +, and then discuss the impacting factors behind the performance behaviours of SMM. Built on such a basis, we shed light on the performance bottlenecks and practical optimizations on SMM from various angles: (1) mitigating the data packing overhead, (2) processing the edge cases properly, (3) selecting a suitable micro-kernel, and (4) adopting a right parallelization method. The result of our work facilitates users to develop efficient SMM optimizations on ARMv8-based many-core architectures, and embed them into real-world applications. Weiling Yang, Jianbin Fang, Dezun Dong |
IPDPS | 2 |
| 2021 | LIBSHALOM: optimizing small and irregular-shaped matrix multiplications on ARMv8 multi-coresabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in highperformance computing. While the mainstream linear algebra libraries can deliver high performance on large and regular-shaped GEMM, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in new HPC applications. Some of the recent works in this direction have made promising progress on x86 architectures and GPUs but still leave much room for improvement on emerging HPC hardware built upon the ARMv8 architecture. We present LibShalom, an open-source library for optimizing small and irregular-shaped GEMMs, explicitly targeting the ARMv8 architecture. LibShalom builds upon the classical Goto algorithm but tailors it to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluate LibShalom by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show that LibShalom can consistently outperform existing solutions across GEMM workloads and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0001 |
SC | 2 |
| 2021 | Performance Evaluation of Memory-Centric ARMv8 Many-Core Architectures: A Case Study with Phytium 2000+
Jianbin Fang, Xiangke Liao, Chun Huang 0006, Dezun Dong |
J. Comput. Sci. Technol. | 1 |
| 2021 | BALS: Blocked Alternating Least Squares for Parallel Sparse Matrix Factorization on GPUsabstractMatrix factorization on sparse matrices has been proven to be an effective approach for data mining and machine learning. However, the prior parallel implementations for matrix factorization fail to capture the internal social property embedded in real-world use cases. This article presents an efficient implementation of the alternative least squares (ALS) algorithm calledBALSbuilt on top of a new sparse matrix format for parallel matrix factorization. The BALS storage format organizes the sparse matrix into 2D tiles to avoid repeated data loads and improve data reuses. We further propose a data reordering technique to sort sparse matrices according to nonzeros. The experimental results show that BALS can yield a superior performance than state-of-the-art implementations, i.e., our BALS generally runs faster than Gates’ implementation over different latent feature sizes, with a speedup of up to 2.08× on K20C, 3.72× on TITAN X and 3.13× on TITAN RTX. When compared with alternative matrix factorization algorithms, our BALS consistently outperforms CDMF, cuMF_CCD, and cuMF_SGD over various latent feature sizes and datasets. The reordering technique can provide an extra improvement of up to 23.68 percent on K20C, 19.87 percent on TITAN X and 20.38 percent on TITAN RTX. Jing Chen 0038, Jianbin Fang, Weifeng Liu 0002, Canqun Yang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Deep Program Structure Modeling Through Multi-Relational Graph-based LearningabstractDeep learning is emerging as a promising technique for building predictive models to support code-related tasks like performance optimization and code vulnerability detection. One of the critical aspects of building a successful predictive model is having the right representation to characterize the model input for the given task. Existing approaches in the area typically treat the program structure as a sequential sequence but fail to capitalize on the rich semantics of data and control flow information, for which graphs are a proven representation structure. Guixin Ye, Zhanyong Tang, Huanting Wang, Dingyi Fang, Jianbin Fang, Songfang Huang, Zheng Wang 0001 |
PACT | 5 |
| 2020 | FlowGAN: A Conditional Generative Adversarial Network for Flow Prediction in Various ConditionsabstractMany flow-related design optimization problems like aircraft and automobile aerodynamic design are solved via computational fluid dynamics (CFD) simulations. However, CFD simulations are known to be resource-demanding and time-consuming. Deep learning (DL) is emerging as a viable means to accelerate CFD simulations by directly predicting the outcomes of multiple simulation iterations. While promising, existing DL-based models have to be re-trained whenever the flow condition changes, which incurs significant training overhead for real-life scenarios with a wide range of flow conditions. This paper presents FLOWGAN, a novel conditional generative adversarial network for accurate prediction of flow fields in various conditions. FlowGAN is designed to directly obtain the generation of solutions to flow fields in various conditions based on observations rather than re-training. As FlowGAN does not rely on knowledge of the underlying governing equations, it can quickly adapt to various flow conditions and avoid the need for expensive re-training. We evaluate FlowGAN by applying it to scenarios of simulating both the whole flow field and selected regions of interest (RoI). Compared to the state-of-the-art DL based methods, FlowGAN significantly reduces the prediction errors by 2.27% while exhibiting a better generalization ability. Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Shizhao Chen, Jianbin Fang, Zhenghua Wang, Zheng Wang 0001 |
ICTAI | 5 |
| 2020 | NUMA-Aware Optimization of Sparse Matrix-Vector Multiplication on ARMv8-Based Many-Core Architectures
Xiaosong Yu, Huihui Ma, Zhengyu Qu, Jianbin Fang, Weifeng Liu 0002 |
NPC | 4 |
| 2020 | Parallel programming models for heterogeneous many-cores: a comprehensive survey
Jianbin Fang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
CCF Trans. High Perform. Comput. | 1 |
| 2020 | clMF: A fine-grained and portable alternating least squares algorithm for parallel matrix factorization
Jing Chen 0038, Jianbin Fang, Weifeng Liu 0002, Tao Tang 0001, Canqun Yang |
Future Gener. Comput. Syst. | 2 |
| 2020 | Deep Learning Research and Development Platform: Characterizing and Scheduling with QoS Guarantees on GPU ClustersabstractDeep learning (DL) has been widely adopted in various domains of artificial intelligence (AI), achieving dramatic developments in industry and academia. Besides giant AI companies, numerous small and medium-sized enterprises, institutes, and universities (EIUs) have focused on the research and development (R&D) of DL. Considering the high cost of datacenters and high performance computing (HPC) systems, EIUs prefer adopting off-the-shelf GPU clusters as a DL R&D platform for multiple users and developers to process diverse DL workloads. In such scenarios, the scheduling of multiple DL tasks on a shared GPU cluster is both significant and challenging in terms of efficiently utilizing limited resources. Existing schedulers cannot predict the resource requirements of diverse DL workloads, leading to the under-utilization of computing resources and a decline in user satisfaction. This paper proposes GENIE, a QoS-aware dynamic scheduling framework for a shared GPU cluster, which achieves users' QoS guarantee and high system utilization. In accordance with an exhaustive characterization, GENIE analyzes the key factors that affect the performance of DL tasks and proposes a prediction model derived from lightweight profiling to estimate the processing rate and response latency for diverse DL workloads. Based on the prediction models, we propose a QoS-aware scheduling algorithm to identify the best placements for DL tasks and schedule them on the shared cluster. Experiments on a GPU cluster and large-scale simulations demonstrate that GENIE achieves a QoS-guarantee percentage improvement of up to 67.4 percent and a makespan reduction of up to 28.2 percent, compared to other baseline schedulers. Zhaoyun Chen, Wei Quan 0004, Mei Wen, Jianbin Fang, Jie Yu 0008, Chunyuan Zhang, Lei Luo 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Optimizing Streaming Parallelism on Heterogeneous Many-Core ArchitecturesabstractAs many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Toward fault-tolerant hybrid programming over large-scale heterogeneous clusters via checkpointing/restart optimization
Cheng Chen 0005, Yunfei Du 0001, Ke Zuo, Jianbin Fang, Canqun Yang |
J. Supercomput. | 4 |
| 2018 | MOCL: an efficient openCL implementation for the matrix-2000 architectureabstractThis paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture. Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001 |
CF | 3 |
| 2018 | Proteus: network-aware web browsing on heterogeneous mobile systemsabstractWe present Proteus, a novel network-aware approach for optimizing web browsing on heterogeneous multi-core mobile systems. It employs machine learning techniques to predict which of the heterogeneous cores to use to render a given webpage and the operating frequencies of the processors. It achieves this by first learning offline a set of predictive models for a range of typical networking environments. A learnt model is then chosen at runtime to predict the optimal processor configuration, based on the web content, the network status and the optimization goal. We evaluate Proteus by implementing it into the open-source Chromium browser and testing it on two representative ARM big.LITTLE mobile multi-core platforms. We apply Proteus to the top 1,000 popular websites across seven typical network environments. Proteus achieves over 80% of best available performance. It obtains, on average, over 17% (up to 63%), 31% (up to 88%), and 30% (up to 91%) improvement respectively for load time, energy consumption and the energy delay product, when compared to two state-of-the-art approaches. Jie Ren 0007, Jianbin Fang, Yansong Feng 0002, Dongxiao Zhu, Zhunchen Luo, Jie Zheng 0005, Zheng Wang 0001 |
CoNEXT | 3 |
| 2018 | Auto-tuning Streamed Applications on Intel Xeon PhiabstractMany-core accelerators, as represented by the XeonPhi coprocessors and GPGPUs, allow software to exploit spatial and temporal sharing of computing resources to improve the overall system performance. To unlock this performance potential requires software to effectively partition the hardware resource to maximize the overlap between host-device communication and accelerator computation, and to match the granularity of task parallelism to the resource partition. However, determining the right resource partition and task parallelism on a per program, per dataset basis is challenging. This is because the number of possible solutions is huge, and the benefit of choosing the right solution may be large, but mistakes can seriously hurt the performance. In this paper, we present an automatic approach to determine the hardware resource partition and the task granularity for any given streamed application, targeting the Intel XeonPhi architecture. Instead of hand-crafting the heuristic for which the process will have to repeat for each hardware generation, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs; we then use the learned model to predict the resource partition and task granularity for any unseen programs at runtime. We apply our approach to 23 representative parallel applications and evaluate it on a CPU-XeonPhi mixed heterogenous many-core platform. Our approach achieves, on average, a 1.6x (upto 5.6x) speedup, which translates to 94.5% of the performance delivered by a theoretically perfect predictor. Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zheng Wang 0001 |
IPDPS | 2 |
| 2018 | Moving from exascale to zettascale computing: challenges and techniquesabstractHigh-performance computing (HPC) is essential for both traditional and emerging scientific fields, enabling scientific activities to make progress. With the development of high-performance computing, it is foreseeable that exascale computing will be put into practice around 2020. As Moore’s law approaches its limit, high-performance computing will face severe challenges when moving from exascale to zettascale, making the next 10 years after 2020 a vital period to develop key HPC techniques. In this study, we discuss the challenges of enabling zettascale computing with respect to both hardware and software. We then present a perspective of future HPC technology evolution and revolution, leading to our main recommendations in support of zettascale computing in the coming future. Xiangke Liao, Kai Lu 0001, Canqun Yang, Jin-wen Li, Yuan Yuan 0034, Libo Huang 0002, Pingjing Lu, Jianbin Fang, Jie Shen 0003 |
Frontiers Inf. Technol. Electron. Eng. | 9 |
| 2018 | Orchestrating parallel detection of strongly connected components on GPUs
Xuhao Chen 0001, Cheng Chen 0005, Jie Shen 0003, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zhiying Wang 0003 |
Parallel Comput. | 4 |
| 2018 | Benchmarking the GPU memory at the warp level
Minquan Fang, Jianbin Fang, Haifang Zhou, Jianxing Liao, Yuangang Wang |
Parallel Comput. | 2 |
| 2017 | Efficient and high-quality sparse graph coloring on GPUsabstractSummary Graph coloring has been broadly used to discover concurrency in parallel computing. To speed up graph coloring for large‐scale datasets, parallel algorithms have been proposed to leverage modern GPUs. Existing GPU implementations either have limited performance or yield unsatisfactory coloring quality (too many colors assigned). We present a work‐efficient parallel graph coloring implementation on GPUs with good coloring quality. Our approach uses the speculative greedy scheme, which inherently yields better quality than the method of finding maximal independent set . To achieve high performance on GPUs, we refine the algorithm to leverage efficient operators and alleviate conflicts. We also incorporate common optimization techniques to further improve performance. Our method is evaluated with both synthetic and real‐world sparse graphs on the NVIDIA GPU. Experimental results show that our proposed implementation achieves averaged 4.1 × (up to 8.9 × ) speedup over the serial implementation. It also outperforms the existing GPU implementation from the NVIDIA CUSPARSE library (2.2 × average speedup), while yielding much better coloring quality than CUSPARSE. Xuhao Chen 0001, Pingfan Li, Jianbin Fang, Tao Tang 0001, Zhiying Wang 0003, Canqun Yang |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | An Energy-Efficient Implementation of LU Factorization on Heterogeneous SystemsabstractEnergy consumption is increasingly becoming a critical issue in HPC. There is a broad consensus that future exascale-computing will be strongly constrained by energy consumption. Heterogeneous systems usually feature higher energy efficiency than homogeneous ones since the former employ coprocessors that provide higher GFlops/Watt than CPUs. Thus, it is of great importance to better utilize the coprocessors from an energy-efficiency standpoint. Dense LU factorization (LU) is a critical kernel that is widely used to solve dense linear algebra problems. However, existingheterogeneous implementations are typically designed to be CPU-centered, which rely highly on CPUs and thus suffer from large data transfer overheads via PCIe, hurting the energy efficiency of the entire computer system. We present a coprocessor-resident implementation of LU for a heterogeneous platform to improve energy efficiency without impeding performance by relieving the CPUs from performing unnecessary computations and reducing excessive data transfers via PCIe. In addition, several optimizations are judiciously employed to overlap the computation and communication between the CPUs and coprocessors. Validation on the Tianhe-2 supercomputer shows that our LU implementation gains higher performance, achieves higher energy efficiency, and features a better scalability than Intel MKL. Canqun Yang, Cheng Chen 0005, Tao Tang 0001, Xuhao Chen 0001, Jianbin Fang, Jingling Xue |
ICPADS | 5 |
| 2016 | Evaluating Multi-core and Many-Core Architectures through Accelerating an Alternating Direction Implicit CFD SolverabstractIn this paper, we accelerate a double-precision alternating direction implicit (ADI) solver for three-dimensional compressible Navier-Stokes equations from our in-house computational fluid dynamics (CFD) software on the latest multi-core and many-core architectures (Intel Ivy Bridge CPU, Intel Xeon Phi 7110P coprocessor and NVIDIA Kepler K20c GPU). For the GPU platform, both the OpenACC-based and the CUDA-based versions of the ADI solver are developed. To achieve high performance, we use a series of optimization techniques. For the Ivy Bridge CPU and Xeon Phi, we focus on three categories of optimization techniques: thread parallelism for multi-/many-core scaling, data parallelism to exploit the SIMD mechanism and improving on-chip data reuse, to maximize the performance. Also, we provide an in-depth analysis on the performance differences between Ivy Bridge and Xeon Phi. Our numerical experiments show that the proposed CUDA-based ADI solver can achieve a speedup of 9.7 on a Kepler GPU in contrast to a single naive serial version and our optimization techniques can improve the performance of the ADI solver by 2.5x on two Ivy Bridge CPUs and 1.7x on the Intel Xeon Phi coprocessor. We also notice that the OpenACC-based version runs around 29% slower than the CUDA-based one with careful manual optimizations. Besides, we systematically evaluate the programmability of the three platforms. Our insights facilitate the programmers to select a right platform with a suitable programming model according to their target applications. Liang Deng, Jianbin Fang, Fang Wang 0004, Hanli Bai |
ISPDC | 2 |
| 2016 | Streaming Applications on Heterogeneous Platforms
Zhaokui Li, Jianbin Fang, Tao Tang 0001, Xuhao Chen 0001, Canqun Yang |
NPC | 2 |
| 2015 | Realistic Performance Characterization of CFD Applications on Intel Many Integrated Core ArchitectureabstractThis paper studies the performance characteristics of computational fluid dynamics (CFD) applications on Intel Many Integrated Core (MIC) architecture. Three CFD applications, BT-MZ, LM3D and HOSTA, are evaluated on Intel Knights Corner (KNC) coprocessor, the first public MIC product. The results show that the pure OpenMP scalability of these applications is not sufficient to utilize the potential of a KNC coprocessor. While utilizing the hybrid MPI/OpenMP programming model helps to improve the parallel scalability, the maximum parallel speedup relative to a single thread is still not satisfactory. The OpenCL version of BT-MZ performs better than the OpenMP version but is not comparable to the MPI version and the hybrid MPI/OpenMP version. At the micro-architecture level, while the three CFD applications achieve reasonable instruction execution rates and L1 data cache hit rates, use a large percent of vector instructions, they have low arithmetic density, incur very high branch misprediction rates and do not utilize the Vector Processing Unit efficiently. As a result, they achieve very low single thread floating-point efficiency. For these applications to attain competitive performance on the MIC architecture as on the Xeon processors, both the parallel scalability and the single thread performance should be improved, which is a difficult task. Yonggang Che, Chuanfu Xu, Jianbin Fang, Yongxian Wang, Zhenghua Wang |
Comput. J. | 3 |
| 2015 | Evaluating vector data type usage in OpenCL kernelsabstractSummary Open Computing Language (OpenCL) is an open, functionally portable programming model for a large range of highly parallel processors. To provide users with access to the underlying platforms, OpenCL has explicit support for features such as local memory and vector data types (VDTs). However, these are often low‐level, hardware‐specific features, which can be detrimental to performance on different platforms. In this paper, we focus on VDTs and investigate their usage in a systematic way. First, we propose two different approaches (inter‐vdtandintra‐vdt) to use VDTs in OpenCL kernels, and show how to translate scalar OpenCL kernels to vectorized ones. After obtaining vectorized code, we evaluate the performance effects of using VDTs with two types of benchmarks: micro‐benchmarks and macro‐benchmarks. With micro‐benchmarks, we study the execution model of VDTs and the role of the compiler‐aided vectorizer on five devices. With macro‐benchmarks, we explore the changes of memory access patterns before and after using VDTs, and the resulting performance impact. Not only our evaluation provides insights into how OpenCL's VDTs are mapped on different processors, but it also indicates that using such data types introduces changes in both computation and memory accesses. Based on the lessons learned, we discuss how to deal with performance portability in the presence of VDTs. Copyright © 2014 John Wiley & Sons, Ltd. Jianbin Fang, Ana Lucia Varbanescu, Xiangke Liao, Henk J. Sips |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Grover: Looking for Performance Improvement by Disabling Local Memory Usage in OpenCL KernelsabstractDue to the diversity of processor architectures and application memory access patterns, the performance impact of using local memory in OpenCL kernels has become unpredictable. For example, enabling the use of local memory for an OpenCL kernel can be beneficial for the execution on a GPU, but can lead to performance losses when running on a CPU. To address this unpredictability, we propose an empirical approach: by disabling the use of local memory in OpenCL kernels, we enable users to compare the kernel versions with and without local memory, and further choose the best performing version for a given platform. To this end, we have designed Grover, a method to automatically remove local memory usage from OpenCL kernels. In particular, we create a correspondence between the global and local memory spaces, which is used to replace local memory accesses by global memory accesses. We have implemented this scheme in the LLVM framework as a compiling pass, which automatically transforms an OpenCL kernel with local memory to a version without it. We have validated Grover with 11 applications, and found that it can successfully disable local memory usage for all of them. We have compared the kernels with and without local memory on three different processors, and found performance improvements for more than a third of the test cases after Grover disabled local memory usage. We conclude that such a compiler pass can be beneficial for performance, and, because it is fully automated, it can be used as an auto-tuning step for OpenCL kernels. Jianbin Fang, Henk J. Sips, Pekka Jääskeläinen, Ana Lucia Varbanescu |
ICPP | 1 |
| 2014 | Balancing CPU-GPU Collaborative High-Order CFD Simulations on the Tianhe-1A SupercomputerabstractHOSTA is an in-house high-order CFD software that can simulate complex flows with complex geometries. Large scale high-order CFD simulations using HOSTA require massive HPC resources, thus motivating us to port it onto modern GPU accelerated supercomputers like Tianhe-1A. To achieve a greater speedup and fully tap the potential of Tianhe-1A, we collaborate CPU and GPU for HOSTA instead of using a naive GPU-only approach. We present multiple novel techniques to balance the loads between the store-poor GPU and the store-rich CPU, and overlap the collaborative computation and communication as far as possible. Taking CPU and GPU load balance into account, we improve the maximum simulation problem size per Tianhe-1A node for HOSTA by 2.3X, meanwhile the collaborative approach can improve the performance by around 45% compared to the GPU-only approach. Scalability tests show that HOSTA can achieve a parallel efficiency of above 60% on 1024 Tianhe-1A nodes. With our method, we have successfully simulated China's large civil airplane configuration C919 containing 150M grid cells. To our best knowledge, this is the first paper that reports a CPUGPU collaborative high-order accurate aerodynamic simulation result with such a complex grid geometry. Chuanfu Xu, Lilun Zhang, Xiaogang Deng, Jianbin Fang, Guangxue Wang, Yonggang Che, Yongxian Wang, Wei Liu 0013 |
IPDPS | 4 |
| 2014 | Test-driving Intel Xeon PhiabstractBased on Intel's Many Integrated Core (MIC) architecture, Intel Xeon Phi is one of the few truly many-core CPUs - featuring around 60 fairly powerful cores, two levels of caches, and graphic memory, all interconnected by a very fast ring. Given its promised ease-of-use and high performance, we took Xeon Phi out for a test drive. In this paper, we present this experience at two different levels: (1) the microbenchmark level, where we stress "each nut and bolt" of Phi in the lab, and (2) the application level, where we study Phi's performance response in a real-life environment. At the microbenchmarking level, we show the high performance of five components of the architecture, focusing on their maximum achieved performance and the prerequisites to achieve it. Next, we choose a medical imaging application (Leukocyte Tracking) as a case study. We observed that it is rather easy to get functional code and start benchmarking, but the first performance numbers can be far from satisfying. Our experience indicates that a simple data structure and massive parallelism are critical for Xeon Phi to perform well. When compiler-driven parallelization and/or vectorization fails, programming Xeon Phi for performance can become very challenging. Jianbin Fang, Henk J. Sips, Lilun Zhang, Chuanfu Xu, Yonggang Che, Ana Lucia Varbanescu |
ICPE | 1 |
| 2013 | Sesame: A User-Transparent Optimizing Framework for Many-Core ProcessorsabstractWith the integration of more computational cores and deeper memory hierarchies on modern processors, the performance gap between naively parallel zed code and optimized code becomes much larger than ever before. Very often, bridging the gap involves architecture-specific optimizations. These optimizations are difficult to implement by application programmers, who typically focus on the basic functionality of their code. Therefore, in this thesis, I focus on answering the following research question: "How can we address architecture-specific optimizations in a programmer-friendly way?'' As an answer, I propose an optimizing framework for parallel applications running on many-core processors (\textit{Sesame}). Taking a simple parallel zed code provided by the application programmers as input, Sesame chooses and applies the most suitable architecture-specific optimizations, aiming to improve the overall application performance in a user-transparent way. In this short paper, I present the motivation for designing and implementing Sesame, its structure and its modules. Furthermore, I describe the current status of Sesame, discussing our promising results in source-to-source vectorization, automated usage of local memory, and auto-tuning for implementation-specific parameters. Finally, I discuss my work-in-progress and sketch my ideas for finalizing Sesame's development and testing. Jianbin Fang, Ana Lucia Varbanescu, Henk J. Sips |
CCGRID | 1 |
| 2013 | ELMO: A User-Friendly API to Enable Local Memory in OpenCL KernelsabstractRecent parallel architectures are equipped with local memory, which simplifies hardware design at the cost of increased program complexity due to explicit management. To simplify this extra-burden that programmers have, we introduce an easy-to-use API, ELMO, that improves productivity while preserving high performance of local memory operations. Specifically, ELMO is a generic API that covers different local memory use-cases. We also present prototype implementations for these APIs and perform multiple GPU-inspired optimizations to maximize their performance. Experimental results on the NVIDIA Quadro5000 GPU show that performance is significantly improved by using ELMO on native implementations: the achieved speedup ranges from 1.3x to 3.7x. Furthermore, using ELMO we still achieve performance comparable (if not better) with that of hand-tuned applications, while the code is shorter, clearer, and safer. Jianbin Fang, Ana Lucia Varbanescu, Jie Shen 0003, Henk J. Sips |
PDP | 1 |
| 2013 | Performance Traps in OpenCL for CPUsabstractWith its design concept of cross-platform portability, OpenCL can be used not only on GPUs (for which it is quite popular), but also on CPUs. Whether porting GPU programs to CPUs, or simply writing new code for CPUs, using OpenCL brings up the performance issue, usually raised in one of two forms: "OpenCL is not performance portable!" or "Why using OpenCL for CPUs after all?!". We argue that both issues can be addressed by a thorough study of the factors that impact the performance of OpenCL on CPUs. This analysis is the focus of this paper. Specifically, starting from the two main architectural mismatches between many-core CPUs and the OpenCL platform-parallelism granularity and the memory model-we identify eight such performance "traps" that lead to performance degradation in OpenCL for CPUs. Using multiple code examples, from both synthetic and real-life benchmarks, we quantify the impact of these traps, showing how avoiding them can give up to 10 times better performance. Furthermore, we point out that the solutions we provide for avoiding these traps are simple and generic code transformations, which can be easily adopted by either programmers or automated tools. Therefore, we conclude that a certain degree of OpenCL inter-platform performance portability, while indeed not a given, can be achieved by simple and generic code transformations. Jie Shen 0003, Jianbin Fang, Henk J. Sips, Ana Lucia Varbanescu |
PDP | 2 |
| 2013 | An application-centric evaluation of OpenCL on multi-core CPUs
Jie Shen 0003, Jianbin Fang, Henk J. Sips, Ana Lucia Varbanescu |
Parallel Comput. | 2 |
| 2012 | Accelerating Cost Aggregation for Real-Time Stereo MatchingabstractReal-time stereo matching, which is important in many applications like self-driving cars and 3-D scene reconstruction, requires large computation capability and high memory bandwidth. The most time-consuming part of stereo-matching algorithms is the aggregation of information (i.e. costs) over local image regions. In this paper, we present a generic representation and suitable implementations for three commonly used cost aggregators on many-core processors. We perform typical optimizations on the kernels, which leads to significant performance improvement (up to two orders of magnitude). Finally, we present a performance model for the three aggregators to predict the aggregation speed for a given pair of input images on a given architecture. Experimental results validate our model with an acceptable error margin (an average of 10.4%). We conclude that GPU-like many-cores are excellent platforms for accelerating stereo matching. Jianbin Fang, Ana Lucia Varbanescu, Jie Shen 0003, Henk J. Sips, Gorkem Saygili, Laurens van der Maaten |
ICPADS | 1 |
| 2011 | A Comprehensive Performance Comparison of CUDA and OpenCLabstractThis paper presents a comprehensive performance comparison between CUDA and OpenCL. We have selected 16 benchmarks ranging from synthetic applications to real-world ones. We make an extensive analysis of the performance gaps taking into account programming models, ptimization strategies, architectural details, and underlying compilers. Our results show that, for most applications, CUDA performs at most 30% better than OpenCL. We also show that this difference is due to unfair comparisons: in fact, OpenCL can achieve similar performance to CUDA under a fair comparison. Therefore, we define a fair comparison of the two types of applications, providing guidelines for more potential analyses. We also investigate OpenCL's portability by running the benchmarks on other prevailing platforms with minor modifications. Overall, we conclude that OpenCL's portability does not fundamentally affect its performance, and OpenCL can be a good alternative to CUDA. Jianbin Fang, Ana Lucia Varbanescu, Henk J. Sips |
ICPP | 1 |