Yonggang Che

dblp:99/4718 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-6906-4940ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DFA-SpTRSV: A Depth-First Asynchronous Algorithm for Sparse Triangular Solver
Zhimeng Han, Chuanfu Xu, Haozhong Qiu, Yue Ding 0001, Yonggang Che
IPDPS10
2026 Cascaded spectral operator transformer with mixture-of-experts for urban wind field prediction
Jie Li 0002, Xinhai Chen 0001, Yonggang Che, Qingyang Zhang 0009
Eng. Appl. Artif. Intell.7
2026 Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core Processor
abstract
Transformer-based large language models (LLM) are increasingly deployed in high-performance computing environments, where the attention mechanism often becomes a key bottleneck during inference. Although state-of-the-art attention algorithms (e.g., FlashAttention) achieve high efficiency on GPUs, they are ill-suited to emerging heterogeneous many-core processors. In this work, we focus on MT-3000, a representative architecture deployed in the new-generation Tianhe supercomputer, and identify three principal challenges in realizing high-performance attention: complex multi-tier memory requiring manual data movement, excessive reduction overhead caused by sub-tile softmax operations, and static execution pipelines that fail to adapt to inference phases and sequence lengths. To overcome these challenges, we propose DeferAttention , a high-performance attention implementation designed for the MT-3000 many-core processor. DeferAttention introduces a novel deferred-reduction attention strategy to decouple reduction from the fused compute pipeline, enabling more efficient aggregation over large tiles. Moreover, DeferAttention adopts a memory-centric operator design, including data tiling, multi-level software pipelining, and modular micro-kernels, to maximize data reuse and execution throughput. Finally, to support runtime-adaptive execution, DeferAttention integrates a lightweight kernel selection strategy guided by an analytical cost model. Experimental results show that DeferAttention achieves up to 98% of the theoretical peak at the micro-kernel level and 85% at the operator level, outperforming baseline implementations and significantly accelerating end-to-end inference.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che
ACM Trans. Archit. Code Optim.4
2026 Towards Efficient Symmetric Sparse Matrix-Vector Multiplication on Multi-Cores
abstract
Exploiting matrix symmetry to halve memory footprint offers a substantial opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently, i.e., either are non-scalable or yield poor performance for large high-bandwidth irregular matrices. This article extends DCS-SpMV , a D ivide-and- C onquer (DC) based shared-memory implementation of S ymmetric SpMV. The key idea of DCS-SpMV is to recursively divide and reorder the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. The DC approach naturally transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution DCH-SpMV that executes these two parts using DCS-SpMV and the standard SpMV, respectively. We also develop a machine learning model for DCH-SpMV to predict the optimal number of DC recursions on a given matrix and architecture. In this work, we further optimize the hybrid DC implementation by reducing data conflicts before the DC preprocessing. First, we present a conflict-pruning strategy to decouple certain highly dense columns or rows from the conflict graph of a symmetric matrix. Second, we implement a heuristic to adaptively select the lower or upper triangular part of a symmetric matrix, leading to fewer data conflicts. Our optimizations not only facilitate the DC preprocessing, but also improve the performance of DCH-SpMV. We evaluate our work on both x86 and ARM multi-core CPUs using 298 symmetric sparse matrices from the SuiteSparse Matrix Collection. Our new optimizations improve the performance of previous version [ 42 ] by up to 4.89×, demonstrating significant speedup over the state-of-the-art approaches including the vendor-tuned Intel oneMKL library.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che
ACM Trans. Archit. Code Optim.10
2025 Me-MPK: Accelerating Krylov Subspace Solvers via Memory-efficient Matrix-Power Kernel
abstract
This paper focuses on optimizing the Matrix-Power Kernel (MPK), which relies on a series of Sparse Matrix-Vector multiplications (SpMVs) using the same sparse matrix. MPK is a crucial component of Krylov subspace methods for solving large sparse linear systems in various fields, including circuit simulations. MPK offers a potential for matrix reuse in cache, which can accelerate memory-bound sparse solvers. Additionally, many sparse matrices encountered in applications are symmetric, allowing us to reduce the memory footprint for SpMVs by half. However, reusing the matrix introduces data dependencies between subsequent SpMVs, and symmetric SpMVs can result in data conflicts during shared-memory parallelization. Previous research has often focused on either matrix reuse or symmetry, failing to leverage both aspects effectively. This paper proposes a unified, memory-efficient approach called Me-MPK that takes advantage of both cache reuse and matrix symmetry for MPK on shared-memory multi-core systems. We first introduce a unified dependency graph for a sparse matrix, which represents all potential data dependencies and conflicts. Next, we perform architecture-aware recursive partitioning on this graph to create subgraphs and formulate a separating subgraph that decouples all dependencies and conflicts among the subgraphs. These independent subgraphs are then scheduled for parallel execution of SpMV or symmetric SpMV in a specified order to optimize cache reuse. We apply Me-MPK in two s-Step Krylov subspace solvers, and our evaluations show that Me-MPK significantly outperforms the current state-of-the-art solutions, delivering an average speedup of up to 2.00X and 1.86X on X86 and ARM CPUs, respectively. As a result, we achieve overall speedup in the sparse solvers of up to $\mathbf{1. 6 5 X}$ and $\mathbf{1. 5 8 X}$.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Shengguo Li, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002
DAC11
2025 Gated Cross-Attention Network for Depth Completion
abstract
Depth completion is a popular research direction in the field of depth estimation. The fusion of color and depth features is the critical challenge in this task, mainly due to the asymmetry between the rich scene details in color images and the sparse pixels in depth maps. To tackle this issue, we design an efficient Gated Cross-Attention Network that propagates confidence via a gating mechanism, simultaneously extracting and refining key information in both color and depth branches to achieve local spatial feature fusion. Additionally, we incorporate a Transformer-based attention network in low-dimensional space to effectively fuse global features and increase the network’s receptive field. At the same time, we use the Ray Tune mechanism with the AsyncHyperBandScheduler and the HyperOptSearch algorithm to automatically search for the optimal number of module iterations, which also allows us to achieve performance comparable to state-of-the-art methods. We conduct experiments on both indoor and outdoor scene datasets. Our fast network ranked first among real-time methods (below 30ms and 100ms), and our accurate network ranked first among all methods on the KITTI official website at the time of submission.
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang
ICASSP4
2025 Optimizing Direct Convolutions on High-Performance Multi-Core DSPs
abstract
Convolution operations form the computational backbone of deep learning inference but often become performance bottlenecks on conventional architectures. While multi-core Digital Signal Processors (DSPs) offer energy-efficient alternatives through long vector units and software-managed memory hierarchies, existing convolution optimizations designed for CPUs/GPUs underperform due to architectural mismatches in memory systems and execution pipelines. We present mtConv, an optimizing convolution method for multi-core DSPs. mtConv achieves high performance by exploiting data reuse, managing on-chip memory, and designing efficient micro-kernels. It maximizes the overlap between computation and communication to hide data transfer latency, leveraging DSPs’ long vector units and hierarchical scratchpad memories. We evaluate mtConv against state-of-the-art convolution optimizations on DSPs. Experimental results show that mtConv delivers the best overall performance across various convolution layers, achieving up to 93.25% of the hardware’s peak performance on a single DSP core and 92.31% when using all 8 cores of a single DSP cluster.
Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Chun Huang 0006, Jie Ren 0007
ICPP5
2025 Hierarchical Neural Architecture Search for Fast and Accurate Depth Completion
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang, Yu-Lin He
ICMR4
2025 Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core Processor
abstract
Optimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Jie Ren 0007
SC4
2025 DCSolver: Accelerating Sparse Iterative Solvers via Divide-and-Conquer on GPUs
abstract
Sparse iterative solvers are commonly used in various fields. However, certain essential kernels of these solvers, such as sparse triangular solves (SpTRSV), present significant challenges for efficient parallelization due to data dependencies . Previous methods, like level-scheduling or multi-coloring, typically involve creating a Task Dependency Graph (TDG) to represent data dependencies and identify independent sets from the TDG for parallel execution. However, these approaches often result in limited parallelism with substantial synchronization overheads or negatively impact the solver convergence rate. This article introduces DCSolver , a Divide-and-Conquer (DC) framework designed to efficiently parallelize sparse solvers with data dependencies on GPUs. To achieve this, we break down the solver TDG into independent subgraphs, allowing us to exploit both coarse-grained and fine-grained parallelism. To efficiently allocate GPU threads for subgraphs with varying degrees of parallelism, we have developed an adaptive in-warp scheduling strategy. Additionally, we propose a hybrid parallelization scheme in DCSolver, which involves employing different parallel approaches for different DC recursions to achieve a more optimal balance between parallelism and convergence for solvers. To evaluate the effectiveness of DCSolver, we apply it to two preconditioned Krylov subspace solvers and an unstructured mesh Computational Fluid Dynamics (CFD) solver. Our results show that when compared with the state-of-the-art methods, DCSolver accelerates the time-to-solution of solvers by an average speedup of up to 26.19X.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002
ACM Trans. Archit. Code Optim.10
2025 nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs
abstract
Convolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures.
Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007
IEEE Trans. Computers10
2025 Automatic GPU memory access optimization for AoSoA-based application in OP2 framework
Zongjing Chen, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
J. Supercomput.3
2024 MARO: Enabling Full MPI Automatic Refactoring in DSL-Based Programming Framework
Zongjing Chen, Yonggang Che, Chuanfu Xu
ICA3PP (4)3
2024 Optimizing Stencil Computation on Multi-core DSPs
abstract
Stencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart.
Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007
ICPP6
2024 Optimizing General Matrix Multiplications on Modern Multi-core DSPs
abstract
General Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware.
Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001
IPDPS9
2024 Towards Scalable Unstructured Mesh Computations on Shared Memory Many-Cores
abstract
Due to data conflicts or data dependences, exploiting shared memory parallelism on unstructured mesh applications is highly challenging. The prior approaches are neither general nor scalable on emerging many-core processors. This paper presents a general and scalable shared memory approach for unstructured mesh computations. We recursively divide and reorder an unstructured mesh to construct a task dependency tree (TDT), where massive parallelism is exposed and data conflicts as well as data dependences are respected. We propose two recursion strategies to support popular programming models on both CPUs and GPUs for TDT. We evaluate our approach by applying it to an industrial unstructured Computational Fluid Dynamics (CFD) software. Experimental results show that our approach significantly outperforms the prior shared memory approaches, delivering up to 8.1× performance improvement over the engineer-tuned implementations.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Yonggang Che, Shizhao Chen, Jie Liu 0002
PPoPP9
2024 A Conflict-aware Divide-and-Conquer Algorithm for Symmetric Sparse Matrix-Vector Multiplication
abstract
Exploiting matrix symmetry to halve memory footprint offers an opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently. This paper proposes DCS-SpMV, a Divide-and-Conquer (DC) algorithm for efficient Symmetric SpMV. The key idea is to recursively divide the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. Our DC algorithm transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution that executes these two parts using DCS-SpMV and traditional SpMV respectively. We develop a machine learning model to predict an optimal hybrid implementation for a given matrix and architecture. We evaluate our work on both X86 and ARM CPUs, demonstrating significant performance improvement over the state-of-the-art.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Shizhao Chen, Yonggang Che, Jie Liu 0002
SC9
2024 Extending OP2 framework to support portable parallel programming of complex applications
Zongjing Chen, Kangjin Huang, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
CCF Trans. High Perform. Comput.3
2024 Evaluating performance portability of five shared-memory programming models using a high-order unstructured CFD solver
Liang Deng, Yonggang Che, Yueqing Wang
J. Parallel Distributed Comput.3
2024 Improving CUDA performance of an unstructured high-order CFD application under OP2 framework
Kangjin Huang, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
J. Supercomput.2
2023 PowerDis: Fine-Grained Power Monitoring Through Power Disaggregation Model
Xinxin Qi, Juan Chen 0001, Rongyu Deng, Yuan Yuan 0034, Yonggang Che
ICA3PP (4)7
2023 Developing a proxy application for an industrial unstructured CFD software: preliminary results
abstract
As programming models and architectures evolve in the exa-scale era, porting large-scale HPC applications are becoming increasingly difficult and expensive. In HPC community, mini-apps are often developed to mimic real-world applications, and it offers an easy way to benchmark new HPC platforms. In this paper, we design and implement a mini-app MiniFS as a proxy for an industry-level unstructured Computational Fluid Dynamics (CFD) software FlowStar. The main purpose of MiniFS is to evaluate different shared memory approaches on emerging multi/many-core architectures, because data conflicts and data dependencies in unstructured CFD pose tough challenges for shared memory parallelization. Results show that our mini-app can represent the performance characteristic of the original application. However, the existing approaches are unscalable on modern multi/many-cores. It is imperative to develop novel scalable shared memory approaches for unstructured CFD.
Chuanfu Xu, Jian Zhang 0115, Liang Deng, Haozhong Qiu, Weixi Dai, Yongzhen Lin, Yue Ding 0001, Yonggang Che
ICPADS10
2020 Memory Access Optimization of High-Order CFD Stencil Computations on GPU
Shengxiang Wang, Zhuoqian Li, Yonggang Che
PDCAT3
2019 Collaborating CPUs and MICs for Large-Scale LBM Multiphase Flow Simulations
Chuanfu Xu, Dali Li, Yonggang Che, Zhenghua Wang
NPC4
2018 Petascale scramjet combustion simulation on the Tianhe-2 heterogeneous supercomputer
Yonggang Che, Meifang Yang, Chuanfu Xu, Yutong Lu
Parallel Comput.1
2015 Realistic Performance Characterization of CFD Applications on Intel Many Integrated Core Architecture
abstract
This paper studies the performance characteristics of computational fluid dynamics (CFD) applications on Intel Many Integrated Core (MIC) architecture. Three CFD applications, BT-MZ, LM3D and HOSTA, are evaluated on Intel Knights Corner (KNC) coprocessor, the first public MIC product. The results show that the pure OpenMP scalability of these applications is not sufficient to utilize the potential of a KNC coprocessor. While utilizing the hybrid MPI/OpenMP programming model helps to improve the parallel scalability, the maximum parallel speedup relative to a single thread is still not satisfactory. The OpenCL version of BT-MZ performs better than the OpenMP version but is not comparable to the MPI version and the hybrid MPI/OpenMP version. At the micro-architecture level, while the three CFD applications achieve reasonable instruction execution rates and L1 data cache hit rates, use a large percent of vector instructions, they have low arithmetic density, incur very high branch misprediction rates and do not utilize the Vector Processing Unit efficiently. As a result, they achieve very low single thread floating-point efficiency. For these applications to attain competitive performance on the MIC architecture as on the Xeon processors, both the parallel scalability and the single thread performance should be improved, which is a difficult task.
Yonggang Che, Chuanfu Xu, Jianbin Fang, Yongxian Wang, Zhenghua Wang
Comput. J.1
2014 Balancing CPU-GPU Collaborative High-Order CFD Simulations on the Tianhe-1A Supercomputer
abstract
HOSTA is an in-house high-order CFD software that can simulate complex flows with complex geometries. Large scale high-order CFD simulations using HOSTA require massive HPC resources, thus motivating us to port it onto modern GPU accelerated supercomputers like Tianhe-1A. To achieve a greater speedup and fully tap the potential of Tianhe-1A, we collaborate CPU and GPU for HOSTA instead of using a naive GPU-only approach. We present multiple novel techniques to balance the loads between the store-poor GPU and the store-rich CPU, and overlap the collaborative computation and communication as far as possible. Taking CPU and GPU load balance into account, we improve the maximum simulation problem size per Tianhe-1A node for HOSTA by 2.3X, meanwhile the collaborative approach can improve the performance by around 45% compared to the GPU-only approach. Scalability tests show that HOSTA can achieve a parallel efficiency of above 60% on 1024 Tianhe-1A nodes. With our method, we have successfully simulated China's large civil airplane configuration C919 containing 150M grid cells. To our best knowledge, this is the first paper that reports a CPUGPU collaborative high-order accurate aerodynamic simulation result with such a complex grid geometry.
Chuanfu Xu, Lilun Zhang, Xiaogang Deng, Jianbin Fang, Guangxue Wang, Yonggang Che, Yongxian Wang, Wei Liu 0013
IPDPS7
2014 Test-driving Intel Xeon Phi
abstract
Based on Intel's Many Integrated Core (MIC) architecture, Intel Xeon Phi is one of the few truly many-core CPUs - featuring around 60 fairly powerful cores, two levels of caches, and graphic memory, all interconnected by a very fast ring. Given its promised ease-of-use and high performance, we took Xeon Phi out for a test drive. In this paper, we present this experience at two different levels: (1) the microbenchmark level, where we stress "each nut and bolt" of Phi in the lab, and (2) the application level, where we study Phi's performance response in a real-life environment. At the microbenchmarking level, we show the high performance of five components of the architecture, focusing on their maximum achieved performance and the prerequisites to achieve it. Next, we choose a medical imaging application (Leukocyte Tracking) as a case study. We observed that it is rather easy to get functional code and start benchmarking, but the first performance numbers can be far from satisfying. Our experience indicates that a simple data structure and massive parallelism are critical for Xeon Phi to perform well. When compiler-driven parallelization and/or vectorization fails, programming Xeon Phi for performance can become very challenging.
Jianbin Fang, Henk J. Sips, Lilun Zhang, Chuanfu Xu, Yonggang Che, Ana Lucia Varbanescu
ICPE5
2014 Microarchitectural performance comparison of Intel Knights Corner and Intel Sandy Bridge with CFD applications
Yonggang Che, Lilun Zhang, Yongxian Wang, Chuanfu Xu, Wei Liu 0013, Zhenghua Wang
J. Supercomput.1
2009 MPTD: A Scalable and Flexible Performance Prediction Framework for Parallel Systems
Chuanfu Xu, Yonggang Che, Zhenghua Wang
APPT2
2009 A Framework for Effective Memory Optimization of High Performance Computing Applications
abstract
Memory wall is an important factor that influences program performance, and its alleviation relies on memory optimization of the program. Static approaches optimize memory performance based on analytical models that are hard to achieve because of increasing architecture complexity and code structures. Execution-driven approaches like iterative compilation achieve it by executing different versions of the program on actual platforms and select the one that renders best performance, outperforming static compilation approaches significantly. But the expensive compilation cost has limited their application scope to embedded applications and a small group of math kernels. This paper proposes a different approach-Combining Model and Iterative Compilation for Effective Memory Optimization (MICEMemO). Such an approach first constructs a memory optimization model based on hardware performance counters to decide how and when to apply transformations, and then selects the optimal transformation parameters using genetic algorithms. Experimental results show that our performance counter based approach can greatly reduce programs' memory access time and influence ratio for memory reference, improve programs' memory performance, therefore, effectively alleviate the problem of memory wall.
Pingjing Lu, Yonggang Che, Zhenghua Wang
HPCC2
2008 An Effective Iterative Compilation Search Algorithm for High Performance Computing Applications
abstract
The performance gap for high performance applications has been widening over time. High level program transformations are critical to improve the applications' performance, many of which concern the determination of optimal values for transformation parameters, such as loop unrolling and blocking. Traditional compilers select these parameters based on static analytical models. However, complex computer architectures and code behaviors greatly limit the strength of optimizing compilers. Iterative compilation approach determines these parameter values by executing the program with different parameter values and selects the one with the shortest runtime, outperforming static compilation approaches significantly, which makes it a hot research topic in the high performance computing research community. But itpsilas quite time consuming because of the huge optimization space. Therefore, an effective search strategy is crucial for iterative compilation. This paper investigates the Nelder-Mead simplex algorithm for iterative compilation optimization parameter search. Experimental results indicate Nelder-Mead simplex based search strategy can produce parameter values with better performance and lower cost.
Pingjing Lu, Yonggang Che, Zhenghua Wang
HPCC2
2008 Analyzing the Efficiency and Bottleneck of Scientific Programs on Imagine Stream Processor by Simulation
abstract
Imagine stream processor has shown high performance and efficiency for media applications. Its potential for scientific applications is of great interest to the high performance computing community. This paper investigates this subject from a new angle. It roughly classifies the scientific programs into three classes based on their computation to memory access ratios. For each class, typical programs are programmed with StreamC/KernelC stream language and simulated based on the cycle-accurate simulator of Imagine. In-depth analysis is carried out for the performance data, with special attentions on the performance bottlenecks. The performance data obtained on Imagine are compared against data on two general-purpose x86 processors. The results show that programs with no DRAM accesses attain high floating point performance and efficiencies on Imagine. These programs' performance is only restricted by limited ILP (Instruction-Level Parallelism) and load imbalance across ALUs. Programs with computation to memory operation ratios O(n) attain absolute floating point performance on Imagine comparable to that obtained on general-purpose processors, but their floating-point efficiencies are not satisfactory. It is essential to optimize these programs for high SRF (Stream Register File) and LRF (Local Register File) reuse and high ILP on Imagine. Programs with lower computation to memory operation ratios attain much lower floating-point performance and efficiencies on Imagine, compared to those obtained on x86 processors.
Yonggang Che, Chuanfu Xu, Zhenghua Wang
ISPA1
2004 Locality Optimizations for Jacobi Iteration on Distributed Parallel Systems
Yonggang Che, Zhenghua Wang, Laurence T. Yang
ISPA1