Jingwei Sun 0001

dblp:66/7761-1 · DBLP profile ↗
← Back
29ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0001-5098-1503ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints
abstract
Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant overhead from CPU-GPU data transfers. While prior work has explored prefetching techniques to mitigate this bottleneck, these methods require costly fallback mechanisms when predictions fail. Since expert transfers cannot be canceled once initiated, the correct experts need to be loaded on demand sequentially, introducing additional latency. To address this, we present CommitMoE, a novel approach featuring a Commit Router that makes execution decisions based on expert predictions without fallback mechanisms. Our key insight reveals that router certainty strongly correlates with prediction accuracy, while in low-certainty scenarios, the model output demonstrates inherent robustness to expert selection. Leveraging this insight to design a systems-level solution, CommitMoE achieves 1.3× to 9.4× faster inference across different environments and datasets compared to state-of-the-art offloading frameworks while maintaining model quality.
Jingwei Sun 0001, Junqing Lin, Guangzhong Sun
AAAI2
2026 Dual-Verbalizer with Label Correlation Modeling for Few-Shot Multi-Label Text Classification
Zhiyi Tian, Guangzhong Sun, Jingwei Sun 0001
DASFAA (4)4
2026 HOPO: Accelerating Multimodal Neural Networks Inference via Holistic Parallelism Optimization
abstract
Multimodal neural networks (MMNNs) feature multi-branch topologies that offer opportunities for inference parallelism. However, mainstream machine learning compilers prioritize intra-operator parallelism as the primary optimization target. While effective for chain-structured models, this approach inhibits the inter-operator parallelism inherent in the MMNNs, leading to suboptimal inference efficiency. Through a systematic analysis, we identify two fundamental sources of inefficiency: (1) during graph scheduling, a compromise among synchronization overhead, resource contention, and GPU utilization; and (2) during compilation, a structural conflict between greedily maximizing intra-operator parallelism and the severe resource competition induced by concurrent operators.
Jingwei Sun 0001, Guangzhong Sun, Jing Li 0047
ICS2
2026 Enhancing HPC Batch Job Scheduling via Imitation Learning-Based Search
Zechun Zhou, Jingwei Sun 0001, Mingfei Ye, Guangzhong Sun
IPDPS2
2026 ClusterFi: Enabling Concurrent WiFi Backscatter Communication via Collided Signal Clustering
Weiqi Wu, Wei Gong 0001, Jingwei Sun 0001, Guangzhong Sun
IWQoS3
2026 Complexity-Aware and Response Time Enhanced Knowledge Tracing
Wangqian Li, Weiqi Wu, Jingwei Sun 0001, Guangzhong Sun
KSEM (1)3
2026 Fast compiler autotuning framework using design of experiments
Chenghua Xu, Jingwei Sun 0001, Mengna Sai, Fuxin Zhang, Guangzhong Sun, Weiwu Hu
CCF Trans. High Perform. Comput.2
2025 Introducing Graph Context into Language Models through Parameter-Efficient Fine-Tuning for Lexical Relation Mining
abstract
Lexical relation refers to the way words are related within a language. Prior work has demonstrated that pretrained language models (PLMs) can effectively mine lexical relations between word pairs. However, they overlook the potential of graph structures composed of lexical relations, which can be integrated with the semantic knowledge of PLMs. In this work, we propose a parameter-efficient fine-tuning method through graph context, which integrates graph features and semantic representations for lexical relation classification (LRC) and lexical entailment (LE) tasks. Our experiments show that graph features can help PLMs better understand more complex lexical relations, establishing a new state-of-the-art for LRC and LE. Finally, we perform an error analysis, identifying the bottlenecks of language models in lexical relation mining tasks and providing insights for future improvements.
Zhiyi Tian, Jingwei Sun 0001, Guangzhong Sun
ACL (1)4
2025 A Fast Sparse Triangular Solve for Structured-grid Problems on Heterogeneous Processors
abstract
Structured-grid problems are common in scientific computing, particularly in applications like fluid dynamics and electromagnetic simulation. One of the key kernels in solving these problems is Sparse Triangular Solve (SpTRSV), which often becomes a performance bottleneck due to its low computing intensity and inherent internal data dependencies. In structured-grid SpTRSV, the regularity of non-zero distributions and the high parallelism of sparse matrices present opportunities to harness the architectural strengths of modern heterogeneous processors. However, existing SpTRSV algorithms fail to fully exploit these advantages, due to their mismatches in data dependencies, computational order, and memory layouts. In this paper, we introduce a novel SpTRSV algorithm tailored for structured-grids on modern heterogeneous processors. Our approach introduces a two-level blocking strategy to enhance data locality and reduce communication overhead, while a vertical tiling-based pipeline balances parallelism with computational granularity. Additionally, we design hardware-specific adaptive scheduling strategies to accommodate varying degrees of parallelism across distinct architectures. The algorithm has been implemented on two types of heterogeneous processors, NVIDIA GPUs and SW26010-Pro, with hardware-specific optimizations to further improve the performance. Experimental results show that our implementations achieve speedups of more than 1.87 × over state-of-the-art baselines and provide efficient end-to-end solutions with lightweight preprocessing.
Zhengding Hu, Yi Zong, Jingwei Sun 0001, Wei Xue 0003, Guangzhong Sun
ICPP3
2025 WinRS: Accelerate Winograd Backward-Filter Convolution with Tiny Workspace
abstract
Winograd algorithm powerfully accelerates Convolutional Neural Networks. However, for backward-filter convolution (BFC), existing implementations often struggle to achieve both high throughput and low memory usage, due to challenges from large filters and small outputs. We propose WinRS, a fast, memory-efficient, and flexible BFC algorithm. WinRS reduces N-D large filters into 1D formats and precisely splits them to match the fastest kernels. These fully-fused kernels execute BFC in on-chip memory with tiny workspace, leveraging the superior acceleration potential of 1D Winograd. WinRS adaptively balances workloads into an optimal number of block groups, maximizing hardware utilization in small-output cases. When ported to FP16 on Tensor Cores, WinRS achieves 3.27 × throughput of its FP32 CUDA-Core version. In experiments, WinRS achieves 1.05 × to 4.7 × speedup over cuDNN GEMM using comparable workspace; WinRS uses less than 4% workspace of cuDNN FFT and Winograd, and exhibits higher throughput with memory- and FLOP-bound workloads.
Junshi Chen 0003, Jingwei Sun 0001, Zhuopin Xu, Jun Shi 0007, Qi Wang 0131
ICPP3
2025 ConCo: Optimizing Compilation of Concurrent Tensor Programs on Shared GPU
abstract
Serving multiple inference tasks of deep neural networks (DNNs) concurrently on a shared GPU is an established method for maximizing hardware resource.Although DNN compilers effectively generate optimal kernel code for individual DNN inferences, they fall short in optimizing for concurrent tasks.This paper presents ConCo, a concurrencyaware compilation scheme designed to optimize the execution of concurrent DNN inference tasks on a shared GPU.ConCo dynamically generates multiple code variants, each tailored to different GPU resource constraints, and efficiently selects optimal variants at runtime according to concurrent workload characteristics.To mitigate the substantial overhead associated with multi-variant compilation, ConCo employs an optimal-code-sharing strategy, significantly accelerating compilation by leveraging commonalities across resource configurations.Evaluations demonstrate that ConCo improves inference throughput by up to 1.2× and reduces job completion time by up to 69.85% compared to existing solutions.
Jiamin Lu, Jingwei Sun 0001, Guangzhong Sun
ICS2
2025 Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their extensive parameter scales pose significant challenges for practical deployment. Unstructured pruning has emerged as an effective model compression strategy with minimal performance loss, which introduces fine-grained sparsity for weight parameters. While existing methods employ a layer-wise pruning strategy to avoid the complexity of global pruning for billion-scale LLMs, they require appropriate sparsity allocation for the layer-wise pruning objectives and often lead to suboptimal solutions for the overall model. In this paper, we propose Lua-LLM ($\textbf{L}$earning $\textbf{u}$nstructured-sparsity $\textbf{a}$llocation in LLMs), a learning-based global pruning framework that explores the optimal unstructured sparsity allocation. Unlike existing pruning methods, which primarily focus on allocating per-layer sparsity, Lua-LLM achieves flexible allocation for both layer-wise and intra-layer sparsity. Furthermore, Lua-LLM leverages a soft Top-K operator to approximate the importance-based mask selection mechanism, enabling efficient binary mask learning. Experimental results on LLaMA and OPT families demonstrate significant performance improvements over existing methods.
Mingge Lu, Jingwei Sun 0001, Junqing Lin, Zechun Zhou, Guangzhong Sun
NeurIPS2
2025 GNNPilot: A Holistic Framework for High-Performance Graph Neural Network Computations on GPUs
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools for graph-based machine learning tasks, but their performance is often constrained by inefficient sparse operators and limited hardware utilization during multi-operator workflows. This article presents GNNPilot, a holistic optimization framework that addresses these challenges through three key innovations. First, we introduce two packing strategies for gather operators, including neighbor packing for load balancing in sparser graphs, and bin packing with a new sparse format for enhanced data locality in denser graphs. Second, we propose dynamic parallelization methods and a novel row panel-based kernel fusion technique to optimize complex multi-operator GNN models. Third, we develop a lightweight sampling-based auto-tuning mechanism that adapts the framework’s optimization strategies to varying input characteristics. Built upon tensor expression-based intermediate representations, GNNPilot maintains the flexibility to optimize both popular and customized GNN models. Extensive experiments across diverse GNN models and graph datasets demonstrate that GNNPilot achieves substantial speedups over state-of-the-art implementations in both the performance of single operators and the efficiency of end-to-end inference. These results establish GNNPilot as an efficient and adaptive solution for accelerating GNN computations on modern GPU architectures.
Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun
ACM Trans. Archit. Code Optim.2
2024 A Learning-path based Supervised Method for Concept Prerequisite Relations Extraction in Educational Data
abstract
In educational data mining, concept prerequisite relations extraction determines which concepts need to be learned before learning another concept. It plays a crucial role in pedagogical practices, such as learning path planning and curriculum design. Deep neural networks, especially graph neural networks, have recently made significant strides in concept prerequisite relations extraction. However, existing methods face two primary limitations. (1) Methods with better performance construct heterogeneous complete graphs, leading to higher model complexity and training cost. Meanwhile, the performance of low-complexity methods is inferior to the former. (2) A disregard for temporal context, essential for learning, limits both the performance and the application of these methods. To address these issues, we propose a novel graph-based approach, called Learning-path based Concept Prerequisite Relations Extraction (LCPRE). LCPRE constructs a lightweight sparse graph in a simple manner, which reduces complexity from quadratic to linear and captures the temporal feature through learning-path, a comprehensible learning approach from one concept to another. Experimental results on three benchmark datasets demonstrate that LCPRE outperforms existing methods, establishing a new state-of-the-art in concept prerequisite relations extraction.
Yiyu Xu, Jingwei Sun 0001, Guangzhong Sun
CIKM4
2024 Siesta: Synthesizing Proxy Applications for MPI Programs
abstract
Proxy applications (proxy-apps) are basic tools for evaluating the performance of specific workloads on high-performance computing (HPC) systems. Since the development of high-fidelity proxy-apps, which exhibit similar performance characteristics as corresponding production applications, is labor-intensive, synthetic proxy-apps are created as a useful supplement to manually developed proxy-apps. To thoroughly resemble performance characteristics of HPC applications represented by Message Passing Interface (MPI) programs, we propose Siesta, a novel framework to automatically synthesize proxy-apps based on communication-computation traces. Given an MPI program, Siesta synthesizes parameterized code snippets to mimic computation behaviors in different execution periods, and combines the code snippets and MPI function records into an event trace. It then extracts program behavior patterns from the trace as grammars and finally transforms the grammars into a synthetic proxy-app. We evaluate the proposed methods on representative MPI programs with various environments. The results show that our synthetic proxy-apps can precisely approximate the performance characteristics of MPI programs,
Jiyu Luo, Qingguo Xu, Jingwei Sun 0001, Guangzhong Sun
CLUSTER4
2024 DProbe: Profiling and Predicting Multi-tenant Deep Learning Workloads for GPU Resource Scaling
Zechun Zhou, Jingwei Sun 0001, Hengquan Mei, Guangzhong Sun
Euro-Par (1)2
2024 PckGNN: Optimizing Aggregation Operators with Packing Strategies in Graph Neural Networks
abstract
Graph Neural Network (GNN) is one of the most prominent machine learning models. It involves a substantial amount of graph-based aggregation operators, which can be abstracted as sparse matrix computation kernels. Due to the irregularity of the graph adjacency matrix, the aggregation has long been the performance bottleneck of GNN. According to our measurements, existing GNN implementations fall short of achieving optimal performance due to their insufficient consideration of load balancing and data locality. To bridge these performance gaps, we propose PckGNN, which aims to accelerate GNN aggregation operators on GPUs with packing strategies. PckGNN categorizes graph matrices into two types based on different sparsity levels and conducts two packing strategies respectively. For sparser matrices, Neighbor Packing enhances load balancing through a moderate-grained non-zero grouping approach. For denser matrices, Bin Packing exposes more potentials of cache data reuse by bin partitioning, non-zero extracting, format converting and two-level scheduling. Experimental results on SpMM and SDDMM show that PckGNN achieves speedups of 1.46x ∼ 6.14x over existing implementations. When applied GNN inference of three typical models, it achieves speedups of more than 1.29x over the state-of-the-art frameworks.
Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun
IPDPS2
2024 AdaSAM: Boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks
Hao Sun 0019, Li Shen 0008, Qihuang Zhong, Liang Ding 0006, Shixiang Chen, Jingwei Sun 0001, Jing Li 0047, Guangzhong Sun, Dacheng Tao
Neural Networks6
2024 AG-SpTRSV: An Automatic Framework to Optimize Sparse Triangular Solve on GPUs
abstract
Sparse Triangular Solve (SpTRSV) has long been an essential kernel in the field of scientific computing. Due to its low computational intensity and internal data dependencies, SpTRSV is hard to implement and optimize on graphics processing units (GPUs). Based on our experimental observations, existing implementations on GPUs fail to achieve the optimal performance due to their suboptimal parallelism setups and code implementations plus lack of consideration of the irregular data distribution. Moreover, their algorithm design lacks the adaptability to different input matrices, which may involve substantial manual efforts of algorithm redesigning and parameter tuning for performance consistency. In this work, we propose AG-SpTRSV, an automatic framework to optimize SpTRSV on GPUs, which provides high performance on various matrices while eliminating the costs of manual design. AG-SpTRSV abstracts the procedures of optimizing an SpTRSV kernel as a scheme and constructs a comprehensive optimization space based on it. By defining a unified code template and preparing code variants, AG-SpTRSV enables fine-grained dynamic parallelism and adaptive code optimizations to handle various tasks. Through computation graph transformation and multi-hierarchy heuristic scheduling, AG-SpTRSV generates schemes for task partitioning and mapping, which effectively address the issues of irregular data distribution and internal data dependencies. AG-SpTRSV searches for the best scheme to optimize the target kernel for the specific matrix. A learned lightweight performance model is also introduced to reduce search costs and provide an efficient end-to-end solution. Experimental results with SuiteSparse Matrix Collection on NVIDIA Tesla A100 and RTX 3080 Ti show that AG-SpTRSV outperforms state-of-the-art implementations with geometric average speedups of 2.12x ∼ 3.99x. With the performance model enabled, AG-SpTRSV can provide an efficient end-to-end solution, with preprocessing times ranging from 3.4 to 245 times of the execution time.
Zhengding Hu, Jingwei Sun 0001, Guangzhong Sun
ACM Trans. Archit. Code Optim.2
2024 LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance implementation generated from advanced tensor compilers, they often take a long time to iteratively search tuning configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this article, we propose LO-SpMM to efficiently generate high-performance SpMM implementations for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of the GPU architecture, and a rank-based cost model, LO-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Besides, rather than generating complete SpMM implementations for evaluation, LO-SpMM constructs simplified proxies to quickly estimate performance, thereby substantially reducing compilation and execution costs. Experimental results show that LO-SpMM can reduce the search time by 281× at most, while the performance of generated SpMM implementations is comparable to or better than the state-of-the-art sparse tensor compiling solutions.
Junqing Lin, Jingwei Sun 0001, Honghe Zhang, Xianzhi Yu, Guangzhong Sun
ACM Trans. Archit. Code Optim.2
2023 GPU Occupancy Prediction of Deep Learning Models Using Graph Neural Network
abstract
Over the past few years, deep learning has been rapidly adopted in many fields. Among the various hardware accelerators specifically for deep learning computation, graphics processing units (GPUs) are mainly used. GPU occupancy—the average ratio of active warps to maximum supported warps on all streaming multiprocessors—is an essential indicator of how well GPUs are utilized. Predicting the GPU occupancy of deep learning models is critical for boosting both job runtime performance and platform resource efficiency. However, GPU occupancy prediction is challenging due to the complex factors hidden in framework runtimes and diverse architectures and hyperparameters of models. In this paper, we propose DNN-occu to predict the GPU occupancy of deep learning models. Our key observation is that models can be represented as directed acyclic computation graphs. DNN-occu extracts a set of occupancy-related features from the computational semantics of the graph nodes and edges. It also employs a novel graph neural network for better feature encoding and prediction generalization. The experiments on various configurations of real-world deep learning models show that DNN-occu achieves high accuracy for occupancy prediction (with an overall error of 9.271%) and has a strong generalization ability for unseen models. In addition, we apply DNN-occu in a trace-driven simulation of deep learning workload scheduling and achieve up to a 31.45% increase in overall GPU utilization and a 19.71% reduction in makespan.
Hengquan Mei, Huaizhi Qu, Jingwei Sun 0001, Yanjie Gao, Haoxiang Lin, Guangzhong Sun
CLUSTER3
2023 EC-SpMM: Efficient Compilation of SpMM Kernel on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance code generated from recent tensor compilers, they often take a long time for iteratively searching candidate configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this paper, we propose EC-SpMM to efficiently generate high-performance SpMM kernels for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of GPU architecture, and a rank-based cost model, EC-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Experimental results show that EC-SpMM can reduce the compilation time by a factor of 35 ×, while the performance of generated SpMM kernels is comparable or even better, compared with the state-of-the-art sparse tensor compiling solution.
Junqing Lin, Honghe Zhang, Jingwei Sun 0001, Xianzhi Yu, Guangzhong Sun
ICPP4
2022 Multi-Net strategy: Accelerating physics-informed neural networks for solving partial differential equations
abstract
Abstract Partial differential equations (PDEs) are the most ubiquitous tools for modeling natural science problems and have long received attention. Physics‐informed neural networks (PINNs) are emerging approaches to approximately solve PDEs. PINNs use automatic differentiation technology to construct the residual of PDEs in the loss function to encode physics conservation laws. We call this process the Single‐Net strategy. Due to the dependency of automatic differentiation among different orders of derivatives, the efficiency of PINNs under the Single‐Net strategy is limited. To address this issue, we propose the Multi‐Net strategy to decouple the dependency. Compared with the Single‐Net strategy, the Multi‐Net strategy reduces the training time of PINNs, and meanwhile, keeps the prediction accuracy. The effectiveness of the proposed strategy is demonstrated through time complexity analysis and a collection of experiments on Burgers equation, advection‐dispersion equation, Kdv equation, and Allen–Cahn equation.
Yunzhuo Wang, Jianfeng Li 0005, Liangying Zhou, Jingwei Sun 0001, Guangzhong Sun
Softw. Pract. Exp.4
2022 Lossy Compression of Communication Traces Using Recurrent Neural Networks
abstract
In high performance computing (HPC) systems, collecting and replaying communication traces are fundamental approaches to analyze performance. With increasingly large-scale HPC systems and applications, tracing tools can produce huge trace data that is costly and challenging to store and analyze. Due to the inherent repetition of behaviors of HPC applications, domain-aware data compression methods can effectively reduce the storage cost of trace data. This study proposes LCR (Lossy Compression and Replay), a framework that aggressively compresses and replays MPI communication traces. Differing from existing trace compression methods, which explicitly identify loop and synchronization structures of communication events, LCR models traces as time series and compactly represents them by lightweight recurrent neural networks. Experimental results demonstrate that LCR can further reduce the size of irregular traces by three orders of magnitude at most, compared with existing structural methods. Meanwhile, LCR accurately reproduces performance and communication patterns of original MPI programs.
Jingwei Sun 0001, Hao Sun 0019, Huancheng Lin, Guangzhong Sun
IEEE Trans. Parallel Distributed Syst.1
2021 An Efficient Channel-level Pruning for CNNs without Fine-tuning
abstract
The success of CNNs in various applications is accompanied by a significant increase in the computation and parameter storage costs. Model compression techniques are able to remove a significant fraction of network parameters to reduce the costs. Channel pruning is among the predominant approaches to compress networks. The typical framework of existing channel pruning methods consists of three steps: training a dense network, pruning redundant parameters, and fine-tuning to increase model accuracy. However, the fine-tuning step usually takes a long time, which is even close to the training costs. Besides, the decoupling of training and pruning leads to that many unimportant parameters that will be removed later still need to be trained. To address these issues, we propose the Dynamic Mask-based Channel Pruning (DMCP) method in this study. The algorithm zeros out unimportant channels with mask vectors and prunes the redundant weights during training. It achieves comparable model accuracy with existing methods, without the fine-tuning step. Moreover, DMCP prunes parameters and reconfigures models during training, so that the number of operations for training useless parameters is reduced. In our evaluations, DMCP removes up to 82% parameters for VGG16, VGG19, and ResNet18 on CIFAR10 dataset, and it reduces up to 58% floating point operation (FLOPs) of training.
Zhongtian Xu, Jingwei Sun 0001, Guangzhong Sun
IJCNN2
2021 Performance Analysis of Graph Neural Network Frameworks
abstract
Graph neural networks (GNNs) are effective models to address learning problems on graphs and have been successfully applied to numerous domains. To improve the productivity of implementing GNNs, various GNN programming frameworks have been developed. Both the effectiveness (accuracy, loss, etc) and the performance (latency, bandwidth, etc) are essential metrics to evaluate the implementation of GNNs. There are many comparative studies related to the effectiveness of different GNN models on domain tasks. However, the performance characteristics of different GNN frameworks are still lacking. In this study, we evaluate the effectiveness and performance of six popular GNN models, GCN, GIN, GAT, GraphSAGE, MoNet, and GatedGCN, across several common benchmarks under two popular GNN frameworks, PyTorch Geometric and Deep Graph Library. We analyze the training time, GPU utilization, and memory usage of different evaluation settings and the performance of models across different hardware configurations under the two frameworks. Our evaluation provides in-depth observations of performance bottlenecks of GNNs and the performance differences between the two popular GNN frameworks. Our work helps GNN researchers understand the performance differences of the popular GNN frameworks, and gives guidelines for developers to find potential performance bugs of frameworks and optimization possibilities of GNNs.
Jingwei Sun 0001, Hao Sun 0019, Guangzhong Sun
ISPASS2
2020 An Active Learning Method for Empirical Modeling in Performance Tuning
abstract
Tuning performance of scientific applications is a challenging problem since performance can be a complicated nonlinear function with respect to application parameters. Empirical performance modeling is a useful approach to approximate the function and enable efficient heuristic methods to find sub-optimal parameter configurations. However, empirical performance modeling requires a large number of samples from the parameter space, which is resource and time-consuming. To address this issue, existing work based on active learning techniques proposed PBU Sampling method considering performance before uncertainty, which iteratively performs performance biased sampling to model the high-performance subspace instead of the entire space before evaluating the most uncertain samples to reduce redundancy. Compared with uniformly random sampling, this approach can reduce the number of samples, but it still involves redundant sampling that potentially can be improved.We propose a novel active learning based method to exploit the information of evaluated samples and explore possible high-performance parameter configurations. Specifically, we adopt a Performance Weighted Uncertainty (PWU) sampling strategy to identify the configurations with either high performance or high uncertainty and determine which ones are selected for evaluation. To evaluate the effectiveness of our proposed method, we construct random forest to predict the execution time of kernels from SPAPT suite and two typical scientific parallel applications kripke, hypre. Experimental results show that compared with existing methods, our proposed method can reduce the cost of modeling by up to 21x and 3x on average meanwhile hold the same prediction accuracy.
Jiepeng Zhang, Jingwei Sun 0001, Wenju Zhou, Guangzhong Sun
IPDPS2
2020 Automated Performance Modeling of HPC Applications Using Machine Learning
abstract
Automated performance modeling and performance prediction of parallel programs are highly valuable in many use cases, such as in guiding task management and job scheduling, offering insights of application behaviors, and assisting resource requirement estimation. The performance of parallel programs is affected by numerous factors, including but not limited to hardware, applications, algorithms, and input parameters, thus an accurate performance prediction is often a challenging and daunting task. In this article, we focus on automatically predicting the execution time of parallel programs (more specifically, MPI programs) with different inputs, at different scales, and without domain knowledge. We model the correlation between the execution time and domain-independent runtime features. These features include values of variables, counters of branches, loops, and MPI communications. Through automatically instrumenting an MPI program, each execution of the program will output a feature vector and its corresponding execution time. After collecting data from executions with different inputs, a random forest machine learning approach is used to build an empirical performance model, which can predict the execution time of the program given a new input. A transfer learning method is used to reuse an existing performance model and improve the prediction accuracy on a new platform that lacks historical execution data. Our experiments and analyses of three parallel applications, Graph500, GalaxSee, and SMG2000, on three different systems confirm that our method performs well, with less than 20 percent prediction error on average.
Jingwei Sun 0001, Guangzhong Sun, Shiyan Zhan, Jiepeng Zhang, Yong Chen 0001
IEEE Trans. Computers1
2016 SPLZ: An efficient algorithm for single source shortest path problem using compression method
Jingwei Sun 0001, Guangzhong Sun
GeoInformatica1