Weiling Yang

dblp:64/5493 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-7167-4086ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers
abstract
Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer latency that exceeds GPU computation by several folds. We present PreScope, a prediction-driven expert scheduling system that addresses three key challenges: inaccurate activation prediction, PCIe bandwidth competition, and cross-device scheduling complexity. Our solution includes: 1) Learnable Layer-Aware Predictor (LLaPor) that captures layer-specific expert activation patterns; 2) Prefetch-Aware Cross-Layer Scheduling (PreSched) that generates globally optimal plans balancing prefetching costs and loading overhead; 3) Asynchronous I/O Optimizer (AsyncIO) that decouples I/O from computation, eliminating waiting bubbles. PreScope achieves 141% higher throughput and 74.6% lower latency than state-of-the-art solutions.
Enda Yu, Dezun Dong, Zhaoning Zhang 0001, Zhe Bai, Weiling Yang, Haojie Wang 0004, Dongsheng Li 0001, Yongwei Wu 0001, Xiangke Liao
ICS5
2026 Demystifying ARM SME to Optimize General Matrix Multiplications
Chencheng Deng, Weiling Yang, Jianbin Fang, Dezun Dong
IPDPS2
2026 (F1,F)-partition of plane graphs without 4- and 5-cycles and without ext-triangular 7-cycles
Xian'an Jin, Tianlong Ma, Weiling Yang
Discret. Appl. Math.3
2025 nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs
abstract
Convolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures.
Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007
IEEE Trans. Computers1
2024 Optimizing Attention by Exploiting Data Reuse on ARM Multi-core CPUs
abstract
Transformers reign supreme in natural language processing, representing a milestone innovation in deep learning. For high-performance model inference, optimizing the time-consuming attention module is crucial. Owing to the irregular-shaped matrix workloads and intricate data access patterns, the attention operator is bounded by memory bandwidth. Existing works utilize kernel fusion to reduce memory access overhead, resulting in promising performance enhancements. However, these efforts primarily focus on GPU or X86 architectures, leaving ARM multi-cores, commonly encountered in emerging HPC systems, insufficiently explored. We present MEATTEN, a memory-efficient attention fusion scheme and batched approach to exploit ARM multi-core CPUs effectively. It builds on fused micro-kernels and a new data layout suitable for SIMD vectorization. An analytic model is used to guide loop permutation, tiling, and batched parallelization according to the on-chip hierarchical memory architecture and workload characterization. We apply MEATTEN to three representative ARM multi-cores against state-of-the-art libraries and compilers. Experimental results demonstrate that our approach consistently outperforms prior approaches across various evaluation scenarios and platforms.
Weiling Yang, Dezun Dong, Xing Su 0004
ICS2
2024 Optimizing Full-Spectrum Matrix Multiplications on ARMv8 Multi-Core CPUs
abstract
General Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. While the mainstream Basic Linear Algebra Subprograms (BLAS) libraries can deliver good performance on large and regular-shaped GEMMs, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in emerging HPC applications. Recent research has focused on improving GEMM performance on GPUs, but there is still significant room for improvement on emerging HPC hardware based on multi-core CPUs. We presentLibShalom2, an open-source library to optimize full-spectrum GEMMs, taking small, irregular-shaped, and large-scale regular-shaped matrices.LibShalom2explicitly targets the ARMv8 architecture, which is becoming common in HPC systems.LibShalom2is designed to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluateLibShalom2by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show thatLibShalom2consistently outperforms existing solutions across full-spectrum GEMM workloads and hardware architectures. We also show thatLibShalom2delivers an average speedup of 2.2x for real-life neural network workloads.
Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0079
IEEE Trans. Parallel Distributed Syst.1
2023 Characterize and Optimize Dense Linear Solver on Multi-core CPUs
abstract
The dense linear solver is an essential subroutine in high-performance computing. Typical parallel implementations either adopt the fork-join or task parallel programming models. Blocked algorithms built upon the fork-join paradigm focus on optimizing cache locality, leaving significant synchronization overhead. Following the data-driven execution model, tile-based algorithms formed on the task parallel paradigm effectively relieve the pain and exhibit superior load balancing. Nevertheless, they introduce redundant memory access expenses, plaguing the CPU execution. In this paper, we first characterize and quantify the impact of the performance bottlenecks in-depth and then propose a series of optimizations. Specifically, we reduce the idle time of threads by merging LU factorization with the subsequent lower triangular solver to improve parallelism. Moreover, we eliminate tile-based matrix format transformation and diminish duplicated data packing operations to lower memory access overhead. Performance evaluation is conducted on two modern multi-core systems, Intel Xeon Gold(R) 6252N and HiSilicon Kunpeng 920. The evaluation results demonstrate the superiority of our proposed solver over state-of-the-art open-source implementations, achieving performance gains of up to 11.5% and 12.2% on the respective platforms.
Xing Su 0004, Dezun Dong, Weiling Yang
ICPADS4
2023 Optimizing Direct Convolutions on ARM Multi-Cores
abstract
Convolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms.
Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001
SC2
2021 Characterizing Small-Scale Matrix Multiplications on ARMv8-based Many-Core Architectures
abstract
General Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. There is a large body of work on evaluating and optimizing large-scale matrix multiplication, but how well the small-scale matrix multiplication (SMM) performs is largely unknown, especially for the ARMv8-based many-core architectures. In this work, we evaluate and characterize the performance of SMM subroutines on Phytium 2000 +, an ARMv8-based 64-core architecture. The evaluation work is extensively performed with the mainstream open-source libraries including OpenBLAS, BLIS, BALSFEO, and Eigen. Given various experimental settings, we observe how well the small-scale GEMM routines perform on Phytium 2000 +, and then discuss the impacting factors behind the performance behaviours of SMM. Built on such a basis, we shed light on the performance bottlenecks and practical optimizations on SMM from various angles: (1) mitigating the data packing overhead, (2) processing the edge cases properly, (3) selecting a suitable micro-kernel, and (4) adopting a right parallelization method. The result of our work facilitates users to develop efficient SMM optimizations on ARMv8-based many-core architectures, and embed them into real-world applications.
Weiling Yang, Jianbin Fang, Dezun Dong
IPDPS1
2021 LIBSHALOM: optimizing small and irregular-shaped matrix multiplications on ARMv8 multi-cores
abstract
General Matrix Multiplication (GEMM) is a key subroutine in highperformance computing. While the mainstream linear algebra libraries can deliver high performance on large and regular-shaped GEMM, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in new HPC applications. Some of the recent works in this direction have made promising progress on x86 architectures and GPUs but still leave much room for improvement on emerging HPC hardware built upon the ARMv8 architecture. We present LibShalom, an open-source library for optimizing small and irregular-shaped GEMMs, explicitly targeting the ARMv8 architecture. LibShalom builds upon the classical Goto algorithm but tailors it to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluate LibShalom by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show that LibShalom can consistently outperform existing solutions across GEMM workloads and hardware architectures.
Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0001
SC1