EDBT 2026 Demo / reviewers in the wild / expert
Xing Su 0004
dblp:76/8056-4
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-7514-1495ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Attention by Exploiting Data Reuse on ARM Multi-core CPUsabstractTransformers reign supreme in natural language processing, representing a milestone innovation in deep learning. For high-performance model inference, optimizing the time-consuming attention module is crucial. Owing to the irregular-shaped matrix workloads and intricate data access patterns, the attention operator is bounded by memory bandwidth. Existing works utilize kernel fusion to reduce memory access overhead, resulting in promising performance enhancements. However, these efforts primarily focus on GPU or X86 architectures, leaving ARM multi-cores, commonly encountered in emerging HPC systems, insufficiently explored. We present MEATTEN, a memory-efficient attention fusion scheme and batched approach to exploit ARM multi-core CPUs effectively. It builds on fused micro-kernels and a new data layout suitable for SIMD vectorization. An analytic model is used to guide loop permutation, tiling, and batched parallelization according to the on-chip hierarchical memory architecture and workload characterization. We apply MEATTEN to three representative ARM multi-cores against state-of-the-art libraries and compilers. Experimental results demonstrate that our approach consistently outperforms prior approaches across various evaluation scenarios and platforms. Weiling Yang, Dezun Dong, Xing Su 0004 |
ICS | 4 |
| 2024 | Optimizing Full-Spectrum Matrix Multiplications on ARMv8 Multi-Core CPUsabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. While the mainstream Basic Linear Algebra Subprograms (BLAS) libraries can deliver good performance on large and regular-shaped GEMMs, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in emerging HPC applications. Recent research has focused on improving GEMM performance on GPUs, but there is still significant room for improvement on emerging HPC hardware based on multi-core CPUs. We presentLibShalom2, an open-source library to optimize full-spectrum GEMMs, taking small, irregular-shaped, and large-scale regular-shaped matrices.LibShalom2explicitly targets the ARMv8 architecture, which is becoming common in HPC systems.LibShalom2is designed to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluateLibShalom2by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show thatLibShalom2consistently outperforms existing solutions across full-spectrum GEMM workloads and hardware architectures. We also show thatLibShalom2delivers an average speedup of 2.2x for real-life neural network workloads. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0079 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Characterize and Optimize Dense Linear Solver on Multi-core CPUsabstractThe dense linear solver is an essential subroutine in high-performance computing. Typical parallel implementations either adopt the fork-join or task parallel programming models. Blocked algorithms built upon the fork-join paradigm focus on optimizing cache locality, leaving significant synchronization overhead. Following the data-driven execution model, tile-based algorithms formed on the task parallel paradigm effectively relieve the pain and exhibit superior load balancing. Nevertheless, they introduce redundant memory access expenses, plaguing the CPU execution. In this paper, we first characterize and quantify the impact of the performance bottlenecks in-depth and then propose a series of optimizations. Specifically, we reduce the idle time of threads by merging LU factorization with the subsequent lower triangular solver to improve parallelism. Moreover, we eliminate tile-based matrix format transformation and diminish duplicated data packing operations to lower memory access overhead. Performance evaluation is conducted on two modern multi-core systems, Intel Xeon Gold(R) 6252N and HiSilicon Kunpeng 920. The evaluation results demonstrate the superiority of our proposed solver over state-of-the-art open-source implementations, achieving performance gains of up to 11.5% and 12.2% on the respective platforms. Xing Su 0004, Dezun Dong, Weiling Yang |
ICPADS | 2 |
| 2021 | LIBSHALOM: optimizing small and irregular-shaped matrix multiplications on ARMv8 multi-coresabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in highperformance computing. While the mainstream linear algebra libraries can deliver high performance on large and regular-shaped GEMM, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in new HPC applications. Some of the recent works in this direction have made promising progress on x86 architectures and GPUs but still leave much room for improvement on emerging HPC hardware built upon the ARMv8 architecture. We present LibShalom, an open-source library for optimizing small and irregular-shaped GEMMs, explicitly targeting the ARMv8 architecture. LibShalom builds upon the classical Goto algorithm but tailors it to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluate LibShalom by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show that LibShalom can consistently outperform existing solutions across GEMM workloads and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0001 |
SC | 4 |
| 2019 | SCP: Shared Cache Partitioning for High-Performance GEMMabstractGEneral Matrix Multiply (GEMM) is the most fundamental computational kernel routine in the BLAS library. To achieve high performance, in-memory data must be prefetched into fast on-chip caches before they are used. Two techniques, software prefetching and data packing, have been used to effectively exploit the capability of on-chip least recent used (LRU) caches, which are popular in traditional high-performance processors used in high-end servers and supercomputers. However, the market has recently witnessed a new diversity in processor design, resulting in high-performance processors equipped with shared caches with non-LRU replacement policies. This poses a challenge to the development of high-performance GEMM in a multithreaded context. As several threads try to load data into a shared cache simultaneously, interthread cache conflicts will increase significantly. We present a Shared Cache Partitioning (SCP) method to eliminate interthread cache conflicts in the GEMM routines, by partitioning a shared cache into physically disjoint sets and assigning different sets to different threads. We have implemented SCP in the OpenBLAS library and evaluated it on Phytium 2000+, a 64-core AArch64 processor with private LRU L1 caches and shared pseudo-random L2 caches (per four-core cluster). Our evaluation shows that SCP has effectively reduced the conflict misses in both L1 and L2 caches in a highly optimized GEMM implementation, resulting in an improvement of its performance by 2.75% to 6.91%. Xing Su 0004, Xiangke Liao, Hao Jiang 0001, Canqun Yang, Jingling Xue |
ACM Trans. Archit. Code Optim. | 1 |
| 2017 | Automatic generation of fast BLAS3-GEMM: a portable compiler approach
Xing Su 0004, Xiangke Liao, Jingling Xue |
CGO | 1 |
| 2016 | Galaxyfly: A Novel Family of Flexible-Radix Low-Diameter Topologies for Large-Scales Interconnection NetworksabstractInterconnection network plays an essential role in the architecture of large-scale high performance computing (HPC) systems. In the paper, we construct a novel family of low-diameter topologies, Galaxyfly, using techniques of algebraic graphs over finite fields. Galaxyfly is guaranteed to retain a small constant diameter while achieving a flexible tradeoff between network scale and bisection bandwidth. Galaxyfly lowers the demands for high radix of network routers and is able to utilize routers with merely moderate radix to build exascale interconnection networks. We present effective congestion-aware routing algorithms for Galaxyfly by exploring its algebraic property. We conduct extensive simulations and analysis to evaluate the performance, cost and power consumption of Galaxyfly against state-of-the-art topologies. The results show that our design achieves better performance than most existing topologies under various routing algorithms and traffic patterns, and is cost-effective to deploy for exascale HPC systems. Dezun Dong, Xiangke Liao, Xing Su 0004, Cunlu Li |
ICS | 4 |
| 2015 | Design and Implementation of a Highly Efficient DGEMM for 64-Bit ARMv8 Multi-core ProcessorsabstractThis paper presents the design and implementation of a highly efficient Double-precision General Matrix Multiplication (DGEMM) based on Open BLAS for 64-bit ARMv8 eight-core processors. We adopt a theory-guided approach by first developing a performance model for this architecture and then using it to guide our exploration. The key enabler for a highly efficient DGEMM is a highly-optimized inner kernel GEBP developed in assembly language. We have obtained GEBP by (1) maximizing its compute-to-memory access ratios across all levels of the memory hierarchy in the ARMv8 architecture with its performance-critical block sizes being determined analytically, and (2) optimizing its computations through exploiting loop unrolling, instruction scheduling and software-implemented register rotation and taking advantage of A64 instructions to support efficient FMA operations, data transfers and prefetching. We have compared our DGEMM implemented in Open BLAS with another implemented in ATLAS (also in terms of a highly-optimized GEBP in assembly). Our implementation outperforms the one in ALTAS by improving the peak performance (efficiency) of DGEMM from 3.88 Gflops (80.9%) to 4.19 Gflops (87.2%) on one core and from 30.4 Gflops (79.2%) to 32.7 Gflops (85.3%) on eight cores. These results translate into substantial performance (efficiency) improvements by 7.79% on one core and 7.70% on eight cores. In addition, the efficiency of our implementation on one core is very close to the theoretical upper bound 91.5% obtained from micro-benchmarking. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Feng Wang 0050, Hao Jiang 0001, Ke Zuo, Xing Su 0004, Jingling Xue, Canqun Yang |
ICPP | 4 |