Leisheng Li

dblp:152/9325 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator
1.012026
AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization
1.012026
AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › convolution optimization
winograd convolution
1.012026
AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.312026
AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026

Methods — techniques the papers use, named apart from their topics

winograd convolution · 2.0micro-kernel optimization · 2.0
YearPublicationVenuePosition
2026 AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors
abstract
As Convolutional Neural Networks (CNNs) continue to gain traction in deep learning, Winograd convolution has emerged as a key algorithm to enhance computational efficiency. Although ARM-based CPUs are increasingly prevalent in mobile devices, embedded systems and HPC servers, existing 2D Winograd convolution implementations for ARM often leave room for improvement in transformation efficiency, computational throughput, and overall versatility. Furthermore, the lack of tailored 3D Winograd convolution implementations for ARM architectures stems from the additional complexity of supporting higher-dimensional kernels. AirWino introduces a set of novel optimizations covering transformations, data layouts, micro-kernel computations, and parallelization strategies for both 2D and 3D Winograd convolution. It supports FP32 and FP16 precisions with filter sizes of 3 and 5, targeting a broad range of applications. Evaluations on four distinct ARM platforms show that AirWino consistently outperforms state-of-the-art libraries across various experimental scenarios and hardware configurations, highlighting its efficiency and portability.
Haoyuan Gui, Ximeng Fu, Leisheng Li
AAAI6
2025 Optimization of Generalized Eigensolver for Dense Symmetric Matrices on AMD GPU
Zitong Su, Wenjing Ma, Leisheng Li
J. Comput. Sci. Technol.6
2025 DH_Aligner: A fast short-read aligner on multicore platforms with AVX vectorization
Qiao Sun 0005, Leisheng Li, Huiyuan Li 0002
J. Parallel Distributed Comput.3
2018 Bandwidth Reduced Parallel SpMV on the SW26010 Many-Core Platform
abstract
SpMV (Sparse Matrix-Vector multiplication), in its simplest form y = Ax, multiplies a sparse matrix with a dense vector and is a widely used computing primitive in the domain of HPC. On the newly SW26010 many-core platform, we propose a highly efficient CSR (Compressed Storage Row) based implementation of parallel SpMV, referred to as SWCSR-SpMV in the sequel. SpMV in the CSR format can be trivially parallelized but its performance is majorly impeded by memory access efficiency, and therefore to leverage high-throughput memory access mechanism while avoiding redundant bandwidth usage becomes the major goal of designing high performance SpMV on the target platform. The original problem is sequentially partitioned into row-slices, each of which can reside in the fast scratchpad memory, so that the loaded x'es can be reused; meanwhile, a dynamic look-ahead scheme is applied to avoid redundant memory access; we split the many-core mesh into smaller communication scope to facilitate the sharing of the common data across the working threads via the high speed on-mesh data bus. Beyond the above, to leverage massive parallelism balanced workload is ensured by both static and dynamic means. Performance evaluation is done on a benchmark of 36 frequently used sparse matrices in the fields of graph computing, data mining, computational fluid dynamics, etc.. While the performance upper-bound is defined by the ratio between the minimal data access volume required against the practically optimal bandwidth, ignoring the computing overhead, SWCSR-SpMV can achieve an efficiency of nearly 87%, maintaining over 75% for 1/3 of the testing matrices. SWCSR-SpMV is further applied in a PETSc based application, a 1.75x-2.6x speedup is sustained in a multi-process environment on the Sunway TaiHuLight supercomputer.
Qiao Sun 0005, Changyou Zhang, Changmao Wu, Leisheng Li
ICPP5
2017 Large-Scale Parallelization of Smoothed Particle Hydrodynamics Method on Heterogeneous Cluster
abstract
This paper implements a Smoothed Particle Hydrodynamics simulation code and distributes it on a heterogeneous cluster. The theoretical analysis results show that treating GPU as equivalent peer of CPU rather than an assistant or a substitute is the most efficient way of using a CPU+GPU compute node. However, it raises complex challenges of heterogeneous cooperation. Our strategies of hybrid-level domain decomposition, multi-stage thread pool and data transfer optimization are addressed and the respective effects are showed. Steady balance ratio of 1.1 between 6 nodes and GPU workload ratio of 0.43 between 2CPU+1GPU cards are observed. Average speedups of 2.9x, 1.81x, 1.1x over the single GPU version are obtained on three heterogeneous systems. The weak scaling tests are carried out on 1 to 16384 CPU-GPU nodes and the overall parallel efficiency of 70.3% is achieved.
Yingrui Wang, Leisheng Li, Rong Tian
ICPP2
2016 Fast Parallel Stream Compaction for IA-Based Multi/many-core Processors
abstract
Stream compaction, frequently found in a large variety of applications, serves as a general primitive that reduces an input stream to a subset containing only the wanted elements so that the follow-on computation can be done efficiently. In this paper, we propose a fast parallel stream compaction for IA-based multi-/many-core processors. Unlike the previously studied algorithms that depend heavily on a black-box parallel scan, we open the black-box in the proposed algorithm and manually tailor it so that both the workload and the memory footprint is significantly reduced. By further eliminating the conditional statements and applying automatic code generation/optimization for performance-critical kernels, the proposed parallel stream compaction achieves high performance in different cases and for various data types across different IA-based multi/manycore platforms. Experimental results on three typical IA-based processors, including a quad-core Core-i7 CPU, a dual-socket 8-core Xeon CPU, and a 61-core Xeon Phi accelerator show that the proposed implementation outperforms the referenced parallel counterpart in the state-of-art library Thrust. On top of the above, we apply it in the random forest based data classifier to show its potential to boost the performance of real-world applications.
Qiao Sun 0005, Chao Yang 0002, Changmao Wu, Leisheng Li, Fangfang Liu 0004
CCGrid4
2016 Accelerating the Simulation of Thermal Convection in the Earth's Outer Core on Tianhe-2
abstract
Numerical simulation of thermal convection in the Earth's outer core requires extreme-scale computing due to the large temporal and spatial disparity, extreme physical parameters, rapid rotation and spherical geometry. In this work, the numerical simulation of the thermal convection in the Earth's outer core for CPU-MIC heterogeneous many-core systems is studied. Firstly, starting from a legacy parallel code based on the PETSc software package, a framework of the numerical simulation built on CPU-MIC heterogeneous many-core systems has been developed. Secondly, a sparse linear solver for CPUMIC heterogeneous many-core systems, which focuses on solving the two linear systems of the simulation, is presented and optimized. Thirdly, some computational kernels of the simulation, including sparse matrix-vector multiplication (SpMV) and polynomial preconditioner on distributed memory Xeon Phiaccelerated systems are implemented and optimized. In addition, in order to reduce the cost of data movement, we use methods to minimize the memory access, the PCI-E data transfer, and the MPI communication. Finally, some optimized measures are taken to the extended code. Experiments on Tianhe-2 Supercomputer show that as compared to the original code, our Xeon Phiaccelerated design is able to deliver 6.93x and 6.00x speedups for single MIC device and 64 MIC devices, respectively.
Changmao Wu, Fangfang Liu 0004, Chao Yang 0002, Ligang Li, Yutong Lu, Leisheng Li, Yunfei Du 0001
ICPADS7
2016 GPU Acceleration of Smoothed Particle Hydrodynamics for the Navier-Stokes Equations
abstract
Although there exist much work on GPU acceleration on the SPH method, the focus so far has been on the Euler equations in fluid mechanics. This paper presents GPU acceleration on the SPH method for the Navier-Stokes equations for both solid and fluid mechanics. We investigate and compare three CPU-GPU coupling models in terms of one large-scale parallel application code: (1) CPU?GPU (to only run hotspots on GPU), (2) GPU-alone (to run the whole of simulation on GPU), and (3) CPU||GPU (to treat CPU and GPU as equivalent processors). A common issue to the three models, "easy code transplant onto GPU", is emphasized. Optimizations on particle indexing and particle interaction on GPU, which are of unique importance to a SPH code, are addressed. Numerical experiments are finally performed and 4x, 10x, 16x speedups are observed for the three coupling models, respectively, with reference to single CPU core. Among the three, the fastest model -- Xthe "CPU||GPU" model -- Xfurther undergoes scalability tests on a cluster of 6 heterogeneous nodes and shows 90+% parallel efficiency.
Yingrui Wang, Leisheng Li, Rong Tian
PDP2
2014 petaPar: A Scalable Petascale Framework for Meshfree/Particle Simulation
abstract
Since high performance computing sustained petaflops in 2008, numerical simulation entered a new era to use 10K to 100K processor cores in one single run of parallel computing. In pursuit of petascale computing, the challenges of scalability must be addressed. Petapar is a highly scalable simulation framework which implements two popular meshfree/particle methods, the smoothed particle hydrodynamics (SPH) and the material point method (MPM). The parallelization starts from the regular-grid-based domain decomposition. The scalability of the code is assured by fully overlapping of communication and computation, and a dynamic load balancing strategy. Petapar supports both flat MPI and MPI+Pthreads hybrid parallelization. The code is tested on Titan, which ranked first on the Top500 supercomputer list when our research work has been done in November 2012. Experiment results show that petaPar linearly scales up to 260K CPU cores with an excellent parallel efficiency of 100% and 96% for the SPH and MPM, respectively.
Leisheng Li, Yingrui Wang, Zhitao Ma, Rong Tian
ISPA1