Lucas Wilkinson

dblp:354/1441 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2023
0009-0000-2052-1290ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 50% High-performance computing · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop optimization
0.712023
Register Tiling for Unstructured Sparsity in Neural Network Inference · Proc. ACM Program. Lang. 2023
Compilers and program optimization › loop transformation
register tiling
0.712023
Register Tiling for Unstructured Sparsity in Neural Network Inference · Proc. ACM Program. Lang. 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator › inference accelerator
neural network inference accelerator
0.712023
Register Tiling for Unstructured Sparsity in Neural Network Inference · Proc. ACM Program. Lang. 2023
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.712023
Register Tiling for Unstructured Sparsity in Neural Network Inference · Proc. ACM Program. Lang. 2023

Methods — techniques the papers use, named apart from their topics

unroll-and-sparse-jam · 1.3data compression · 1.3
YearPublicationVenuePosition
2023 Register Tiling for Unstructured Sparsity in Neural Network Inference
abstract
Unstructured sparse neural networks are an important class of machine learning (ML) models, as they compact model size and reduce floating point operations. The execution time of these models is frequently dominated by the sparse matrix multiplication (SpMM) kernel, C = A × B , where A is a sparse matrix, and B and C are dense matrices. The unstructured sparsity pattern of matrices in pruned machine learning models along with their sparsity ratio has rendered useless the large class of libraries and systems that optimize sparse matrix multiplications. Reusing registers is particularly difficult because accesses to memory locations should be known statically. This paper proposes Sparse Register Tiling, a new technique composed of an unroll-and-sparse-jam transformation followed by data compression that is specifically tailored to sparsity patterns in ML matrices. Unroll-and-sparse-jam uses sparsity information to jam the code while improving register reuse. Sparse register tiling is evaluated across 2396 weight matrices from transformer and convolutional models with a sparsity range of 60-95% and provides an average speedup of 1.72× and 2.65× over MKL SpMM and dense matrix multiplication, respectively, on a multicore CPU processor. It also provides an end-to-end speedup of 2.12× for MobileNetV1 with 70% sparsity on an ARM processor commonly used in edge devices.
Lucas Wilkinson, Kazem Cheshmi, Maryam Mehri Dehnavi
Proc. ACM Program. Lang.1