Victor Ferrari

dblp:149/7668 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0002-2051-2877ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 87% Memory systems · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
code generation
0.712023
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization
0.712023
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023
Memory systems › memory access optimization
cache blocking
0.212023
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023

Methods — techniques the papers use, named apart from their topics

vector-based packing · 1.3convolution slicing analysis · 1.3MLIR/LLVM · 1.3
YearPublicationVenuePosition
2023 Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions
abstract
Convolution is one of the most computationally intensive operations that must be performed for machine learning model inference. A traditional approach to computing convolutions is known as the Im2Col + BLAS method. This article proposes SConv: a direct-convolution algorithm based on an MLIR/LLVM code-generation toolchain that can be integrated into machine-learning compilers. This algorithm introduces: (a) Convolution Slicing Analysis (CSA)—a convolution-specific 3D cache-blocking analysis pass that focuses on tile reuse over the cache hierarchy; (b) Convolution Slicing Optimization—a code-generation pass that uses CSA to generate a tiled direct-convolution macro-kernel; and (c) Vector-based Packing—an architecture-specific optimized input-tensor packing solution based on vector-register shift instructions for convolutions with unitary stride. Experiments conducted on 393 convolutions from full ONNX-MLIR machine learning models indicate that the elimination of the Im2Col transformation and the use of fast packing routines result in a total packing time reduction, on full model inference, of 2.3×–4.0× on Intel x86 and 3.3×–5.9× on IBM POWER10. The speed-up over an Im2Col + BLAS method based on current BLAS implementations for end-to-end machine-learning model inference is in the range of 11%–27% for Intel x86 and 11%–34% for IBM POWER10 architectures. The total convolution speedup for model inference is 13%–28% on Intel x86 and 23%–39% on IBM POWER10. SConv also outperforms BLAS GEMM, when computing pointwise convolutions in more than 82% of the 219 tested instances.
Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, José E. Moreira, Guido Araujo
ACM Trans. Archit. Code Optim.1
2022 Improving Convolution via Cache Hierarchy Tiling and Reduced Packing
abstract
Convolution is one of the most computationally intensive machine learning model operations, usually solved by the known Im2Col + BLAS method. This work proposes a novel convolution-algorithm to improve upon Im2Col + BLAS by introducing (a) CSA: a convolution specific 3D cache-blocking analysis that focuses on tile reuse over the cache hierarchy, (b) CSO: a macro-kernel that follows CSA to compute the convolution by tiling it, (c) a specialized microkernel that seeks to achieve peak hardware performance, and (d) packing routines for the input tensor and filters to bridge the gap between tiling and micro-kernel. Our approach speeds up end-to-end machine learning model inference by up to 26% and 21% for x86 and POWER10 architectures, respectively.
Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, Guido Araujo
PACT1