VLDB 2026 Research / reviewers in the wild / expert
Victor Ferrari
dblp:149/7668
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0002-2051-2877ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 87% Memory systems · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
code generation |
0.7 | 1 | 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization |
0.7 | 1 | 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.7 | 1 | 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023 |
Memory systems › memory access optimization
cache blocking |
0.2 | 1 | 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions · ACM Trans. Archit. Code Optim. 2023 |
Methods — techniques the papers use, named apart from their topics
vector-based packing · 1.3convolution slicing analysis · 1.3MLIR/LLVM · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Advancing Direct Convolution Using Convolution Slicing Optimization and ISA ExtensionsabstractConvolution is one of the most computationally intensive operations that must be performed for machine learning model inference. A traditional approach to computing convolutions is known as the Im2Col + BLAS method. This article proposes SConv: a direct-convolution algorithm based on an MLIR/LLVM code-generation toolchain that can be integrated into machine-learning compilers. This algorithm introduces: (a) Convolution Slicing Analysis (CSA)—a convolution-specific 3D cache-blocking analysis pass that focuses on tile reuse over the cache hierarchy; (b) Convolution Slicing Optimization—a code-generation pass that uses CSA to generate a tiled direct-convolution macro-kernel; and (c) Vector-based Packing—an architecture-specific optimized input-tensor packing solution based on vector-register shift instructions for convolutions with unitary stride. Experiments conducted on 393 convolutions from full ONNX-MLIR machine learning models indicate that the elimination of the Im2Col transformation and the use of fast packing routines result in a total packing time reduction, on full model inference, of 2.3×–4.0× on Intel x86 and 3.3×–5.9× on IBM POWER10. The speed-up over an Im2Col + BLAS method based on current BLAS implementations for end-to-end machine-learning model inference is in the range of 11%–27% for Intel x86 and 11%–34% for IBM POWER10 architectures. The total convolution speedup for model inference is 13%–28% on Intel x86 and 23%–39% on IBM POWER10. SConv also outperforms BLAS GEMM, when computing pointwise convolutions in more than 82% of the 219 tested instances. Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, José E. Moreira, Guido Araujo |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | Improving Convolution via Cache Hierarchy Tiling and Reduced PackingabstractConvolution is one of the most computationally intensive machine learning model operations, usually solved by the known Im2Col + BLAS method. This work proposes a novel convolution-algorithm to improve upon Im2Col + BLAS by introducing (a) CSA: a convolution specific 3D cache-blocking analysis that focuses on tile reuse over the cache hierarchy, (b) CSO: a macro-kernel that follows CSA to compute the convolution by tiling it, (c) a specialized microkernel that seeks to achieve peak hardware performance, and (d) packing routines for the input tensor and filters to bridge the gap between tiling and micro-kernel. Our approach speeds up end-to-end machine learning model inference by up to 26% and 21% for x86 and POWER10 architectures, respectively. Victor Ferrari, Rafael C. F. Sousa, Márcio Machado Pereira, João P. L. de Carvalho, José Nelson Amaral, Guido Araujo |
PACT | 1 |