Alexandre de Limas Santana

dblp:340/6516 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-3203-3662ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 38% Processor architecture and microarchitecture · 38% Memory systems · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization
0.712023
Efficient Direct Convolution Using Long SIMD Instructions · PPoPP 2023
Processor architecture and microarchitecture › SIMD
SIMD instructions
0.712023
Efficient Direct Convolution Using Long SIMD Instructions · PPoPP 2023
Memory systems
cache
0.212023
Efficient Direct Convolution Using Long SIMD Instructions · PPoPP 2023
Memory systems › cache › cache miss
cache conflict misses
0.212023
Efficient Direct Convolution Using Long SIMD Instructions · PPoPP 2023

Methods — techniques the papers use, named apart from their topics

multi-block direct convolution · 0.7memory layout transformation · 0.7bounded direct convolution · 0.7
YearPublicationVenuePosition
2026 Just-in-Time Convolution and GEMM Code Generation for SIMD Architectures
Alexandre de Limas Santana, Adrià Armejach, Marc Casas
IPDPS1
2023 Efficient Direct Convolution Using Long SIMD Instructions
abstract
This paper demonstrates that state-of-the-art proposals to compute convolutions on architectures with CPUs supporting SIMD instructions deliver poor performance for long SIMD lengths due to frequent cache conflict misses. We first discuss how to adapt the state-of-the-art SIMD direct convolution to architectures using long SIMD instructions and analyze the implications of increasing the SIMD length on the algorithm formulation. Next, we propose two new algorithmic approaches: the Bounded Direct Convolution (BDC), which adapts the amount of computation exposed to mitigate cache misses, and the Multi-Block Direct Convolution (MBDC), which redefines the activation memory layout to improve the memory access pattern. We evaluate BDC, MBDC, the state-of-the-art technique, and a proprietary library on an architecture featuring CPUs with 16,384-bit SIMD registers using ResNet convolutions. Our results show that BDC and MBDC achieve respective speed-ups of 1.44× and 1.28× compared to the state-of-the-art technique for ResNet-101, and 1.83× and 1.63× compared to the proprietary library.
Alexandre de Limas Santana, Adrià Armejach, Marc Casas
PPoPP1
2021 PackStealLB: A scalable distributed load balancer based on work stealing and workload discretization
Vinicius Freitas, Laércio Lima Pilla, Alexandre de Limas Santana, Márcio Castro 0001, Johanne Cohen
J. Parallel Distributed Comput.3
2021 ARTful: A model for user-defined schedulers targeting multiple high-performance computing runtime systems
abstract
Abstract Global schedulers are components in parallel runtime libraries that distribute the application's workload across physical resources. More often than not, applications showcase dynamic load imbalance and require customized scheduling solutions to avoid wasting resources. Some libraries lack support for user‐defined schedulers and developers resort to unofficial extensions that are harder to reuse and maintain. We propose a global scheduler software design, entitled ARTful model, to create user‐defined solutions with minimal alterations in the runtime library. Our model uses a component‐based design to separate components from the runtime library and the scheduling policy implementation. The ARTful modeldescribes the interface of a portable scheduler library, allowing policies to operate on different runtime libraries. We study the overhead induced by our design through our ARTful library implementation metaprogramming‐oriented global scheduling library using workload‐aware scheduling policies. We experiment with two different policies from OpenMP and Charm++ runtime systems, also presenting evaluations of the policies outside of their original library context. We observe that our portable schedulers can sometimes perform decisions faster than their native counterparts with negligible overhead in the execution times of synthetic applications and molecular dynamics kernels.
Alexandre de Limas Santana, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla, Jean-François Méhaut
Softw. Pract. Exp.1
2018 A Batch Task Migration Approach for Decentralized Global Rescheduling
abstract
Effectively mapping tasks of High Performance Computing (HPC) applications on parallel systems is crucial to assure substantial performance gains. As platforms and applications grow, load imbalance becomes a priority issue. Even though centralized rescheduling has been a viable solution to mitigate this problem, its efficiency is not able to keep up with the increasing size of shared memory platforms. To efficiently solve load imbalance today, and in the years to come, we should prioritize decentralized strategies developed for large scale platforms. In this paper, we propose our Batch Task Migration approach to improve decentralized global rescheduling, ultimately reducing communication costs and preserving task locality. We implemented and evaluated our approach in two different parallel platforms, using both synthetic workloads and a molecular dynamics (MD) benchmark. Our solution was able to achieve speedups of up to 3.75 and 1.15 on rescheduling time, when compared to other centralized and distributed approaches, respectively. Moreover, it improved the execution time of MD by factors up to 1.34 and 1.22 when compared to a scenario without load balancing on two different platforms.
Vinicius Freitas, Alexandre de Limas Santana, Márcio Castro 0001, Laércio Lima Pilla
SBAC-PAD2