EDBT 2026 Demo / reviewers in the wild / expert
Pedro J. Martínez-Ferrer
dblp:238/4518
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-6097-6676ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A task-based data-flow methodology for programming heterogeneous systems with multiple accelerator APIs
Aleix Boné, Alejandro Aguirre 0005, David Álvarez 0006, Pedro J. Martínez-Ferrer, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 4 |
| 2025 | Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
Pedro J. Martínez-Ferrer, Albert-Jan Nicholas Yzelman, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 1 |
| 2023 | Assessing Saiph, a task-based DSL for high-performance computational fluid dynamics
Sandra Macià, Pedro J. Martínez-Ferrer, Eduard Ayguadé, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 2 |
| 2023 | Improving the performance of classical linear algebra iterative methods via hybrid parallelism
Pedro J. Martínez-Ferrer, Tufan Arslan, Vicenç Beltran 0001 |
J. Parallel Distributed Comput. | 1 |
| 2022 | A Native Tensor-Vector Multiplication Algorithm for High Performance ComputingabstractTensor computations are important mathematical operations for applications that rely on multidimensional data. The tensor–vector multiplication (TVM) is the most memory-bound tensor contraction in this class of operations. This article proposes an open-source TVM algorithm which is much simpler and efficient than previous approaches, making it suitable for integration in the most popular BLAS libraries available today. Our algorithm has been written from scratch and features unit-stride memory accesses, cache awareness, mode obliviousness, full vectorization and multi-threading as well as NUMA awareness for non-hierarchically stored dense tensors. Numerical experiments are carried out on tensors up to order 10 and various compilers and hardware architectures equipped with traditional DDR and high bandwidth memory (HBM). For large tensors the average performance of the TVM ranges between 62% and 76% of the theoretical bandwidth for NUMA systems with DDR memory and remains independent of the contraction mode. On NUMA systems with HBM the TVM exhibits some mode dependency but manages to reach performance figures close to peak values. Finally, the higher-order power method is benchmarked with the proposed TVM kernel and delivers on average between 58% and 69% of the theoretical bandwidth for large tensors. Pedro J. Martínez-Ferrer, Albert-Jan Nicholas Yzelman, Vicenç Beltran 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | HDOT - An approach towards productive programming of hybrid applications
Jan Ciesko, Pedro J. Martínez-Ferrer, Raúl Peñacoba Veigas, Xavier Teruel, Vicenç Beltran 0001 |
J. Parallel Distributed Comput. | 2 |