EDBT 2026 Demo / reviewers in the wild / expert
João V. F. Lima
dblp:98/8779 · also Joao Vicente Ferreira Lima
· DBLP profile ↗
15ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-2670-6963ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance and Cost Evaluation of StarPU on AWS: Case Studies With Dense Linear Algebra Kernels and N-Body SimulationsabstractABSTRACT Task‐based programming interfaces introduce a paradigm in which computations are decomposed into fine‐grained units of work known as “tasks”. StarPU is a runtime system originally developed to support task‐based parallelism on on‐premise heterogeneous architectures by abstracting low‐level hardware details and efficiently managing resource scheduling. It enables developers to express applications as task graphs with explicit data dependencies, which are then dynamically scheduled across available processing units, such as CPUs and GPUs. In recent years, major cloud providers have begun offering virtual machines equipped with both CPUs and GPUs, allowing researchers to deploy and execute parallel workloads in virtual heterogeneous clusters. However, the performance and cost effectiveness of executing StarPU‐based applications in public cloud environments remain unclear, particularly due to variability in hardware configurations, network performance, ever‐changing pricing models, and computing performance due to virtualization and multi‐tenancy. In this paper, we evaluate the performance and cost‐efficiency of StarPU on Amazon Elastic Compute Cloud (EC2) using dense linear algebra kernels and N‐Body simulations as case studies. Our experiments consider different cluster configurations, including powerful and more expensive instances with four NVIDIA GPUs per node (which we refer to as “fat nodes”), and less powerful and lower‐cost instances with a single NVIDIA GPU per node (which we refer to as “thin nodes”). Our results show that arithmetic precision affects the performance–cost trade‐off for dense linear algebra applications, whereas N‐Body simulations consistently achieve better cost‐efficiency on thin‐node clusters. These findings underscore the challenges of optimizing HPC workloads for performance and cost in cloud environments. Vanderlei Munhoz, Vinícius Garcia Pinto, João V. F. Lima, Márcio Castro 0001, Daniel Cordeiro, Emilio Francesquini |
Concurr. Comput. Pract. Exp. | 3 |
| 2025 | Reproducibility Report for SC25 Paper XaaS Containers: Performance-Portable Representation With Source and IR ContainersabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper XaaS Containers: Performance-Portable Representation With Source and IR Containers by Marcin Copik et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee. João V. F. Lima |
SC | 1 |
| 2023 | NAS Parallel Benchmarks with Python: a performance and programming effort analysis focusing on GPUs
Daniel Di Domenico, João V. F. Lima, Gerson G. H. Cavalheiro |
J. Supercomput. | 2 |
| 2023 | An evaluation of relational and NoSQL distributed databases on a low-power cluster
Lucas Ferreira Da Silva, João V. F. Lima |
J. Supercomput. | 2 |
| 2022 | NAS Parallel Benchmark Kernels with Python: A performance and programming effort analysis focusing on GPUsabstractGPU devices are currently seen as one of the trending topics for parallel computing. Commonly, GPU applications are developed with programming tools based on compiled languages, like C/C++ and Fortran. This paper presents a performance and programming effort analysis employing the Python high-level language to implement the NAS Parallel Benchmark kernels targeting GPUs. We used Numba environment to enable CUDA support in Python, a tool that allows us to implement a GPU application with pure Python code. Our experimental results showed that Python applications reached a performance similar to C++ programs employing CUDA and better than C++ using OpenACC for most NPB kernels. Furthermore, Python codes required less operations related to the GPU framework than CUDA, mainly because Python needs a lower number of statements to manage memory allocations and data transfers. However, our Python versions demanded more operations than OpenACC implementations. Daniel Di Domenico, Gerson G. H. Cavalheiro, João V. F. Lima |
PDP | 3 |
| 2021 | Collaborative execution of fluid flow simulation using non-uniform decomposition on heterogeneous architectures
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux |
J. Parallel Distributed Comput. | 3 |
| 2020 | XKBlas: a High Performance Implementation of BLAS-3 Kernels on Multi-GPU ServerabstractIn the last ten years, GPUs have dominated the market considering the computing/power metric and numerous research works have provided Basic Linear Algebra Subprograms implementations accelerated on GPUs. Several software libraries have been developed for exploiting performance of systems with accelerators, but the real performance may be far from the platform peak performance. This paper presents XKBlas that aims to improve performance of BLAS-3 kernels on multi-GPU systems. At low level, we model computation as a set of tasks accessing data on different resources. At high level, the API design favors non-blocking calls as uniform concept to overlap latency, even by fine grain computation. Unit benchmark of BLAS-3 kernels showed that XKBlas outperformed most implementations including the overhead of dynamic task's creation and scheduling. XKBlas outperformed BLAS implementations such as cuBLAS-XT, PaRSEC, BLASX and Chameleon/StarPU. João V. F. Lima |
PDP | 2 |
| 2019 | A Dynamic Task-Based D3Q19 Lattice-Boltzmann Method for Heterogeneous ArchitecturesabstractNowadays computing platforms expose a significant number of heterogeneous processing units such as multicore processors and accelerators. The task-based programming model has been a de facto standard model for such architectures since its model simplifies programming by unfolding parallelism at runtime based on data-flow dependencies between tasks. Many studies have proposed parallel strategies over heterogeneous platforms with accelerators. However, to the best of our knowledge, no dynamic task-based strategy of the Lattice-Boltzmann Method (LBM) has been proposed to exploit CPU+GPU computing nodes. In this paper, we present a dynamic task-based D3Q19 LBM implementation using three runtime systems for heterogeneous architectures: OmpSs, StarPU, and XKaapi. We detail our implementations and compare performance over two heterogeneous platforms. Experimental results demonstrate that our task-based approach attained up to 8.8 of speedup over an OpenMP parallel loop version. João V. F. Lima, Gabriel Freytag, Vinícius Garcia Pinto, Claudio Schepke, Philippe Olivier Alexandre Navaux |
PDP | 1 |
| 2019 | Non-uniform Partitioning for Collaborative Execution on Heterogeneous ArchitecturesabstractSince the demand for computing power increases, new architectures arise to obtain better performance. An important class of integrated devices is heterogeneous architectures, which join different specialized hardware into a single chip, composing a System on Chip - SoC. Within this context, effectively splitting tasks between the different architectures is primal to obtain efficiency and performance. In this work, we evaluate two heterogeneous architectures: one composed of a general-purpose CPU and a graphics processing unit (GPU) integrated into a single chip (AMD Kaveri SoC), and another composed by a general-purpose CPU and a Field Programmable Gate Array (FPGA) integrated into a single chip (Intel Arria 10 SoC). We investigate how data partitioning affects the performance of each device in a collaborative execution through the decomposition of the data domain. As a case study, we apply the technique in the well-known Lattice Boltzmann Method (LBM), analyzing the performance of five kernels in both architectures. Our experimental results show that non-uniform partitioning improves LBM kernels performance by up to 11.40% and 15.15% on AMD Kaveri and Intel Arria 10, respectively. Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 3 |
| 2015 | Design and analysis of scheduling strategies for multi-CPU and multi-GPU architectures
João V. F. Lima, Vincent Danjean, Bruno Raffin, Nicolas Maillard |
Parallel Comput. | 1 |
| 2014 | Scheduling Data Flow Program in XKaapi: A New Affinity Based Algorithm for Heterogeneous Architectures
Raphaël Bleuse, João V. F. Lima, Grégory Mounié, Denis Trystram |
Euro-Par | 3 |
| 2013 | XKaapi: A Runtime System for Data-Flow Task Programming on Heterogeneous ArchitecturesabstractMost recent HPC platforms have heterogeneous nodes composed of multi-core CPUs and accelerators, like GPUs. Programming such nodes is typically based on a combination of OpenMP and CUDA/OpenCL codes; scheduling relies on a static partitioning and cost model. We present the XKaapi runtime system for data-flow task programming on multi-CPU and multi-GPU architectures, which supports a data-flow task model and a locality-aware work stealing scheduler. XKaapi enables task multi-implementation on CPU or GPU and multi-level parallelism with different grain sizes. We show performance results on two dense linear algebra kernels, matrix product (GEMM) and Cholesky factorization (POTRF), to evaluate XKaapi on a heterogeneous architecture composed of two hexa-core CPUs and eight NVIDIA Fermi GPUs. Our conclusion is two-fold. First, fine grained parallelism and online scheduling achieve performance results as good as static strategies, and in most cases outperform them. This is due to an improved work stealing strategy that includes locality information; a very light implementation of the tasks in XKaapi; and an optimized search for ready tasks. Next, the multi-level parallelism on multiple CPUs and GPUs enabled by XKaapi led to a highly efficient Cholesky factorization. Using eight NVIDIA Fermi GPUs and four CPUs, we measure up to 2.43 TFlop/s on double precision matrix product and 1.79 TFlop/s on Cholesky factorization; and respectively 5.09 TFlop/s and 3.92 TFlop/s in single precision. João V. F. Lima, Nicolas Maillard, Bruno Raffin |
IPDPS | 2 |
| 2013 | Preliminary Experiments with XKaapi on Intel Xeon Phi CoprocessorabstractThis paper presents preliminary performance comparisons of parallel applications developed natively for the Intel Xeon Phi accelerator using three different parallel programming environments and their associated runtime systems. We compare Intel OpenMP, Intel CilkPlus and XKaapi together on the same benchmark suite and we provide comparisons between an Intel Xeon Phi coprocessor and a Sandy Bridge Xeon-based machine. Our benchmark suite is composed of three computing kernels: a Fibonacci computation that allows to study the overhead and the scalability of the runtime system, a NQueens application generating irregular and dynamic tasks and a Cholesky factorization algorithm. We also compare the Cholesky factorization with the parallel algorithm provided by the Intel MKL library for Intel Xeon Phi. Performance evaluation shows our XKaapi data-flow parallel programming environment exposes the lowest overhead of all and is highly competitive with native OpenMP and CilkPlus environments on Xeon Phi. Moreover, the efficient handling of data-flow dependencies between tasks makes our XKaapi environment exhibit more parallelism for some applications such as the Cholesky factorization. In that case, we observe substantial gains with up to 180 hardware threads over the state of the art MKL, with a 47% performance increase for 60 hardware threads. João V. F. Lima, François Broquedis, Bruno Raffin |
SBAC-PAD | 1 |
| 2012 | Exploiting Concurrent GPU Operations for Efficient Work Stealing on Multi-GPUsabstractThe race for Exascale computing has naturally led the current technologies to converge to multi-CPU/multi-GPU computers, based on thousands of CPUs and GPUs interconnected by PCI-Express buses or interconnection networks. To exploit this high computing power, programmers have to solve the issue of scheduling parallel programs on hybrid architectures. And, since the performance of a GPU increases at a much faster rate than the throughput of a PCI bus, data transfers must be managed efficiently by the scheduler. This paper targets multi-GPU compute nodes, where several GPUs are connected to the same machine. To overcome the data transfer limitations on such platforms, the available soft wares compute, usually before the execution, a mapping of the tasks that respects their dependencies and minimizes the global data transfers. Such an approach is too rigid and it cannot adapt the execution to possible variations of the system or to the application's load. We propose a solution that is orthogonal to the above mentioned: extensions of the Xkaapi software stack that enable to exploit full performance of a multi-GPUs system through asynchronous GPU tasks. Xkaapi schedules tasks by using a standard Work Stealing algorithm and the runtime efficiently exploits concurrent GPU operations. The runtime extensions make it possible to overlap the data transfers and the task executions on current generation of GPUs. We demonstrate that the overlapping capability is at least as important as computing a scheduling decision to reduce completion time of a parallel program. Our experiments on two dense linear algebra problems (Matrix Product and Cholesky factorization) show that our solution is highly competitive with other soft wares based on static scheduling. Moreover, we are able to sustain the peak performance (approx. 310 GFlop/s) on DGEMM, even for matrices that cannot be stored entirely in one GPU memory. With eight GPUs, we archive a speed-up of 6.74 with respect to single-GPU. The performance of our Cholesky factorization, with more complex dependencies between tasks, outperforms the state of the art single-GPU MAGMA code. João V. F. Lima, Nicolas Maillard, Vincent Danjean |
SBAC-PAD | 1 |
| 2010 | Challenges and Issues of Supporting Task Parallelism in MPI
Márcia C. Cera, João V. F. Lima, Nicolas Maillard, Philippe Olivier Alexandre Navaux |
EuroMPI | 2 |