EDBT 2026 Demo / reviewers in the wild / expert
Tiago Carneiro 0001
dblp:07/9766 · also Tiago Carneiro Pessoa
· DBLP profile ↗
9ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-6145-8352ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Portable PGAS-Based GPU-Accelerated Branch-And-Bound Algorithms at ScaleabstractABSTRACT The Branch‐and‐Bound (B&B) technique plays a key role in solving many combinatorial optimization problems, enabling efficient problem‐solving and decision‐making in a wide range of applications. It incrementally constructs a tree by building candidates to the solutions and abandoning a candidate as soon as it determines that it cannot lead to an optimal solution. With modern problems growing increasingly large, accelerating B&B algorithms through parallelization has become a critical challenge for handling large solution spaces. At the same time, modern parallel computing systems themselves are becoming larger, more heterogeneous, and more diverse, requiring programming approaches capable of effectively exploiting such complexity. To address these challenges, this work presents a GPU‐accelerated B&B algorithm based on the Partitioned Global Address Space (PGAS) programming model, implemented using the Chapel language. The PGAS‐based design is motivated by the high‐level abstraction provided by this programming model, which favors programmability, whereas vendor‐neutral GPU features of the Chapel language favor GPU portability. The algorithm uses a pool‐based approach for generality and exploits a dynamic load balancing mechanism for performance scalability. Extensive experimentation on the N‐Queens and permutation flowshop scheduling problems demonstrated both code performance and code portability of the proposed algorithm on several GPU architectures compared to optimized CUDA‐based implementations. Moreover, the strong scaling efficiency of the proposed algorithm is investigated on a TOP500 pre‐exascale supercomputer up to 1024 GPUs. Guillaume Helbecque, Ezhilmathi Krishnasamy, Tiago Carneiro 0001, Nouredine Melab, Pascal Bouvry |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Investigating Portability in Chapel for Tree-Based Optimization on GPU-Powered Clusters
Tiago Carneiro 0001, Engin Kayraklioglu, Guillaume Helbecque, Nouredine Melab |
Euro-Par (3) | 1 |
| 2023 | Parallel distributed productivity-aware tree-search using ChapelabstractAbstract With the recent arrival of the exascale era, modern supercomputers are increasingly big making their programming much more complex. In addition to performance, software productivity is a major concern to choose a programming language, such as Chapel, designed for exascale computing. In this paper, we investigate the design of a parallel distributed tree‐search algorithm, namely P3D‐DFS, and its implementation using Chapel. The design is based on the Chapel's DistBag data structure, revisited by: (1) redefining the data structure for Depth‐First tree‐Search (DFS), henceforth renamed DistBag‐DFS; (2) redesigning the underlying load balancing mechanism. In addition, we propose two instantiations of P3D‐DFS considering the Branch‐and‐Bound (B&B) and Unbalanced Tree Search (UTS) algorithms. In order to evaluate how much performance is traded for productivity, we compare the Chapel‐based implementations of B&B and UTS to their best‐known counterparts based on traditional OpenMP (intra‐node) and MPI+X (inter‐node). For experimental validation using 4096 processing cores, we consider the permutation flow‐shop scheduling problem for B&B and synthetic literature benchmarks for UTS. The reported results show that P3D‐DFS competes with its OpenMP baselines for coarser‐grained shared‐memory scenarios, and with its MPI+X counterparts for distributed‐memory settings, considering both performance and productivity‐awareness. In the context of this work, this makes Chapel an alternative to OpenMP/MPI+X for exascale programming. Guillaume Helbecque, Jan Gmys, Nouredine Melab, Tiago Carneiro 0001, Pascal Bouvry |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | A Task Offloading Scheme for WAVE Vehicular Clouds and 5G Mobile Edge ComputingabstractVehicular applications are becoming increasingly popular. Some of them are known to be compute-intensive as to require real-time processing. One technique used to improve the performance of these applications is task offloading, which allows computational tasks to be processed on remote servers. In vehicular environments, these servers can be the vehicles themselves or edge servers coupled to base stations. But it is challenging to apply this technique in these environments, where frequent changes in network topology occur, and there is no central coordination point. Therefore, we propose a scheme to improve the performance of computationally intensive applications while dealing with the mobility challenges of vehicular environments. The proposed scheme runs on multi-interface networks WAVE and 5G and searches and prioritizes, through a lightweight greedy algorithm, for servers with high processing capabilities available and shorter distances. The results show that the proposed scheme reduces up to 54.1 % the total offloading time and increases up to 71.8 % the offloading success rate compared to other schemes. In addition, this is one of the first works to compare different task offloading schemes for vehicles using WAVE and 5G multi-interface networks. Alisson Barbosa de Souza, Paulo A. L. Rego, Paulo Henrique Gonçalves Rocha, Tiago Carneiro 0001, José Neuman de Souza |
GLOBECOM | 4 |
| 2020 | Towards ultra-scale Branch-and-Bound using a high-productivity language
Tiago Carneiro 0001, Jan Gmys, Nouredine Melab, Daniel Tuyttens |
Future Gener. Comput. Syst. | 1 |
| 2019 | Detecting Parkinson's disease with sustained phonation and speech signals using machine learning techniques
Jefferson S. Almeida, Pedro Pedrosa Rebouças Filho, Tiago Carneiro 0001, Wei Wei 0006, Robertas Damasevicius, Rytis Maskeliunas, Victor Hugo C. de Albuquerque |
Pattern Recognit. Lett. | 3 |
| 2018 | GPU-accelerated backtracking using CUDA Dynamic ParallelismabstractSummary New GPGPU technologies, such as CUDA Dynamic Parallelism (CDP), can help dealing with recursive patterns of computation, such as divide‐and‐conquer, used by backtracking algorithms. In this paper, we propose a GPU‐accelerated backtracking algorithm using CDP that extends a well‐known parallel backtracking model. The search starts on CPU, processing the search tree until a first cutoff depth. Based on this partial backtracking tree, the algorithm analyzes the memory requirements of subsequent kernel generations. The proposed algorithm performs no dynamic allocation of memory on GPU, unlike related works from the literature. The proposed algorithm has been extensively tested using the N‐Queens Puzzle problem and instances of the Asymmetric Traveling Salesman Problem (ATSP) as test‐cases. The proposed CDP algorithm may, under some conditions, outperform its non‐CDP counterpart by a factor up to 25. But, it may also be up to twice slower. The CDP‐based implementation has much better worst case execution times and makes algorithm's performance less dependent on the tuning of parameters. Compared to other CDP‐based strategies from the literature, the proposed algorithm is on average 8× faster. The proposed algorithm is also hybridized with another CDP‐based strategy from the literature. The combination of strategies is in average 4.5× faster than the related strategy. We also identify some difficulties, limitations, and bottlenecks concerning the CDP programming model which may be useful for helping potential users. Tiago Carneiro 0001, Jan Gmys, Francisco Heron de Carvalho Junior, Nouredine Melab, Daniel Tuyttens |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | A GPU-Based Backtracking Algorithm for Permutation Combinatorial Problems
Tiago Carneiro 0001, Jan Gmys, Nouredine Melab, Francisco Heron de Carvalho Junior, Daniel Tuyttens |
ICA3PP | 1 |
| 2011 | A New Parallel Schema for Branch-and-Bound Algorithms Using GPGPUabstractThis work presents a new parallel procedure designed to process combinatorial B&B algorithms using GPGPU. In our schema we dispatch a number of threads that treats intelligently the massively parallel processors of NVIDIA GeForce graphical units. The strategy is to build sequentially a series of initial searches that can map a subspace of the B&B tree by starting a number of limited threads after achieving a specific level of the tree. The search is then processed massively by DFS. The whole subspace is optimized accordingly to memory and limits of threads and blocks available by the GPU. We compare our results with its OpenMP and Serial versions of the same search schema using explicitly enumeration (all possible solutions) to the Asymmetrical Travelling Salesman Problem's instances. We also show the great superiority of our GPGPU based method. Tiago Carneiro 0001, Albert Einstein Fernandes Muritiba, Marcos Negreiros, Gustavo A. L. de Campos |
SBAC-PAD | 1 |