EDBT 2026 Demo / reviewers in the wild / expert
Cristóbal A. Navarro
dblp:130/3788
· DBLP profile ↗
13ranked-venue papers
6as first author
10since 2021 · last 2027
0000-0001-7090-9904ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Ray tracing cores for general-purpose computing: A literature review
Enzo Meneses, Cristóbal A. Navarro, Héctor Ferrada, Konstantin Verichev, Cristian Salazar-Concha |
Future Gener. Comput. Syst. | 2 |
| 2026 | Advancing RT core-accelerated fixed-radius nearest neighbor search
Enzo Meneses, Hugo Bec, Cristóbal A. Navarro, Benoît Crespin, Felipe A. Quezada, Nancy Hitschfeld-Kahler, Heinich Porro, Maxime Maria |
Future Gener. Comput. Syst. | 3 |
| 2025 | CAT: Cellular Automata on Tensor CoresabstractCellular automata (CA) are simulation models that can produce complex emergent behaviors from simple local rules. Although state-of-the-art GPU solutions are already fast due to their data-parallel nature, their performance can rapidly degrade in CA with a large neighborhood radius. With the inclusion of tensor cores across the entire GPU ecosystem, interest has grown in finding ways to leverage these fast units outside the field of artificial intelligence, which was their original purpose. In this work, we present CAT, a GPU tensor core approach that can accelerate CA in which the cell transition function acts on a weighted summation of its neighborhood. CAT is evaluated theoretically, using an extended PRAM cost model, as well as empirically using the Larger Than Life (LTL) family of CA as case studies. The results confirm that the cost model is accurate, showing that CAT exhibits constant time throughout the entire radius range$1 \leq r \leq 16$, and its theoretical speedups agree with the empirical results. At low radius$r=1,2$, CAT is competitive and is only surpassed by the fastest state-of-the-art GPU solution. Starting from$r=3$, CAT progressively outperforms all other approaches, reaching speedups of up to$101\times$over a GPU baseline and up to$\sim \!14\times$over the fastest state-of-the-art GPU approach. In terms of energy efficiency, CAT is competitive in the range$1 \leq r \leq 4$and from$r \geq 5$it is the most energy efficient approach. As for performance scaling across GPU architectures, CAT shows a promising trend that, if continues for future generations, it would increase its performance at a higher rate than classical GPU solutions. A CPU version of CAT was also explored, using the recently introduced AMX instructions. Although its performance is still below GPU tensor cores, it is a promising approach as it can still outperform some GPU approaches at large radius. The results obtained in this work put CAT as an approach with great potential for scientists who need to study emerging phenomena in CA with a large neighborhood radius, both in the GPU and in the CPU. Cristóbal A. Navarro, Felipe A. Quezada, Enzo Meneses, Héctor Ferrada, Nancy Hitschfeld-Kahler |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Accelerating range minimum queries with ray tracing cores
Enzo Meneses, Cristóbal A. Navarro, Héctor Ferrada, Felipe A. Quezada |
Future Gener. Comput. Syst. | 2 |
| 2024 | An evaluation of GPU filters for accelerating the 2D convex hull
Roberto Carrasco, Héctor Ferrada, Cristóbal A. Navarro, Nancy Hitschfeld-Kahler |
J. Parallel Distributed Comput. | 3 |
| 2023 | A scalable and energy efficient GPU thread map for m-simplex domains
Cristóbal A. Navarro, Felipe A. Quezada, Benjamin Bustos, Nancy Hitschfeld-Kahler, Rolando Kindelan |
Future Gener. Comput. Syst. | 1 |
| 2023 | Modeling GPU Dynamic Parallelism for self similar density workloads
Felipe A. Quezada, Cristóbal A. Navarro, Miguel Romero 0001 |
Future Gener. Comput. Syst. | 2 |
| 2022 | Squeeze: Efficient compact fractals for tensor core GPUs
Felipe A. Quezada, Cristóbal A. Navarro, Nancy Hitschfeld-Kahler, Benjamin Bustos |
Future Gener. Comput. Syst. | 2 |
| 2022 | Fast kNN query processing over a multi-node GPU environment
Ricardo J. Barrientos, Javier A. Riquelme, Ruber Hernández-García, Cristóbal A. Navarro, Wladimir E. Soto-Silva |
J. Supercomput. | 4 |
| 2021 | GPU Tensor Cores for Fast Arithmetic ReductionsabstractThis article proposes a parallel algorithm for computing the arithmetic reduction of n numbers as a set of matrix-multiply accumulate (MMA) operations that are executed simultaneously by GPU tensor cores. The analysis, assuming tensors of size m x m, shows that the proposed algorithm has a parallel running time of T(n) = 5logm2n and a speedup of S = 45log2m2 over a canonical parallel reduction. Experimental performance results on a Tesla V100 GPU show that the tensor-core based approach is energy efficient and runs up to ~ 3:2× and 2× faster than a standard GPU-based reduction and Nvidia's CUB library, respectively, while keeping the numerical error below 1 percent with respect to a double precision CPU reduction. The chained design of the algorithm allows a flexible configuration of GPU thread-blocks and the optimal values found through experimentation agree with the theoretical ones. The results obtained in this work show that GPU tensor cores are relevant not only for Deep Learning or Linear Algebra computations, but also for applications that require the acceleration of large summations. Cristóbal A. Navarro, Roberto Carrasco, Ricardo J. Barrientos, Javier A. Riquelme, Raimundo Vega |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Efficient GPU thread mapping on embedded 2D fractals
Cristóbal A. Navarro, Felipe A. Quezada, Nancy Hitschfeld-Kahler, Raimundo Vega, Benjamin Bustos |
Future Gener. Comput. Syst. | 1 |
| 2018 | Competitiveness of a Non-Linear Block-Space GPU Thread Map for Simplex DomainsabstractThis work presents and studies the efficiency problem of mapping GPU threads onto simplex domains. A non-linear map$\lambda (\omega)$is formulated based on a block-space enumeration principle that reduces the number of thread-blocks by a factor of approximately$2\times$and$6\times$for 2-simplex and 3-simplex domains, respectively, when compared to the standard approach. Performance results show that$\lambda (\omega)$is competitive and even the fastest map when ran in recent GPU architectures such as the Tesla V100, where it reaches up to$1.5\times$of speedup in 2-simplex tests. In 3-simplex tests, it reaches up to$2.3\times$of speedup for small workloads and up to$1.25\times$for larger ones. The results obtained make$\lambda (\omega)$a useful GPU optimization technique with applications on parallel problems that define all-pairs, all-triplets or nearest neighbors interactions in a 2-simplex or 3-simplex domain. Cristóbal A. Navarro, Matthieu Vernier, Benjamin Bustos, Nancy Hitschfeld-Kahler |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Potential benefits of a block-space GPU approach for discrete tetrahedral domainsabstractThe study of data-parallel domain re-organization and thread-mapping techniques are relevant topics as they can increase the efficiency of GPU computations on spatial discrete domains with non-box-shaped geometry. In this work we study the potential benefits of applying a succinct data re-organization of a tetrahedral data-parallel domain of size O(n3) combined with an efficient block-space GPU map of the form g (λ) : N → N3. Results from the analysis suggest that in theory the combination of these two optimizations produce significant performance improvement as block-based data reorganization allows a coalesced one-to-one correspondence at local thread-space while g(λ) produces an efficient block-space spatial correspondence between groups of data and groups of threads, reducing the number of unnecessary threads from O(n3) to O(n2ρ3) with ρ ∊ O(1). From the analysis, we obtained that a block based succinct data re-organization can provide up to 2× improved performance over a linear data organization while the map can be up to 6× more efficient than a bounding box approach. The results from this work can serve as a useful guide for a more efficient GPU computation on tetrahedral domains found in spin lattice, finite element and special n-body problems, among others. Cristóbal A. Navarro, Benjamin Bustos, Nancy Hitschfeld-Kahler |
CLEI | 1 |