VLDB 2026 Research / reviewers in the wild / expert
Javier Cabezas
dblp:96/3926
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-3335-8036ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-authorArtificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 55% High-performance computing · 15% Processor architecture and microarchitecture · 11% |
Topics — the 8 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
heterogeneous parallel programming |
0.3 | 2 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
GPUs and heterogeneous computing
GPU scheduling |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2011 | Assessing Accelerator-Based HPC Reverse Time Migration · IEEE Trans. Parallel Distributed Syst. 2011 |
High-performance computing › scientific computing systems
seismic imaging |
0.1 | 1 | 2011 | Assessing Accelerator-Based HPC Reverse Time Migration · IEEE Trans. Parallel Distributed Syst. 2011 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Integrated circuit design
interconnect |
0.1 | 1 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Methods — techniques the papers use, named apart from their topics
pinned buffers · 0.2peer DMA · 0.2double buffering · 0.2finite difference scheme · 0.1data transfer management · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Combining Multi-Agent Systems and Subjective Logic to Develop Decision Support Systems
César González-Fernández, Javier Cabezas, Alberto Fernández-Isabel, Isaac Martín de Diego |
IPMU (1) | 2 |
| 2020 | Knowledge-based framework for estimating the relevance of scientific articles
Alberto Fernández-Isabel, Adrián Alonso, Javier Cabezas, Isaac Martín de Diego, J. F. J. Viseu Pinheiro |
Expert Syst. Appl. | 3 |
| 2015 | Automatic Parallelization of Kernels in Shared-Memory Multi-GPU NodesabstractIn this paper we present AMGE, a programming framework and runtime system that transparently decomposes GPU kernels and executes them on multiple GPUs in parallel. AMGE exploits the remote memory access capability in modern GPUs to ensure that data can be accessed regardless of its physical location, allowing our runtime to safely decompose and distribute arrays across GPU memories. It optionally performs a compiler analysis that detects array access patterns in GPU kernels. Using this information, the runtime can perform more efficient computation and data distribution configurations than previous works. The GPU execution model allows AMGE to hide the cost of remote accesses if they are kept below 5%. We demonstrate that a thread block scheduling policy that distributes remote accesses through the whole kernel execution further reduces their overhead. Results show 1.98× and 3.89× execution speedups for 2 and 4 GPUs for a wide range of dense computations compared to the original versions on a single GPU. Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B. Jablin, Nacho Navarro, Wen-Mei W. Hwu |
ICS | 1 |
| 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator ApplicationsabstractHeterogeneous parallel computing applications often process large data sets that require multiple GPUs to jointly meet their needs for physical memory capacity and compute throughput. However, the lack of high-level abstractions in previous heterogeneous parallel programming models force programmers to resort to multiple code versions, complex data copy steps and synchronization schemes when exchanging data between multiple GPU devices, which results in high software development cost, poor maintainability, and even poor performance. This paper describes the HPE runtime system, and the associated architecture support, which enables a simple, efficient programming interface for exchanging data between multiple GPUs through either interconnects or cross-node network interfaces. The runtime and architecture support presented in this paper can also be used to support other types of accelerators. We show that the simplified programming interface reduces programming complexity. The research presented in this paper started in 2009. It has been implemented and tested extensively in several generations of HPE runtime systems as well as adopted into the NVIDIA GPU hardware and drivers for CUDA 4.0 and beyond since 2011. The availability of real hardware that support key HPE features gives rise to a rare opportunity for studying the effectiveness of the hardware support by running important benchmarks on real runtime and hardware. Experimental results show that in a exemplar heterogeneous system, peer DMA and double-buffering, pinned buffers, and software techniques can improve the inter-accelerator data communication bandwidth by 2×. They can also improve the execution speed by 1.6× for a 3D finite difference, 2.5× for 1D FFT, and 1.6× for merge sort, all measured on real hardware. The proposed architecture support enables the HPE runtime to transparently deploy these optimizations under simple portable user code, allowing system designers to freely employ devices of different capabilities. We further argue that simple interfaces such as HPE are needed for most applications to benefit from advanced hardware features in practice. Javier Cabezas, Isaac Gelado, John E. Stone, Nacho Navarro, David Blair Kirk, Wen-Mei W. Hwu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Automatic execution of single-GPU computations across multiple GPUsabstractWe present AMGE, a programming framework and runtime system to decompose data and GPU kernels and execute them on multiple GPUs concurrently. AMGE exploits the remote memory access capability of recent GPUs to guarantee data accessibility regardless of its physical location, thus allowing AMGE to safely decompose and distribute arrays across GPU memories. AMGE also includes a compiler analysis to detect array access patterns in GPU kernels. The runtime uses this information to automatically choose the best computation and data distribution configuration. Through effective use of GPU caches, AMGE achieves good scalability in spite of the limited interconnect bandwidth between GPUs. Results show 1.95x and 3.73x execution speedups for 2 and 4 GPUs for a wide range of dense computations compared to the original versions on a single GPU. Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B. Jablin, Nacho Navarro, Wen-Mei W. Hwu |
PACT | 1 |
| 2014 | Enabling preemptive multiprogramming on GPUsabstractGPUs are being increasingly adopted as compute accelerators in many domains, spanning environments from mobile systems to cloud computing. These systems are usually running multiple applications, from one or several users. However GPUs do not provide the support for resource sharing traditionally expected in these scenarios. Thus, such systems are unable to provide key multiprogrammed workload requirements, such as responsiveness, fairness or quality of service. In this paper, we propose a set of hardware extensions that allow GPUs to efficiently support multiprogrammed GPU workloads. We argue for preemptive multitasking and design two preemption mechanisms that can be used to implement GPU scheduling policies. We extend the architecture to allow concurrent execution of GPU kernels from different user processes and implement a scheduling policy that dynamically distributes the GPU cores among concurrently running kernels, according to their priorities. We extend the NVIDIA GK110 (Kepler) like GPU architecture with our proposals and evaluate them on a set of multiprogrammed workloads with up to eight concurrent processes. Our proposals improve execution time of high-priority processes by 15.6x, the average application turnaround time between 1.5x to 2x, and system fairness up to 3.4x. Ivan Tanasic, Isaac Gelado, Javier Cabezas, Alex Ramírez, Nacho Navarro, Mateo Valero |
ISCA | 3 |
| 2011 | Assessing Accelerator-Based HPC Reverse Time MigrationabstractOil and gas companies trust Reverse Time Migration (RTM), the most advanced seismic imaging technique, with crucial decisions on drilling investments. The economic value of the oil reserves that require RTM to be localized is in the order of 10^{13} dollars. But RTM requires vast computational power, which somewhat hindered its practical success. Although, accelerator-based architectures deliver enormous computational power, little attention has been devoted to assess the RTM implementations effort. The aim of this paper is to identify the major limitations imposed by different accelerators during RTM implementations, and potential bottlenecks regarding architecture features. Moreover, we suggest a wish list, that from our experience, should be included as features in the next generation of accelerators, to cope with the requirements of applications like RTM. We present an RTM algorithm mapping to the IBM Cell/B.E., NVIDIA Tesla and an FPGA platform modeled after the Convey HC-1. All three implementations outperform a traditional processor (Intel Harpertown) in terms of performance (10x), but at the cost of huge development effort, mainly due to immature development frameworks and lack of well-suited programming models. These results show that accelerators are well positioned platforms for this kind of workload. Due to the fact that our RTM implementation is based on an explicit high order finite difference scheme, some of the conclusions of this work can be extrapolated to applications with similar numerical scheme, for instance, magneto-hydrodynamics or atmospheric flow simulations. Mauricio Araya-Polo, Javier Cabezas, Mauricio Hanzich, Miquel Pericàs, Félix Rubio, Isaac Gelado, Muhammad Shafiq 0003, Enric Morancho, Nacho Navarro, Eduard Ayguadé, José María Cela, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systemsabstractHeterogeneous computing combines general purpose CPUs with accelerators to efficiently execute both sequential control-intensive and data-parallel phases of applications. Existing programming models for heterogeneous computing rely on programmers to explicitly manage data transfers between the CPU system memory and accelerator memory. Isaac Gelado, Javier Cabezas, Nacho Navarro, John E. Stone, Sanjay J. Patel, Wen-Mei W. Hwu |
ASPLOS | 2 |