EDBT 2026 Demo / reviewers in the wild / expert
Isaac Gelado
dblp:22/4158
· DBLP profile ↗
18ranked-venue papers
3as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
GPUs and heterogeneous computing · 50% Memory systems · 22% Parallel and multicore computing · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU storage access |
0.7 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
Memory systems › cache management
software-managed cache |
0.7 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
Parallel and multicore computing › synchronization
synchronization mechanisms |
0.4 | 1 | 2019 | Throughput-oriented GPU memory allocation · PPoPP 2019 |
GPUs and heterogeneous computing
GPU architecture |
0.3 | 2 | 2017 | Efficient exception handling support for GPUs · MICRO 2017 Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
GPUs and heterogeneous computing
heterogeneous parallel programming |
0.3 | 2 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
GPUs and heterogeneous computing
GPU memory management |
0.3 | 1 | 2017 | Efficient exception handling support for GPUs · MICRO 2017 |
Memory systems › memory management
virtual memory |
0.3 | 1 | 2017 | Efficient exception handling support for GPUs · MICRO 2017 |
Graph data management
graph analytics |
0.2 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
GPUs and heterogeneous computing
GPU scheduling |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
Processor architecture and microarchitecture › microprocessor
commodity processors |
0.2 | 1 | 2013 | Supercomputing with commodity CPUs: are mobile SoCs ready for HPC? · SC 2013 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
Performance modeling and evaluation
performance monitoring |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2011 | Assessing Accelerator-Based HPC Reverse Time Migration · IEEE Trans. Parallel Distributed Syst. 2011 |
High-performance computing › scientific computing systems
seismic imaging |
0.1 | 1 | 2011 | Assessing Accelerator-Based HPC Reverse Time Migration · IEEE Trans. Parallel Distributed Syst. 2011 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Parallel and multicore computing
parallel programming models |
0.1 | 2 | 2010 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Integrated circuit design
interconnect |
0.1 | 1 | 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator Applications · IEEE Trans. Parallel Distributed Syst. 2015 |
Integrated circuit design
system-on-chip |
0.0 | 1 | 2013 | Supercomputing with commodity CPUs: are mobile SoCs ready for HPC? · SC 2013 |
Compilers and program optimization › parallelization
automatic parallelization |
0.0 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Methods — techniques the papers use, named apart from their topics
concurrent programming techniques · 0.4memory fault forwarding · 0.3exception handling · 0.3pinned buffers · 0.2peer DMA · 0.2double buffering · 0.2performance evaluation · 0.2benchmarking · 0.2monte carlo simulation · 0.1memory trace collection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureabstractGraphics Processing Units (GPUs) have traditionally relied on the host CPU to initiate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their dataset to be processed in a pipelined fashion in the GPU. However, emerging applications such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU initiation of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU-initiated storage removes these overheads from the storage control path and, thus, can potentially support these applications at much higher speed. However, there is a lack of systems architecture and software stack that enable efficient GPU-initiated storage access. This work presents a novel system architecture, BaM, that fills this gap. BaM features a fine-grained software cache to coalesce data storage requests while minimizing I/O traffic amplification. This software cache communicates with the storage system via high-throughput queues that enable the massive number of concurrent threads in modern GPUs to make I/O requests at a high rate to fully utilize the storage devices and the system interconnect. Experimental results show that BaM delivers 1.0x and 1.49x end-to-end speed up for BFS and CC graph analytics benchmarks while reducing hardware costs by up to 21.7x over accessing the graph data from the host memory. Furthermore, BaM speeds up data-analytics workloads by 5.3x over CPU-initiated storage access on the same hardware. Zaid Qureshi, Vikram S. Mailthody, Isaac Gelado, Seungwon Min, Amna Masood, Jeongmin Brian Park, Jinjun Xiong, Chris J. Newburn, Dmitri Vainbrand, I-Hsin Chung, Michael Garland, William J. Dally, Wen-Mei W. Hwu |
ASPLOS (2) | 3 |
| 2019 | Accelerating reduction and scan using tensor core unitsabstractDriven by deep learning, there has been a surge of specialized processors for matrix multiplication, referred to as Tensor Core Units (TCUs). These TCUs are capable of performing matrix multiplications on small matrices (usually 4 × 4 or 16 × 16) to accelerate HPC and deep learning workloads. Although TCUs are prevalent and promise increase in performance and/or energy efficiency, they suffer from over specialization as only matrix multiplication on small matrices is supported. In this paper we express both reduction and scan in terms of matrix multiplication operations and map them onto TCUs. To our knowledge, this paper is the first to try to broaden the class of algorithms expressible as TCU operations and is the first to show benefits of this mapping in terms of: program simplicity, efficiency, and performance. We implemented the reduction and scan algorithms using NVIDIA's V100 TCUs and achieved 89% -- 98% of peak memory copy bandwidth. Our results are orders of magnitude faster (up to 100 × for reduction and 3 × for scan) than state-of-the-art methods for small segment sizes (common in HPC and deep learning applications). Our implementation achieves this speedup while decreasing the power consumption by up to 22% for reduction and 16% for scan. Abdul Dakkak, Cheng Li 0014, Jinjun Xiong, Isaac Gelado, Wen-Mei W. Hwu |
ICS | 4 |
| 2019 | Throughput-oriented GPU memory allocationabstractThroughput-oriented architectures, such as GPUs, can sustain three orders of magnitude more concurrent threads than multicore architectures. This level of concurrency pushes typical synchronization primitives (e.g., mutexes) over their scalability limits, creating significant performance bottlenecks in modules, such as memory allocators, that use them. In this paper, we develop concurrent programming techniques and synchronization primitives, in support of a dynamic memory allocator, that are efficient for use with very high levels of concurrency. Isaac Gelado, Michael Garland |
PPoPP | 1 |
| 2017 | Efficient exception handling support for GPUsabstractOperating systems have long relied on the exception handling mechanism to implement numerous virtual memory features and optimizations. However, today's GPUs have a limited support for exceptions, which prevents implementation of such techniques. The existing solution forwards GPU memory faults to the CPU while the faulting instruction is stalled in the GPU pipeline. This approach prevents preemption of the faulting threads, and results in underutilized hardware resources while the page fault is being resolved by the CPU. Ivan Tanasic, Isaac Gelado, Marc Jordà, Eduard Ayguadé, Nacho Navarro |
MICRO | 2 |
| 2015 | Automatic Parallelization of Kernels in Shared-Memory Multi-GPU NodesabstractIn this paper we present AMGE, a programming framework and runtime system that transparently decomposes GPU kernels and executes them on multiple GPUs in parallel. AMGE exploits the remote memory access capability in modern GPUs to ensure that data can be accessed regardless of its physical location, allowing our runtime to safely decompose and distribute arrays across GPU memories. It optionally performs a compiler analysis that detects array access patterns in GPU kernels. Using this information, the runtime can perform more efficient computation and data distribution configurations than previous works. The GPU execution model allows AMGE to hide the cost of remote accesses if they are kept below 5%. We demonstrate that a thread block scheduling policy that distributes remote accesses through the whole kernel execution further reduces their overhead. Results show 1.98× and 3.89× execution speedups for 2 and 4 GPUs for a wide range of dense computations compared to the original versions on a single GPU. Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B. Jablin, Nacho Navarro, Wen-Mei W. Hwu |
ICS | 3 |
| 2015 | Runtime and Architecture Support for Efficient Data Exchange in Multi-Accelerator ApplicationsabstractHeterogeneous parallel computing applications often process large data sets that require multiple GPUs to jointly meet their needs for physical memory capacity and compute throughput. However, the lack of high-level abstractions in previous heterogeneous parallel programming models force programmers to resort to multiple code versions, complex data copy steps and synchronization schemes when exchanging data between multiple GPU devices, which results in high software development cost, poor maintainability, and even poor performance. This paper describes the HPE runtime system, and the associated architecture support, which enables a simple, efficient programming interface for exchanging data between multiple GPUs through either interconnects or cross-node network interfaces. The runtime and architecture support presented in this paper can also be used to support other types of accelerators. We show that the simplified programming interface reduces programming complexity. The research presented in this paper started in 2009. It has been implemented and tested extensively in several generations of HPE runtime systems as well as adopted into the NVIDIA GPU hardware and drivers for CUDA 4.0 and beyond since 2011. The availability of real hardware that support key HPE features gives rise to a rare opportunity for studying the effectiveness of the hardware support by running important benchmarks on real runtime and hardware. Experimental results show that in a exemplar heterogeneous system, peer DMA and double-buffering, pinned buffers, and software techniques can improve the inter-accelerator data communication bandwidth by 2×. They can also improve the execution speed by 1.6× for a 3D finite difference, 2.5× for 1D FFT, and 1.6× for merge sort, all measured on real hardware. The proposed architecture support enables the HPE runtime to transparently deploy these optimizations under simple portable user code, allowing system designers to freely employ devices of different capabilities. We further argue that simple interfaces such as HPE are needed for most applications to benefit from advanced hardware features in practice. Javier Cabezas, Isaac Gelado, John E. Stone, Nacho Navarro, David Blair Kirk, Wen-Mei W. Hwu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Automatic execution of single-GPU computations across multiple GPUsabstractWe present AMGE, a programming framework and runtime system to decompose data and GPU kernels and execute them on multiple GPUs concurrently. AMGE exploits the remote memory access capability of recent GPUs to guarantee data accessibility regardless of its physical location, thus allowing AMGE to safely decompose and distribute arrays across GPU memories. AMGE also includes a compiler analysis to detect array access patterns in GPU kernels. The runtime uses this information to automatically choose the best computation and data distribution configuration. Through effective use of GPU caches, AMGE achieves good scalability in spite of the limited interconnect bandwidth between GPUs. Results show 1.95x and 3.73x execution speedups for 2 and 4 GPUs for a wide range of dense computations compared to the original versions on a single GPU. Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B. Jablin, Nacho Navarro, Wen-Mei W. Hwu |
PACT | 3 |
| 2014 | Energy Efficient HPC on Embedded SoCs: Optimization Techniques for Mali GPUabstractA lot of effort from academia and industry has been invested in exploring the suitability of low-power embedded technologies for HPC. Although state-of-the-art embedded systems-on-chip (SoCs) inherently contain GPUs that could be used for HPC, their performance and energy capabilities have never been evaluated. Two reasons contribute to the above. Primarily, embedded GPUs until now, have not supported 64-bit floating point arithmetic - a requirement for HPC. Secondly, embedded GPUs did not provide support for parallel programming languages such as OpenCL and CUDA. However, the situation is changing, and the latest GPUs integrated in embedded SoCs do support 64-bit floating point precision and parallel programming models. In this paper, we analyze performance and energy advantages of embedded GPUs for HPC. In particular, we analyze ARM Mali-T604 GPU - the first embedded GPUs with OpenCL Full Profile support. We identify, implement and evaluate software optimization techniques for efficient utilization of the ARM Mali GPU Compute Architecture. Our results show that, HPC benchmarks running on the ARM Mali-T604 GPU integrated into Exynos 5250 SoC, on average, achieve speed-up of 8.7X over a single Cortex-A15 core, while consuming only 32% of the energy. Overall results show that embedded GPUs have performance and energy qualities that make them candidates for future HPC systems. Ivan Grasso, Petar Radojkovic, Nikola Rajovic, Isaac Gelado, Alex Ramírez |
IPDPS | 4 |
| 2014 | Enabling preemptive multiprogramming on GPUsabstractGPUs are being increasingly adopted as compute accelerators in many domains, spanning environments from mobile systems to cloud computing. These systems are usually running multiple applications, from one or several users. However GPUs do not provide the support for resource sharing traditionally expected in these scenarios. Thus, such systems are unable to provide key multiprogrammed workload requirements, such as responsiveness, fairness or quality of service. In this paper, we propose a set of hardware extensions that allow GPUs to efficiently support multiprogrammed GPU workloads. We argue for preemptive multitasking and design two preemption mechanisms that can be used to implement GPU scheduling policies. We extend the architecture to allow concurrent execution of GPU kernels from different user processes and implement a scheduling policy that dynamically distributes the GPU cores among concurrently running kernels, according to their priorities. We extend the NVIDIA GK110 (Kepler) like GPU architecture with our proposals and evaluate them on a set of multiprogrammed workloads with up to eight concurrent processes. Our proposals improve execution time of high-priority processes by 15.6x, the average application turnaround time between 1.5x to 2x, and system fairness up to 3.4x. Ivan Tanasic, Isaac Gelado, Javier Cabezas, Alex Ramírez, Nacho Navarro, Mateo Valero |
ISCA | 2 |
| 2013 | Experiences with mobile processors for energy efficient HPCabstractThe performance of High Performance Computing (HPC) systems is already limited by their power consumption. The majority of top HPC systems today are built from commodity server components that were designed for maximizing the compute performance. The Mont-Blanc project aims at using low-power parts from the mobile domain for HPC. In this paper we present our first experiences with the use of mobile processors and accelerators for the HPC domain based on the research that was performed in the project. We show initial evaluation of NVIDIA Tegra 2 and Tegra 3 mobile SoCs and the NVIDIA Quadro 1000M GPU with a set of HPC micro-benchmarks to evaluate their potential for energy-efficient HPC. Nikola Rajovic, Alejandro Rico, James Vipond, Isaac Gelado, Nikola Puzovic, Alex Ramírez |
DATE | 4 |
| 2013 | Supercomputing with commodity CPUs: are mobile SoCs ready for HPC?abstractIn the late 1990s, powerful economic forces led to the adoption of commodity desktop processors in high-performance computing. This transformation has been so effective that the June 2013 TOP500 list is still dominated by x86. Nikola Rajovic, Paul M. Carpenter, Isaac Gelado, Nikola Puzovic, Alex Ramírez, Mateo Valero |
SC | 3 |
| 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processorsabstractWith the emergence of highly multithreaded architectures, performance monitoring techniques face new challenges in efficiently locating sources of performance discrepancies in the program source code. For example, the state-of-the-art performance counters in highly multithreaded graphics processing units (GPUs) report only the overall occurrences of microarchitecture events at the end of program execution. Furthermore, even if supported, any fine-grained sampling of performance counters will distort the actual program behavior and will make the sampled values inaccurate. On the other hand, it is difficult to achieve high resolution performance information at low sampling rates in the presence of thousands of concurrently running threads. In this paper, we present a novel software-based approach for monitoring the memory hierarchy performance in highly multithreaded general-purpose graphics processors. The proposed analysis is based on memory traces collected for snapshots of an application execution. A trace-based memory hierarchy model with a Monte Carlo experimental methodology generates statistical bounds of performance measures without being concerned about the exact inter-thread ordering of individual events but rather studying the behavior of the overall system. The statistical approach overcomes the classical problem of disturbed execution timing due to fine-grained instrumentation. The approach scales well as we deploy an efficient parallel trace collection technique to reduce the trace generation overhead and a simple memory hierarchy model to reduce the simulation time. The proposed scheme also keeps track of individual memory operations in the source code and can quantify their efficiency with respect to the memory system. A cross-validation of our results shows close agreement with the values read from the hardware performance counters on an NVIDIA Tesla C2050 GPU. Based on the high resolution profile data produced by our model we optimized memory accesses in the sparse matrix vector multiply kernel and achieved speedups ranging from 2.4 to 14.8 depending on the characteristics of the input matrices. Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-Mei W. Hwu |
PPoPP | 2 |
| 2011 | Assessing Accelerator-Based HPC Reverse Time MigrationabstractOil and gas companies trust Reverse Time Migration (RTM), the most advanced seismic imaging technique, with crucial decisions on drilling investments. The economic value of the oil reserves that require RTM to be localized is in the order of 10^{13} dollars. But RTM requires vast computational power, which somewhat hindered its practical success. Although, accelerator-based architectures deliver enormous computational power, little attention has been devoted to assess the RTM implementations effort. The aim of this paper is to identify the major limitations imposed by different accelerators during RTM implementations, and potential bottlenecks regarding architecture features. Moreover, we suggest a wish list, that from our experience, should be included as features in the next generation of accelerators, to cope with the requirements of applications like RTM. We present an RTM algorithm mapping to the IBM Cell/B.E., NVIDIA Tesla and an FPGA platform modeled after the Convey HC-1. All three implementations outperform a traditional processor (Intel Harpertown) in terms of performance (10x), but at the cost of huge development effort, mainly due to immature development frameworks and lack of well-suited programming models. These results show that accelerators are well positioned platforms for this kind of workload. Due to the fact that our RTM implementation is based on an explicit high order finite difference scheme, some of the conclusions of this work can be extrapolated to applications with similar numerical scheme, for instance, magneto-hydrodynamics or atmospheric flow simulations. Mauricio Araya-Polo, Javier Cabezas, Mauricio Hanzich, Miquel Pericàs, Félix Rubio, Isaac Gelado, Muhammad Shafiq 0003, Enric Morancho, Nacho Navarro, Eduard Ayguadé, José María Cela, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systemsabstractHeterogeneous computing combines general purpose CPUs with accelerators to efficiently execute both sequential control-intensive and data-parallel phases of applications. Existing programming models for heterogeneous computing rely on programmers to explicitly manage data transfers between the CPU system memory and accelerator memory. Isaac Gelado, Javier Cabezas, Nacho Navarro, John E. Stone, Sanjay J. Patel, Wen-Mei W. Hwu |
ASPLOS | 1 |
| 2009 | Predictive Runtime Code Scheduling for Heterogeneous Architectures
Víctor J. Jiménez, Lluís Vilanova, Isaac Gelado, Marisa Gil, Grigori Fursin, Nacho Navarro |
HiPEAC | 3 |
| 2008 | CUBA: an architecture for efficient CPU/co-processor data communicationabstractData-parallel co-processors have the potential to improve performance in highly parallel regions of code when coupled to a general-purpose CPU. However, applications often have to be modified in non-intuitive and complicated ways to mitigate the cost of data marshalling between the CPU and the co-processor. In some applications the overheads cannot be amortized and co-processors are unable to provide benefit. The additional effort and complexity of incorporating co-processors makes it difficult, if not impossible, to effectively utilize co-processors in large applications. Isaac Gelado, John H. Kelm, Shane Ryoo, Steven S. Lumetta, Nacho Navarro, Wen-Mei W. Hwu |
ICS | 1 |
| 2007 | CIGAR: Application Partitioning for a CPU/Coprocessor Architecture
John H. Kelm, Isaac Gelado, Mark J. Murphy, Nacho Navarro, Steven S. Lumetta, Wen-Mei W. Hwu |
PACT | 2 |
| 2007 | Implicitly Parallel Programming Models for Thousand-Core MicroprocessorsabstractThis paper argues for an implicitly parallel programming model for many-core microprocessors, and provides initial technical approaches towards this goal. In an implicitly parallel programming model, programmers maximize algorithm-level parallelism, express their parallel algorithms by asserting high-level properties on top of a traditional sequential programming language, and rely on parallelizing compilers and hardware support to perform parallel execution under the hood. In such a model, compilers and related tools require much more advanced program analysis capabilities and programmer assertions than what are currently available so that a comprehensive understanding of the input program's concurrency can be derived. Such an understanding is then used to drive automatic or interactive parallel code generation tools for a diverse set of parallel hardware organizations. The chip-level architecture and hardware should maintain parallel execution state in such a way that a strictly sequential execution state can always be derived for the purpose of verifying and debugging the program. We argue that implicitly parallel programming models are critical for addressing the software development crises and software scalability challenges for many-core microprocessors. Wen-Mei W. Hwu, Shane Ryoo, Sain-Zee Ueng, John H. Kelm, Isaac Gelado, Sam S. Stone, Robert E. Kidd, Sara S. Baghsorkhi, Aqeel Mahesri, Stephanie C. Tsao, Nacho Navarro, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel |
DAC | 5 |