VLDB 2026 Research / reviewers in the wild / expert
Rodrigo Huerta
dblp:131/0893
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-0052-7710ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 44% Electronic design automation · 44% Performance modeling and evaluation · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU architecture |
0.9 | 1 | 2025 | Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025 |
Electronic design automation › hardware verification and test
reverse engineering |
0.9 | 1 | 2025 | Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025 |
Performance modeling and evaluation
simulation |
0.3 | 1 | 2025 | Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.9reverse engineering · 0.9microarchitectural modeling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPU Simulation Acceleration via ParallelizationabstractSimulating modern GPU architectures with increased core counts and recent workloads can be challenging, even on powerful computing platforms. In this paper, we present a simple approach to parallelize Accel-sim with minimal code changes using OpenMP. In addition, we introduce PaSSMA, a novel technique to optimize the OpenMP for-loop scheduler performance in the different simulated workloads. Moreover, our parallelization technique is deterministic, so the simulator provides exactly the same results for single-threaded and multithreaded simulations. When we run the simulator with 16 CPU cores, we achieve an average speed-up of 6.4x and reach 10x in some workloads. Rodrigo Huerta, Antonio González 0001 |
ISPASS | 1 |
| 2025 | Dissecting and Modeling the Architecture of Modern GPU CoresabstractGPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations.However, most microarchitectural research in academia relies on simulators that model GPU core architectures based on designs that are more than 15 years old, and differ significantly from modern core architectures.This work reverse engineers the architecture of modern NVIDIA GPU cores, unveiling key aspects of its design and the important role of the compiler in some of its main components.In particular, it reveals how the issue logic works, the structure of the register file and its associated cache, multiple features of the instruction and data memory pipelines.When modeling all these discovered microarchitectural details in a state-of-the-art simulation framework, we show that its accuracy is significantly improved, achieving a 20.58% reduction in mean absolute percentage error (MAPE) on average, which results in a 13.45% MAPE on average with respect to real modern hardware.In addition, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area. Rodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio González 0001 |
MICRO | 1 |