Rodrigo Huerta

dblp:131/0893 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-0052-7710ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 44% Electronic design automation · 44% Performance modeling and evaluation · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU architecture
0.912025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025
Electronic design automation › hardware verification and test
reverse engineering
0.912025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025
Performance modeling and evaluation
simulation
0.312025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025

Methods — techniques the papers use, named apart from their topics

simulation · 0.9reverse engineering · 0.9microarchitectural modeling · 0.9
YearPublicationVenuePosition
2025 GPU Simulation Acceleration via Parallelization
abstract
Simulating modern GPU architectures with increased core counts and recent workloads can be challenging, even on powerful computing platforms. In this paper, we present a simple approach to parallelize Accel-sim with minimal code changes using OpenMP. In addition, we introduce PaSSMA, a novel technique to optimize the OpenMP for-loop scheduler performance in the different simulated workloads. Moreover, our parallelization technique is deterministic, so the simulator provides exactly the same results for single-threaded and multithreaded simulations. When we run the simulator with 16 CPU cores, we achieve an average speed-up of 6.4x and reach 10x in some workloads.
Rodrigo Huerta, Antonio González 0001
ISPASS1
2025 Dissecting and Modeling the Architecture of Modern GPU Cores
abstract
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations.However, most microarchitectural research in academia relies on simulators that model GPU core architectures based on designs that are more than 15 years old, and differ significantly from modern core architectures.This work reverse engineers the architecture of modern NVIDIA GPU cores, unveiling key aspects of its design and the important role of the compiler in some of its main components.In particular, it reveals how the issue logic works, the structure of the register file and its associated cache, multiple features of the instruction and data memory pipelines.When modeling all these discovered microarchitectural details in a state-of-the-art simulation framework, we show that its accuracy is significantly improved, achieving a 20.58% reduction in mean absolute percentage error (MAPE) on average, which results in a 13.45% MAPE on average with respect to real modern hardware.In addition, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area.
Rodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio González 0001
MICRO1