José-Lorenzo Cruz

dblp:76/954 · also Josep-Llorenc Cruz · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-5325-9153ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 43% Electronic design automation · 43% Performance modeling and evaluation · 13%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU architecture
0.912025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025
Electronic design automation › hardware verification and test
reverse engineering
0.912025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025
Performance modeling and evaluation
simulation
0.312025
Dissecting and Modeling the Architecture of Modern GPU Cores · MICRO 2025
Processor architecture and microarchitecture
register file
0.012000
Multiple-banked register file architectures · ISCA 2000
Processor architecture and microarchitecture
superscalar processor
0.012000
Multiple-banked register file architectures · ISCA 2000

Methods — techniques the papers use, named apart from their topics

simulation · 0.9reverse engineering · 0.9microarchitectural modeling · 0.9
YearPublicationVenuePosition
2025 Dissecting and Modeling the Architecture of Modern GPU Cores
abstract
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations.However, most microarchitectural research in academia relies on simulators that model GPU core architectures based on designs that are more than 15 years old, and differ significantly from modern core architectures.This work reverse engineers the architecture of modern NVIDIA GPU cores, unveiling key aspects of its design and the important role of the compiler in some of its main components.In particular, it reveals how the issue logic works, the structure of the register file and its associated cache, multiple features of the instruction and data memory pipelines.When modeling all these discovered microarchitectural details in a state-of-the-art simulation framework, we show that its accuracy is significantly improved, achieving a 20.58% reduction in mean absolute percentage error (MAPE) on average, which results in a 13.45% MAPE on average with respect to real modern hardware.In addition, we show that the software-based dependence management mechanism included in modern NVIDIA GPUs outperforms a hardware mechanism based on scoreboards in terms of performance and area.
Rodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio González 0001
MICRO3
2024 Exploiting beam search confidence for energy-efficient speech recognition
abstract
Abstract With mobile and embedded devices getting more integrated in our daily lives, the focus is increasingly shifting toward human-friendly interfaces, making automatic speech recognition (ASR) a central player as the ideal means of interaction with machines. ASR is essential for many cognitive computing applications, such as speech-based assistants, dictation systems and real-time language translation. Consequently, interest in speech technology has grown in the last few years, with more systems being proposed and higher accuracy levels being achieved, even surpassing human accuracy. However, highly accurate ASR systems are computationally expensive, requiring on the order of billions of arithmetic operations to decode each second of audio, which conflicts with a growing interest in deploying ASR on edge devices. On these devices, efficient hardware acceleration is key for achieving acceptable performance. In this paper, we propose a technique to improve the energy efficiency and performance of ASR systems, focusing on low-power hardware for edge devices. We focus on optimizing the DNN-based acoustic model evaluation, as we have observed it to be the main bottleneck in popular ASR systems, by leveraging run-time information from the beam search. By doing so, we reduce energy and execution time of the acoustic model evaluation by 25.6 and 25.9 %, respectively, with negligible accuracy loss.
Dennis Pinto, José-María Arnau, Marc Riera, José-Lorenzo Cruz, Antonio González 0001
J. Supercomput.4
2011 A take-home exam to assess professional skills
abstract
Professional Skills, such as the ability to communicate effectively or the ability to gather and integrate information, are not easy to teach or to assess. A traditional exam is not the best way of assessing these skills because it is limited both by time and by the resources students are able to consult. Moreover, in a traditional exam it is difficult to assess if professional skills have been acquired in depth. In this paper we propose to substitute the traditional exam by a take-home exam in which students have more time to solve the questions and are not restricted by the sources they can consult, thereby providing a highly educational task in which students experience a deep learning process. We also analyze what kind of questions should be asked to evaluate professional skills, as well as analyzing the potential drawbacks of these kind of exams (such as inappropriate student behavior). Finally, we show the results of one subject at the Barcelona School of Informatics, in which the take-home exam replaced the traditional exam. This course has been taught over 11 terms with good results.
David López 0001, José-Lorenzo Cruz, Fermín Sánchez, Agustín Fernández
FIE2
2000 Multiple-banked register file architectures
abstract
The register file access time is one of the critical delays in current superscalar processors. Its impact on processor performance is likely to increase in future processor generations, as they are expected to increase the issue width (which implies more register ports) and the size of the instruction window (which implies more registers), and to use some kind of multithreading. Under this scenario, the register file access time could be a dominant delay and a pipelined implementation would be desirable to allow for high clock rates.
José-Lorenzo Cruz, Antonio González 0001, Mateo Valero, Nigel P. Topham
ISCA1