Hector Gerardo Muñoz Hernandez

dblp:261/6238 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-0891-235XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Integrating an open-source soft-GPU overlay with RISC-V control and high-bandwidth memory
abstract
Image and signal processing workloads are widely deployed on Graphics Processing Units (GPUs) for high throughput and on Field-Programmable Gate Arrays (FPGAs) for hardware specialization and energy efficiency. Soft GPU overlays on FPGAs aim to combine these advantages, yet existing solutions often depend on fixed hard processors or impose platform constraints that limit portability. This work extends a popular open-source soft GPGPU overlay to integrate a soft RISC-V control plane and enable compatibility with High-Bandwidth Memory (HBM2). The resulting system can be instantiated on FPGA boards without a hard ARM processor, improving portability, simplifying system integration, and broadening deployability. Across representative image and signal processing kernels, the soft GPGPU achieves geometric-mean speedups of 114.60 × over a scalar soft RISC-V core and 19.72 × over a hard ARM core, demonstrating substantial performance benefits while retaining FPGA reconfigurability. HBM2 integration further benefits bandwidth-sensitive workloads by increasing sustained throughput and reducing the performance bottlenecks associated with off-chip memory access. Collectively, these results indicate that GPU-like programmability and performance can be delivered on reconfigurable platforms without reliance on hard CPU subsystems, providing a portable and scalable foundation for embedded vision and DSP acceleration.
Hector Gerardo Muñoz Hernandez, Mahdi Taheri, Muhammad Ali 0010, Keyvan Shahin, Alireza Syavashi, Diana Göhringer, Marc Reichenbach, Christian Herglotz, Michael Hübner 0001
J. Syst. Archit.1
2021 AITIA: Embedded AI Techniques for Industrial Applications
abstract
Motivated by an increasing interest from startups in embedded Artificial Intelligence (AI) and by their limited expertise, the AITIA Project targets the development of embedded AI techniques for industrial applications. This extended abstract presents the motivation and the solutions being developed towards four use cases: smart sensors, network intrusion detection, driver-assistance systems, and Industry 4.0.
Marcelo Brandalero, Mitko Veleski, Hector Gerardo Muñoz Hernandez, Muhammad Ali 0010, Laurens Le Jeune, Toon Goedemé, Nele Mentens, Jurgen Vandendriessche, Lancelot Lhoest, Bruno da Silva 0001, Abdellah Touhafi, Diana Göhringer, Michael Hübner 0001
FPL3
2021 Towards the Efficient Multi-Platform Execution of Deep Neural Networks
abstract
Modern Systems-on-Chip (SoCs) based on Field-Programmable Gate Arrays (FPGAs) offer users significant flexibility in deciding the best approach to implement Convolutional Neural Networks (CNNs): a) in a fixed, hardwired general-purpose processor, or b) using the programmable logic to implement application-specific processing cores. This thesis proposes an automated toolflow that maps Tensorflow/Keras pre-trained models into different possible platforms: ARM core using the Neon extension and a soft-core GPU for FPGA. CNNs are heterogeneous, meaning that convolutional layers, for example, will have different resource access and computation requirements as the Fully Connected (FC) layers, hinting that different hardware may be optimal for different layer types. After evaluating the performance of different CNNs executed in an ARM Cortex-A9 and the soft-core GPU, it was found that convolutional layers were 5.9x faster in the soft-core GPU than in the ARM core. On the other hand, FC layers were executed faster in the ARM core. As a result, this work proposes a collaborative execution of CNNs using these two platforms together, running the convolutional and maxpooling layers in the soft-core GPU and the FC layers in the ARM core, achieving a speedup of 2x against using only the ARM core. Consequently, this thesis is exploring other mixes of hardware platforms or even using partial reconfiguration techniques.
Hector Gerardo Muñoz Hernandez
FPL1
2020 A Machine Learning Methodology for Cache Memory Design Based on Dynamic Instructions
abstract
Cache memories are an essential component of modern processors and consume a large percentage of their power consumption. Its efficacy depends heavily on the memory demands of the software. Thus, finding the optimal cache for a particular program is not a trivial task and usually involves exhaustive simulation. In this article, we propose a machine learning–based methodology that predicts the optimal cache reconfiguration for any given application, based on its dynamic instructions. Our evaluation shows that our methodology reaches 91.1% accuracy. Moreover, an additional experiment shows that only a small portion of the dynamic instructions (10%) suffices to reach 89.71% accuracy.
Osvaldo Navarro, Jones Yudi Mori, Javier Hoffmann, Hector Gerardo Muñoz Hernandez, Michael Hübner 0001
ACM Trans. Embed. Comput. Syst.4