Gregorio Bernabé

dblp:59/5526 · also Gregorio Bernabé García · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-7265-3508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Characterization of machine learning compilers for LLM inference on NVIDIA GPUs
abstract
Abstract AI inference is conflicted between Performance, developer Productivity, and device Portability–the P3 problem. Machine learning compilers (MLCs) aim to address this, but their ecosystem is fragmented, with tools that each prioritize a different issue. This paper evaluates the deployment trade-offs of PyTorch-based LLMs on NVIDIA GPUs using four intertwined prominent MLC tools: , TensorRT, XLA, and ONNX Runtime. A dual methodology is used, leveraging synthetic PyTorch models to isolate optimizations and end-to-end benchmarks with State-of-the-Art (SOTA) models (TinyLlama-1.1B, Llama-2-7B) to measure real-world performance. Findings reveal that the peak performance of Ahead-Of-Time (AOT) compilation requires architecture-specific tools such as TensorRT-LLM, which are necessary for SOTA LLMs but are unusable for PyTorch models. As for Just-In-Time (JIT) solutions such as and its backends, they are flexible and portable, compatible with all tested models, but they do not consistently accelerate LLMs; therefore, the choice of MLC depends on P3 considerations and model architecture.
Alejandro Carmona-Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.2
2025 A Real Time Cardiomyopathy Detection Tool Using Ml Ensemble Models
abstract
Left Ventricular noncompaction (LVNC) is a recently classified form of cardiomyopathy. Although various methods have been proposed for accurately quantifying trabeculae in the left ventricle (LV), consensus on the optimal approach remains elusive. Previous research introduced DL‐LVTQ, a deep learning solution for trabecular quantification based on a UNet 2D convolutional neural network (CNN) architecture and a graphical user interface (GUI) to streamline its use in clinical workflows. Building on this foundation, this work presents LVNC detector, an enhanced application designed to support cardiologists in the automated diagnosis of LVNC. The application integrates two segmentation models: DL‐LVTQ and ViTUNet, the latter inspired by modern hybrid architectures combining convolutional neural networks (CNNs) and transformer‐based designs. These models, implemented within an ensemble framework, leverage advancements in deep learning to improve the accuracy and robustness of magnetic resonance imaging (MRI) segmentation. Key innovations include multithreading to optimize model loading times and ensemble methods to enhance segmentation consistency across MRI slices. Additionally, the platform‐independent design ensures compatibility with Windows and Linux, eliminating complex setup requirements. The LVNC detector delivers an efficient and user‐friendly solution for LVNC diagnosis. It enables real‐time performance and allows cardiologists to select and compare segmentation models for improved diagnostic outcomes. This work demonstrates how state‐of‐the‐art machine learning techniques can seamlessly integrate into clinical practice to reduce human error and expedite diagnostic processes.
Salvador de Haro, Esteban Becerra, Pilar González-Férez, José M. García 0001, Gregorio Bernabé
IET Softw.5
2024 POAS: a framework for exploiting accelerator level parallelism in heterogeneous environments
abstract
Abstract In the era of heterogeneous computing, a new paradigm called accelerator level parallelism (ALP) has emerged. In ALP, accelerators are used concurrently to provide unprecedented levels of performance and energy efficiency. To reach that there are many problems to be solved, one of the most challenging being co-execution. In this paper, we present a new scheduling framework called POAS, a general method for providing co-execution to applications. Our proposal consists of four steps: predict, optimize, adapt and schedule. With POAS, an unseen application can be executed concurrently in ALP with little effort. We evaluate POAS on a heterogeneous environment consisting of CPUs, GPUs (CUDA cores), and XPUs (Tensor cores) on two different fields, namely linear algebra (matrix multiplication benchmark) and deep learning (convolution benchmark). Our experiments prove that POAS provides excellent performance and completes the tasks within a time very close to the optimal time for the hardware and applications used, with a negligible execution time overhead. Moreover, the POAS predictor performed exceptionally well, achieving very low RMSE values for both use cases. Therefore, POAS can be a valuable tool for fully exploiting ALP and improving overall performance over offloading in heterogeneous settings.
Pablo Antonio Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.2
2023 Matching Linear Algebra and Tensor Code to Specialized Hardware Accelerators
abstract
Dedicated tensor accelerators demonstrate the importance of linear algebra in modern applications. Such accelerators have the potential for impressive performance gains, but require programmers to rewrite code using vendor APIs - a barrier to wider scale adoption. Recent work overcomes this by matching and replacing patterns within code, but such approaches are fragile and fail to cope with the diversity of real-world codes.
Pablo Antonio Martínez, Jackson Woodruff, Jordi Armengol-Estapé, Gregorio Bernabé, José M. García 0001, Michael F. P. O'Boyle
CC4
2022 Applying Intel's oneAPI to a machine learning case study
abstract
Abstract Different technologies and approaches exist to work around the performance portability problem. Companies and academia work together to find a way to preserve performance across heterogeneous hardware using a unified language, one language to rule them all. Intel's oneAPI appears with this idea in mind. In this article, we try the new Intel solution to approach heterogeneous programming, choosing machine learning as our case study. More precisely, we choose Caffe, a machine learning framework that was created six years ago. Nevertheless, how would it be to make Caffe again from the beginning, using a fresh new technology like oneAPI? In terms of not only the ease of programming‐because only one source code would be needed to deploy Caffe to CPUs, GPUs, FPGAs, and accelerators (platforms that oneAPI currently supports)‐but also performance, where oneAPI may be capable of taking advantage of specific hardware automatically. Is Intel's oneAPI ready to take the leap?
Pablo Antonio Martínez, Biagio Peccerillo, Sandro Bartolini, José M. García 0001, Gregorio Bernabé
Concurr. Comput. Pract. Exp.5
2022 HDNN: a cross-platform MLIR dialect for deep neural networks
abstract
Abstract This paper presents HDNN, a proof-of-concept MLIR dialect for cross-platform computing specialized in deep neural networks. As target devices, HDNN supports CPUs, GPUs and TPUs. In this paper, we provide a comprehensive description of the HDNN dialect, outlining how this novel approach aims to solve the $$P^3$$ P 3 problem of parallel programming (portability, productivity, and performance). An HDNN program is device-agnostic, i.e., only the device specifier has to be changed to run a given workload in one device or another. Moreover, HDNN has been designed to be a domain-specific language, which ultimately helps programming productivity. Finally, HDNN relies on optimized libraries for heavy, performance-critical workloads. HDNN has been evaluated against other state-of-the-art machine learning frameworks on all the hardware platforms achieving excellent performance. We conclude that the ideas and concepts used in HDNN can be crucial for designing future generation compilers and programming languages to overcome the challenges of the forthcoming heterogeneous computing era.
Pablo Antonio Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.2
2021 Deploying deep learning approaches to left ventricular non-compaction measurement
Jesús M. Rodríguez-de-Vera, Josefa González-Carrillo, José M. García 0001, Gregorio Bernabé
J. Supercomput.4
2020 A highly accurate method for quantifying LVNC cardiomyophaty
Gregorio Bernabé, José D. Casanova, Guillem Casas, Josefa González-Carrillo
AMIA1
2019 A self-optimized software tool for quantifying the degree of left ventricle hyper-trabeculation
Gregorio Bernabé, José D. Casanova, Javier Cuenca 0001, Josefa González-Carrillo
J. Supercomput.1
2018 Parallel implementations of the 3D fast wavelet transform on a Raspberry Pi 2 cluster
Gregorio Bernabé, Raúl Hernández, Manuel E. Acacio
J. Supercomput.1
2014 Improving an autotuning engine for 3D Fast Wavelet Transform on manycore systems
Gregorio Bernabé, Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez
J. Supercomput.1
2013 Optimizing a 3D-FWT Code in a Heterogeneous Cluster of Multicore CPUs and Manycore GPUs
abstract
Clusters of nodes composed of many core GPUs and multicore CPUs are used to solve scientific problems with high computational requirements. The development and optimization of parallel-heterogeneous codes for these systems is a complex task which requires a deep knowledge of the different components of the hybrid, heterogeneous and hierarchical computational system, and also of the scientific problem to be solved and the different programing paradigms to be used for its efficient solution. Techniques for efficient development and optimization of scientific codes for these systems are needed. This paper presents an analysis of the development and optimization of the 3D-Fast Wavelet Transform (3D-FWT) for a heterogeneous cluster of multicores+GPUs. Different parallel programming paradigms (message passing, shared memory and SIMD GPU) are combined to fully exploit the computing capacity of the different computational elements of the cluster, so resulting in an efficient combination of basic codes developed previously for individual components (individual nodes, multicore or GPU) and an important reduction of the compression time of long video sequences.
Gregorio Bernabé, Javier Cuenca 0001, Domingo Giménez
SBAC-PAD1
2009 A Parallel Implementation of the 2D Wavelet Transform Using CUDA
abstract
There is a multicore platform that is currently concentrating an enormous attention due to its tremendous potential in terms of sustained performance: the NVIDIA Tesla boards. These cards intended for general-purpose computing on graphic processing units (GPGPUs) are used as data-parallel computing devices. They are based on the Computed Unified Device Architecture (CUDA) which is common to the latest NVIDIA GPUs. The bottom line is a multicore platform which provides an enormous potential performance benefit driven by a non-traditional programming model. In this paper we try to provide some insight into the peculiarities of CUDA in order to target scientific computing by means of a specific example. In particular, we show that the parallelization of the two-dimensional fast wavelet transform for the NVIDIA Tesla C870 achieves a speedup of 20.8 for an image size of 8192x8192, when compared with the fastest host-only version implementation using OpenMP and including the data transfers between main memory and device memory.
Joaquín Franco, Gregorio Bernabé, Juan Fernández Peinador, Manuel E. Acacio
PDP2
2009 A lossy 3D wavelet transform for high-quality compression of medical video
Gregorio Bernabé, José M. García 0001, José González 0002
J. Syst. Softw.1
2007 An efficient implementation of a 3D wavelet transform based encoder on hyper-threading technology
Gregorio Bernabé, Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José González 0002
Parallel Comput.1