Daniel Jiménez-González

dblp:59/1032 · DBLP profile ↗
← Back
37ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-6064-7883ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021
YearPublicationVenuePosition
2026 OmpSs@FPGA: An Open Source Framework for Programming FPGA Clusters
abstract
Field Programmable Gate Arrays (FPGA) have seen an increase in popularity in High-Performance Computing (HPC) environments due to their high flexibility and energy efficiency. However, the adoption of these devices for HPC application acceleration is relatively new, and their application is still focused on niche applications mainly due to the programmability issues FPGAs have traditionally faced. In this work, we present an open source toolchain with support for different FPGA devices and memory topologies, and for implicit and explicit message-passing multi-node and multi-FPGA programming models, which, to the best of our knowledge, is the first of its kind available. The presented toolchain tries to address the increased complexity of programming newer multi-die devices using High-Bandwidth Memory (HBM) by proposing and evaluating HBM design strategies as well as inter-FPGA communication. In addition to presenting the framework characteristics, we evaluate its performance results over a cluster of FPGAs for a set of HPC application benchmarks delivering performance improvements over previous implementations in both single and multi-node versions. The presented framework is an ongoing effort to solve the lack of mature tools and ecosystems that allow programming large clusters of FPGAs and consequently make scaling applications across multiple FPGA devices a simpler task.
Antonio Filgueras, Ismael el Basli, Juan Miguel De Haro Ruiz, Miquel Vidal, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell
ACM Trans. Reconfigurable Technol. Syst.6
2025 Parallel and Distributed Protein Processing for 3D-protein Pattern Discovery and Clustering
abstract
The discovery and clustering of three-dimensional protein patterns in a set of protein structures, without relying on predefined search patterns, can be highly beneficial for predicting the functions of unknown proteins and facilitating rational multi-target drug design. This work introduces a novel OpenMP parallelization of the 3D-PP algorithm using explicit and nested tasks, which balances data sharing synchronization and load unbalance, improving previous implementations based on implicit tasks. Given the vast number of protein structures available in the Protein Data Bank (over 231,000 from PDB and more than 1,068,000 from AlphaFold), processing large datasets locally can be constrained by available resources such as main memory and secondary storage. To address this challenge, we propose parallel approaches to the 3D-PP algorithm for distributed memory systems (using MPI) and hybrid systems (combining MPI and OpenMP). The evaluated strategies distribute the workload based on the entire protein structure. Experimental results using a dataset of 8,344 protein structures show that the new OpenMP taskified version is 1.6x faster than the previous best OpenMP implementation. Moreover, the hybrid MPI+OpenMP version achieves a speedup of up to 162.5x, demonstrating its scalability and efficiency for large-scale protein structure analysis.
Alejandro Valdés-Jiménez, Gabriel Núñez-Vivanco, Daniel Jiménez-González
eScience3
2025 Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label Propagation
abstract
Task Scheduling is a popular technique for exploiting parallelism in modern computing systems.In particular, HW-accelerated Task Scheduling has been shown to be effective at improving the performance of fine-grained workloads by dynamically assigning tasks to cores based on their data dependencies with minimal overhead, allowing the handling of tasks with execution times in the order of thousands of cycles.However, the performance of applications assisted by accelerated Task Scheduling is limited by the fact that once a task has all its dependencies fulfilled, it is typically executed on the first available core, which might not be locality-optimal.We thus propose a novel approach to Task Scheduling that leverages HW-accelerated Label Propagation (LP), a graph clustering algorithm, to group tasks with intersecting data patterns such that they are executed on the same core.We show that our approach can significantly improve the performance of task-based applications, improving overall program execution times by up to 1.50× while simultaneously reducing average task sizes by up to 1.81×, augmenting both synthetic benchmarks and real-world applications running on a 24-core RISC-V processor mapped to the Alveo U55C FPGA.These gains rely heavily on the low-latency nature of our proposed label propagation accelerator, which will typically cluster dynamic task graphs in under 300 cycles, up to 581× faster than an equivalent software implementation.Furthermore, by ensuring that ideal placement predictions are used as a hint rather than a hard constraint, we allow the system to benefit from improved data locality for memory-intensive applications while also maintaining high core utilization in compute-bound scenarios.Our results hence demonstrate the potential of HW-accelerated label propagation to improve the performance of Task Scheduling systems with low-latency, dynamic data locality optimization.
Lucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo, Giacomo Pedretti, Jim Ignowski, Michael Frank 0008, Xavier Martorell, Daniel Jiménez-González, Carlos Álvarez 0001
MICRO9
2024 The TEXTAROSSA Project: Cool all the Way Down to the Hardware
abstract
The TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research.
Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo
DSD21
2024 Automated parallel execution of distributed task graphs with FPGA clusters
abstract
Over the years, Field Programmable Gate Arrays (FPGA) have been gaining popularity in the High Performance Computing (HPC) field, because their reconfigurability enables very fine-grained optimizations with low energy cost. However, the different characteristics, architectures, and network topologies of the clusters have hindered the use of FPGAs at a large scale. In this work, we present an evolution of OmpSs@FPGA, a high-level task-based programming model and extension to OmpSs-2, that aims at unifying all FPGA clusters by using a message-passing interface that is compatible with FPGA accelerators. These accelerators are programmed with C/C++ pragmas, and synthesized with High-Level Synthesis tools. The new framework includes a custom protocol to exchange messages between FPGAs, agnostic of the architecture and network type. On top of that, we present a new communication paradigm called Implicit Message Passing (IMP), where the user does not need to call any message-passing API. Instead, the runtime automatically infers data movement between nodes. We test classic message passing and IMP with three benchmarks on two different FPGA clusters. One is cloudFPGA, a disaggregated platform with AMD FPGAs that are only connected to the network through UDP/TCP/IP. The other is ESSPER, composed of CPU-attached Intel FPGAs that have a private network at the ethernet level. In both cases, we demonstrate that IMP with OmpSs@FPGA can increase the productivity of FPGA programmers at a large scale thanks to simplifying communication between nodes, without limiting the scalability of applications. We implement the N-body, Heat simulation and Cholesky decomposition benchmarks, and show that FPGA clusters get 2.6x and 2.4x better performance per watt than a CPU-only supercomputer for N-body and Heat.
Juan Miguel De Haro Ruiz, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Tomohiro Ueno, Kentaro Sano, Burkhard Ringlein, François Abel, Beat Weiss
Future Gener. Comput. Syst.3
2024 Enabling HW-Based Task Scheduling in Large Multicore Architectures
abstract
Dynamic Task Scheduling is an enticing programming model aiming to ease the development of parallel programs with intrinsically irregular or data-dependent parallelism. The performance of such solutions relies on the ability of the Task Scheduling HW/SW stack to efficiently evaluate dependencies at runtime and schedule work to available cores. Traditional SW-only systems implicate scheduling overheads of around 30K processor cycles per task, which severely limit the (core count,task granularity) combinations that they might adequately handle. Previous work on HW-accelerated Task Scheduling has shown that such systems might support high performance scheduling on processors with up to eight cores, but questions remained regarding the viability of such solutions to support the greater number of cores now frequently found in high-end SMP systems.The present work presents an FPGA-proven, tightly-integrated, Linux-capable, 30-core RISC-V system with hardware accelerated Task Scheduling. We use this implementation to show that HW Task Scheduling can still offer competitive performance at such high core count, and describe how this organization includes hardware and software optimizations that make it even more scalable than previous solutions. Finally, we outline ways in which this architecture could be augmented to overcome inter-core communication bottlenecks, mitigating the cache-degradation effects usually involved in the parallelization of highly optimized serial code.
Lucas Morais, Carlos Álvarez 0001, Daniel Jiménez-González, Juan Miguel De Haro Ruiz, Guido Araujo, Michael Frank 0008, Alfredo Goldman, Xavier Martorell
IEEE Trans. Computers3
2024 Parallel Algorithm for Discovering and Comparing Three-Dimensional Proteins Patterns
abstract
Identifying conserved (similar) three-dimensional patterns among a set of proteins can be helpful for the rational design of polypharmacological drugs. Some available tools allow this identification from a limited perspective, only considering the available information, such as known binding sites or previously annotated structural motifs. Thus, these approaches do not look for similarities among all putative orthosteric and or allosteric bindings sites between protein structures. To overcome this tech-weakness Geomfinder was developed, an algorithm for the estimation of similarities between all pairs of three-dimensional amino acids patterns detected in any two given protein structures, which works without information about their known patterns. Even though Geomfinder is a functional alternative to compare small structural proteins, it is computationally unfeasible for the case of large protein processing and the algorithm needs to improve its performance. This work presents several parallel versions of the Geomfinder to exploit SMPs, distributed memory systems, hybrid version of SMP and distributed memory systems, and GPU based systems. Results show significant improvements in performance as compared to the original version and achieve up to 24.5x speedup when analyzing proteins of average size and up to 95.4x in larger proteins.
Alejandro Valdés-Jiménez, Miguel Reyes-Parada, Gabriel Núñez-Vivanco, Daniel Jiménez-González
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Improving Performance of HPC Kernels on FPGAs Using High-Level Resource Management
abstract
In state-of-the-art FPGA, especially in chiplet-based devices, place and route has become an important challenge due to an increase in device size and complexity. In the same way, off-chip memory resources have grown in size and number of memory modules. Making efficient use of them has become a difficult task.
Antonio Filgueras, Miquel Vidal, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell
FCCM3
2023 Improving the Discovery and Clustering of Three-Dimensional Protein Patterns with OpenMP
abstract
The discovery of conserved three-dimensional (3D) amino-acid patterns among a set of protein structures can be useful, for instance, to predict the functions of unknown proteins or for the rational design of multi-target drugs. There are several applications that perform a three-dimensional search of patterns in the structures of proteins. However, discovering conserved 3D patterns in a set of proteins with no other baseline patterns is a challenge. In this paper, we analyze and improve a state-of-the-art algorithm, 3D-PP, that implements this discovery. In this algorithm, the 3D patterns are detected and clustered using the root mean square deviation value, measured among each pair of 3D patterns (topological variability indicator). Even when 3D-PP deals with this task, the simultaneous processing of high amounts of proteins becomes a computational challenge with the size and the number of proteins to be evaluated. In this work, we present and analyze different shared memory parallel strategies of 3D-PP, using OpenMP. Those strategies improve the overall performance of the original implementation by reducing parallel load unbalance among threads and overall increasing parallelism. The results show significant performance improvements compared to the original version, achieving up to 13x speedup for a small number of proteins and 17.7× for a larger set.
Alejandro Valdés-Jiménez, Miguel Reyes-Parada, Gabriel Núñez-Vivanco, Fabio Durán-Verdugo, Daniel Jiménez-González
SBAC-PAD5
2022 Towards Reconfigurable Accelerators in HPC: Designing a Multipurpose eFPGA Tile for Heterogeneous SoCs
abstract
The goal of modern high performance computing platforms is to combine low power consumption and high throughput. Within the European Processor Initiative (EPI), such an SoC platform to meet the novel exascale requirements is built and investigated. As part of this project, we introduce an embedded Field Programmable Gate Array (eFPGA), adding flexibility to accelerate various workloads. In this article, we show our approach to design the eFPGA tile that supports the EPI SoC. While eFPGAs are inherently reconfigurable, their initial design has to be determined for tape-out. The design space of the eFPGA is explored and evaluated with different configurations of two HPC workloads, covering control and dataflow heavy applications. As a result, we present a well-balanced eFPGA design that can host several use cases and potential future ones by allocating 1% of the total EPI SoC area. Finally, our simulation results of the architectures on the eFPGA show great performance improvements over their software counterparts.
Tim Hotfilter, Fabian Kreß, Fabian Kempf, Jürgen Becker 0001, Juan Miguel De Haro Ruiz, Daniel Jiménez-González, Miquel Moretó, Carlos Álvarez 0001, Jesús Labarta, Imen Baili
DATE6
2022 OmpSs@cloudFPGA: An FPGA Task-Based Programming Model with Message Passing
abstract
Nowadays, a new parallel paradigm for energy-efficient heterogeneous hardware infrastructures is required to achieve better performance at a reasonable cost on high-performance computing applications. Under this new paradigm, some application parts are offloaded to specialized accelerators that run faster or are more energy-efficient than CPUs. Field-Programmable Gate Arrays (FPGA) are one of those types of accelerators that are becoming widely available in data centers. This paper proposes OmpSs@cloudFPGA, which includes novel extensions to parallel task-based programming models that enable easy and efficient programming of heterogeneous clusters with FPGAs. The programmer only needs to annotate, with OpenMP-like pragmas, the tasks of the application that should be accelerated in the cluster of FPGAs. Next, the proposed programming model framework automatically extracts parts annotated with High-Level Synthesis (HLS) pragmas and synthesizes them into hardware accelerator cores for FPGAs. Additionally, our extensions include and support two novel features: 1) FPGA-to-FPGA direct communication using a Message Passing Interface (MPI) similar Application Programming Interface (API) with one-to-one and collective communications to alleviate host communication channel bottleneck, and 2) creating and spawning work from inside the FPGAs to their own accelerator cores based on an MPI rank-like identification. These features break the classical host-accelerator model, where the host (typically the CPU) generates all the work and distributes it to each accelerator. We also present an evaluation of OmpSs@cloudFPGA for different parallel strategies of the N-Body application on the IBM cloudFPGA research platform. Results show that for cluster sizes up to 56 FPGAs, the performance scales linearly. To the best of our knowledge, this is the best performance obtained for N-body over FPGA platforms, reaching 344 Gpairs/s with 56 FPGAs. Finally, we compare the performance and power consumption of the proposed approach with the ones obtained by a classical execution on the MareNostrum 4 supercomputer, demonstrating that our FPGA approach reduces power consumption by an order of magnitude.
Juan Miguel De Haro Ruiz, Rubén Cano, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, François Abel, Burkhard Ringlein, Beat Weiss
IPDPS4
2021 OmpSs@FPGA Framework for High Performance FPGA Computing
abstract
This article presents the new features of the OmpSs@FPGA framework. OmpSs is a data-flow programming model that supports task nesting and dependencies to target asynchronous parallelism and heterogeneity. OmpSs@FPGA is the extension of the programming model addressed specifically to FPGAs. OmpSs environment is built on top of Mercurium source to source compiler and Nanos++ runtime system. To address FPGA specifics Mercurium compiler implements several FPGA related features as local variable caching, wide memory accesses or accelerator replication. In addition, part of the Nanos++ runtime has been ported to hardware. Driven by the compiler this new hardware runtime adds new features to FPGA codes, such as task creation and dependence management, providing both performance increases and ease of programming. To demonstrate these new capabilities, different high performance benchmarks have been evaluated over different FPGA platforms using the OmpSs programming model. The results demonstrate that programs that use the OmpSs programming model achieve very competitive performance with low to moderate porting effort compared to other FPGA implementations.
Juan Miguel De Haro Ruiz, Jaume Bosch, Antonio Filgueras, Miquel Vidal, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta
IEEE Trans. Computers5
2020 Breaking master-slave model between host and FPGAs
abstract
This paper proposes to enhance current task-based programming models by breaking their current master-slave approach between the main processor and its hardware accelerators. As a proof-of-concept, it presents an extension of the [email protected] toolchain that allows the tasks offloaded into the FPGA to create and synchronize nested tasks on their own without involving the host. Those FPGA spawned tasks may target the host to execute code not suitable for the FPGA, like system calls or I/O operations; or target other kernel accelerators inside the same FPGA. In addition to the programmability benefits of this new feature, the proposed system presents significant performance improvements and a better productivity over the classical master-slave approach.
Jaume Bosch, Miquel Vidal, Antonio Filgueras, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé
PPoPP5
2020 Asynchronous runtime with distributed manager for task-based programming models
Jaume Bosch, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé
Parallel Comput.3
2019 A Hardware Runtime for Task-Based Programming Models
abstract
Task-based programming models such as OpenMP 5.0 and OmpSs are simple to use and powerful enough to exploit task parallelism of applications over multicore, manycore and heterogeneous systems. However, their software-only runtimes introduce relevant overhead when targeting fine-grained tasks, resulting in performance losses. To overcome this drawback, we present a hardware runtime Picos++ that accelerates critical runtime functions such as task dependence analysis, nested task support, and heterogeneous task scheduling. As a proof-of-concept, the Picos++ hardware runtime has been integrated with a compiler infrastructure that supports parallel task-based programming models. A FPGA SoC running Linux OS has been used to implement the hardware accelerated part of Picos++, integrated with a heterogeneous system composed of 4 symmetric multiprocessor (SMP) cores and several hardware functional accelerators (HwAccs) for task execution. Results show significant improvements on energy and performance compared to state-of-the-art parallel software-only runtimes. With Picos++, applications can achieve up to 7.6x speedup and save up to 90 percent of energy, when using 4 threads and up to 4 HwAccs, and even reach a speedup of 16x over the software alternative when using 12 HwAccs and small tasks.
Xubin Tan, Jaume Bosch, Carlos Álvarez 0001, Daniel Jiménez-González, Eduard Ayguadé, Mateo Valero
IEEE Trans. Parallel Distributed Syst.4
2018 LEGaTO: towards energy-efficient, secure, fault-tolerant toolset for heterogeneous computing
abstract
LEGaTO is a three-year EU H2020 project which started in December 2017. The LEGaTO project will leverage task-based programming models to provide a software ecosystem for Made-in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC.
Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel Jiménez-González, Carlos Álvarez 0001, Behzad Salami 0001, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Micha vor dem Berge, Gunnar Billung-Meyer, Stefan Krupop, Wolfgang Christmann, Frank Klawonn, Amani Mihklafi, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Vesna Nowack, Christof Fetzer, Jens Hagemeyer, Thorsten Jungeblut, Nils Kucza, Martin Kaiser, Mario Porrmann, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber
CF7
2018 Application Acceleration on FPGAs with OmpSs@FPGA
abstract
OmpSs@FPGA is the flavor of OmpSs that allows offloading application functionality to FPGAs. Similarly to OpenMP, it is based on compiler directives. While the OpenMP specification also includes support for heterogeneous execution, we use OmpSs and OmpSs@FPGA as prototype implementation to develop new ideas for OpenMP. OmpSs@FPGA implements the tasking model with runtime support to automatically exploit all SMP and FPGA resources available in the execution platform. In this paper, we present the OmpSs@FPGA ecosystem, based on the Mercurium compiler and the Nanos++ runtime system. We show how the applications are transformed to run on the SMP cores and the FPGA. The application kernels defined as tasks to be accelerated, using the OmpSs directives are: 1) transformed by the compiler into kernels connected with the proper synchronization and communication ports, 2) extracted to intermediate files, 3) compiled through the FPGA vendor HLS tool, and 4) used to configure the FPGA. Our Nanos++ runtime system schedules the application tasks on the platform, being able to use the SMP cores and the FPGA accelerators at the same time. We present the evaluation of the OmpSs@FPGA environment with the Matrix Multiplication, Cholesky and N-Body benchmarks, showing the internal details of the execution, and the performance obtained on a Zynq Ultrascale+ MPSoC (up to 128x). The source code uses OmpSs@FPGA annotations and different Vivado HLS optimization directives are applied for acceleration.
Jaume Bosch, Xubin Tan, Antonio Filgueras, Miquel Vidal, Marc Mateu, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta
FPT6
2018 LightDock: a new multi-scale approach to protein-protein docking
abstract
Motivation: Computational prediction of protein-protein complex structure by docking can provide structural and mechanistic insights for protein interactions of biomedical interest. However, current methods struggle with difficult cases, such as those involving flexible proteins, low-affinity complexes or transient interactions. A major challenge is how to efficiently sample the structural and energetic landscape of the association at different resolution levels, given that each scoring function is often highly coupled to a specific type of search method. Thus, new methodologies capable of accommodating multi-scale conformational flexibility and scoring are strongly needed. Results: We describe here a new multi-scale protein-protein docking methodology, LightDock, capable of accommodating conformational flexibility and a variety of scoring functions at different resolution levels. Implicit use of normal modes during the search and atomic/coarse-grained combined scoring functions yielded improved predictive results with respect to state-of-the-art rigid-body docking, especially in flexible cases. Availability and implementation: The source code of the software and installation instructions are available for download at https://life.bsc.es/pid/lightdock/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Brian Jiménez-García, Jorge Roel-Touris, Miguel Romero-Durana, Miquel Vidal, Daniel Jiménez-González, Juan Fernández-Recio
Bioinform.5
2018 An approach to task-based parallel programming for undergraduate students
Eduard Ayguadé, Daniel Jiménez-González
J. Parallel Distributed Comput.2
2017 General Purpose Task-Dependence Management Hardware for Task-Based Dataflow Programming Models
abstract
Task-based programming models such as OpenMP, IntelTBB and OmpSs offer the possibility of expressing dependences among tasks to drive their execution at runtime. Managing these dependences introduces noticeable overheads when targeting fine-grained tasks, diminishing the potential speedups or even introducing performance losses. To overcome this drawback, we present a general purpose hardware accelerator, Picos++, to manage the inter-task dependences efficiently in both time and energy. Our design also includes a novel nested task support. To this end, a new hardware/software co-design is presented to overcome the fact that nested tasks with dependences could result in system deadlocks due to the limited amount of resources in hardware task dependence managers. In this paper we describe a detailed implementation of this design and evaluate a parallel task-based programming model using Picos++ in a Linux embedded system with two ARM Cortex-A9 and a FPGA. The scalability and energy consumption of the real system implemented have been studied and compared against a software runtime. Even in a system limited to 2 threads, using Picos++ results in more than 1.8x speedup and 40% of energy savings in the most demanding parallelizations of real benchmarks. As a matter of fact, a hardware task dependence manager should be able to achieve much higher speedup and provide more energy savings with more threads.
Xubin Tan, Jaume Bosch, Miquel Vidal, Carlos Álvarez 0001, Daniel Jiménez-González, Eduard Ayguadé, Mateo Valero
IPDPS5
2016 AXIOM: A Hardware-Software Platform for Cyber Physical Systems
abstract
Cyber-Physical Systems (CPSs) are widely necessary for many applications that require interactions with the humans and the physical environment. A CPS integrates a set of hardware-software components to distribute, execute and manage its operations. The AXIOM project (Agile, eXtensible, fast I/O Module) aims at developing a hardware-software platform for CPS such that i) it can use an easy parallel programming model and ii) it can easily scale-up the performance by adding multiple boards (e.g., 1 to 10 boards can run in parallel). AXIOM supports task-based programming model based on OmpSs and leverage a high-speed, inexpensive communication interface called AXIOM-Link. Another key aspect is that the board provides programmable logic (FPGA) to accelerate portions of an application. We are using smart video surveillance, and smart home living applications to drive our design.
Somnath Mazumdar, Eduard Ayguadé, Nicola Bettin, Javier Bueno, Sara Ermini, Antonio Filgueras, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Francesco Montefoschi, David Oro, Dionisios N. Pnevmatikatos, Antonio Rizzo, Dimitris Theodoropoulos 0001, Roberto Giorgi
DSD7
2016 Performance analysis of a hardware accelerator of dependence management for task-based dataflow programming models
abstract
Along with the popularity of multicore and manycore, task-based dataflow programming models obtain great attention for being able to extract high parallelism from applications without exposing the complexity to programmers. One of these pioneers is the OpenMP Superscalar (OmpSs). By implementing dynamic task dependence analysis, dataflow scheduling and out-of-order execution in runtime, OmpSs achieves high performance using coarse and medium granularity tasks. In theory, for the same application, the more parallel tasks can be exposed, the higher possible speedup can be achieved. Yet this factor is limited by task granularity, up to a point where the runtime overhead outweighs the performance increase and slows down the application. To overcome this handicap, Picos was proposed to support task-based dataflow programming models like OmpSs as a fast hardware accelerator for fine-grained task and dependence management, and a simulator was developed to perform design space exploration. This paper presents the very first functional hardware prototype inspired by Picos. An embedded system based on a Zynq 7000 All-Programmable SoC is developed to study its capabilities and possible bottlenecks. Initial scalability and hardware consumption studies of different Picos designs are performed to find the one with the highest performance and lowest hardware cost. A further thorough performance study is employed on both the prototype with the most balanced configuration and the OmpSs software-only alternative. Results show that our OmpSs runtime hardware support significantly outperforms the software-only implementation currently available in the runtime system for fine-grained tasks.
Xubin Tan, Jaume Bosch, Daniel Jiménez-González, Carlos Álvarez 0001, Eduard Ayguadé, Mateo Valero
ISPASS3
2016 MInGLE: An Efficient Framework for Domain Acceleration Using Low-Power Specialized Functional Units
abstract
The end of Dennard scaling leads to new research directions that try to cope with the utilization wall in modern chips, such as the design of specialized architectures. Processor customization utilizes transistors more efficiently, optimizing not only for performance but also for power. However, hardware specialization for each application is costly and impractical due to time-to-market constraints. Domain-specific specialization is an alternative that can increase hardware reutilization across applications that share similar computations. This article explores the specialization of low-power processors with custom instructions (CIs) that run on a specialized functional unit. We are the first, to our knowledge, to design CIs for an application domain and across basic blocks, selecting CIs that maximize both performance and energy efficiency improvements. We present the Merged Instructions Generator for Large Efficiency (MInGLE), an automated framework that identifies and selects CIs. Our framework analyzes large sequences of code (across basic blocks) to maximize acceleration potential while also performing partial matching across applications to optimize for reuse of the specialized hardware. To do this, we convert the code into a new canonical representation, the Merging Diagram, which represents the code’s functionality instead of its structure. This is key to being able to find similarities across such large code sequences from different applications with different coding styles. Groups of potential CIs are clustered depending on their similarity score to effectively reduce the search space. Additionally, we create new CIs that cover not only whole-body loops but also fragments of the code to optimize hardware reutilization further. For a set of 11 applications from the media domain, our framework generates CIs that significantly improve the energy-delay product (EDP) and performance speedup. CIs with the highest utilization opportunities achieve an average EDP improvement of 3.8 × compared to a baseline processor modeled after an Intel Atom. We demonstrate that we can efficiently accelerate a domain with partially matched CIs, and that their design time, from identification to selection, stays within tractable bounds.
Cecilia González-Alvarez, Jennifer B. Sartor, Carlos Álvarez 0001, Daniel Jiménez-González, Lieven Eeckhout
ACM Trans. Archit. Code Optim.4
2015 Automatic design of domain-specific instructions for low-power processors
abstract
This paper explores hardware specialization of low-power processors to improve performance and energy efficiency. Our main contribution is an automated framework that analyzes instruction sequences of applications within a domain at the loop body level and identifies exactly and partially-matching sequences across applications that can become custom instructions. Our framework transforms sequences to a new code abstraction, a Merging Diagram, that improves similarity identification, clusters alike groups of potential custom instructions to effectively reduce the search space, and selects merged custom instructions to efficiently exploit the available customizable area. For a set of 11 media applications, our fast framework generates instructions that significantly improve the energy-delay product and speed-up, achieving more than double the savings as compared to a technique analyzing sequences within basic blocks. This paper shows that partially-matched custom instructions, which do not significantly increase design time, are crucial to achieving higher energy efficiency at limited hardware areas.
Cecilia González-Alvarez, Jennifer B. Sartor, Carlos Álvarez 0001, Daniel Jiménez-González, Lieven Eeckhout
ASAP4
2015 The AXIOM Software Layers
abstract
People and objects will soon share the same digital network for information exchange in a world named as the age of the cyber-physical systems. The general expectation is that people and systems will interact in real-time. This poses pressure onto systems design to support increasing demands on computational power, while keeping a low power envelop. Additionally, modular scaling and easy programmability are also important to ensure these systems to become widespread. The whole set of expectations impose scientific and technological challenges that need to be properly addressed. The AXIOM project (Agile, eXtensible, fast I/O Module) will research new hardware/software architectures for cyber-physical systems to meet such expectations. The technical approach aims at solving fundamental problems to enable easy programmability of heterogeneous multi-core multi-board systems. AXIOM proposes the use of the task-based OmpSs programming model, leveraging low-level communication interfaces provided by the hardware. Modular scalability will be possible thanks to a fast interconnect embedded into each module. To this aim, an innovative ARM and FPGA-based board will be designed, with enhanced capabilities for interfacing with the physical world. Its effectiveness will be demonstrated with key scenarios such as Smart Video-Surveillance and Smart Living/Home (domotics).
Carlos Álvarez 0001, Eduard Ayguadé, Javier Bueno, Antonio Filgueras, Daniel Jiménez-González, Xavier Martorell, Nacho Navarro, Dimitris Theodoropoulos 0001, Dionisios N. Pnevmatikatos, Davide Catani, Claudio Scordino, Paolo Gai, Carlos Segura, Carles Fernández, David Oro, Javier Rodríguez Saeta, Pierluigi Passera, Alberto Pomella, Antonio Rizzo, Roberto Giorgi
DSD5
2015 Picos: A hardware runtime architecture support for OmpSs
Fahimeh Yazdanpanah, Carlos Álvarez 0001, Daniel Jiménez-González, Rosa M. Badia, Mateo Valero
Future Gener. Comput. Syst.3
2014 OmpSs@Zynq all-programmable SoC ecosystem
abstract
OmpSs is an OpenMP-like directive-based programming model that includes heterogeneous execution (MIC, GPU, SMP, etc.) and runtime task dependencies management. Indeed, OmpSs has largely influenced the recently appeared OpenMP 4.0 specification. Zynq All-Programmable SoC combines the features of a SMP and a FPGA and benefits DLP, ILP and TLP parallelisms in order to efficiently exploit the new technology improvements and chip resource capacities. In this paper, we focus on programmability and heterogeneous execution support, presenting a successful combination of the OmpSs programming model and the Zynq All-Programmable SoC platforms.
Antonio Filgueras, Eduard Gil, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Jan Langer, Juanjo Noguera, Kees A. Vissers
FPGA3
2014 Hybrid Dataflow/von-Neumann Architectures
abstract
General purpose hybrid dataflow/von-Neumann architectures are gaining attraction as effective parallel platforms. Although different implementations differ in the way they merge the conceptually different computational models, they all follow similar principles: they harness the parallelism and data synchronization inherent to the dataflow model, yet maintain the programmability of the von-Neumann model. In this paper, we classify hybrid dataflow/von-Neumann models according to two different taxonomies: one based on the execution model used for inter- and intrablock execution, and the other based on the integration level of both control and dataflow execution models. The paper reviews the basic concepts of von-Neumann and dataflow computing models, highlights their inherent advantages and limitations, and motivates the exploration of a synergistic hybrid computing model. Finally, we compare a representative set of recent general purpose hybrid dataflow/von-Neumann architectures, discuss their different approaches, and explore the evolution of these hybrid processors.
Fahimeh Yazdanpanah, Carlos Álvarez 0001, Daniel Jiménez-González, Yoav Etsion
IEEE Trans. Parallel Distributed Syst.3
2013 Heterogeneous tasking on SMP/FPGA SoCs: The case of OmpSs and the Zynq
abstract
OmpSs is a directive-based programming model that uses OpenMP-like directives, that allow to execute the tasks annotated on both the SMPs and as FPGA kernels on modern SoC processors, like the Xilinx Zynq platform. OmpSs includes the support for accelerators (MIC, GPUs, FPGAs) and task dependencies, like OpenMP 4.0 will support. In this paper we present our approach for the support of FPGAs and the Zynq SoC, the current status of the implementation, its analysis and performance evaluation.
Antonio Filgueras, Eduard Gil, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Jan Langer, Juanjo Noguera
VLSI-SoC4
2013 Accelerating an application domain with specialized functional units
abstract
Hardware specialization has received renewed interest recently as chips are hitting power limits. Chip designers of traditional processor architectures have primarily focused on general-purpose computing, partially due to time-to-market pressure and simpler design processes. But new power limits require some chip specialization. Although hardware configured for a specific application yields large speedups for low-power dissipation, its design is more complex and less reusable. We instead explore domain-based specialization, a scalable approach that balances hardware’s reusability and performance efficiency. We focus on specialization using customized compute units that accelerate particular operations. In this article, we develop automatic techniques to identify code sequences from different applications within a domain that can be targeted to a new custom instruction that will be run inside a configurable specialized functional unit (SFU). We demonstrate that using a canonical representation of computations finds more common code sequences among applications that can be mapped to the same custom instruction, leading to larger speedups while specializing a smaller core area than previous pattern-matching techniques. We also propose new heuristics to narrow the search space of domain-specific custom instructions, finding those that achieve the best performance across applications. We estimate the overall performance achieved with our automatic techniques using hardware models on a set of nine media benchmarks, showing that when limiting the core area devoted to specialization, the SFU customization with the largest speedups includes both application- and domain-specific custom instructions. We demonstrate that exploring domain-specific hardware acceleration is key to continued computing system performance improvements.
Cecilia González-Alvarez, Jennifer B. Sartor, Carlos Álvarez 0001, Daniel Jiménez-González, Lieven Eeckhout
ACM Trans. Archit. Code Optim.4
2012 Cell-Dock: high-performance protein-protein docking
abstract
SUMMARY: The application of docking to large-scale experiments or the explicit treatment of protein flexibility are part of the new challenges in structural bioinformatics that will require large computer resources and more efficient algorithms. Highly optimized fast Fourier transform (FFT) approaches are broadly used in docking programs but their optimal code implementation leaves hardware acceleration as the only option to significantly reduce the computational cost of these tools. In this work we present Cell-Dock, an FFT-based docking algorithm adapted to the Cell BE processor. We show that Cell-Dock runs faster than FTDock with maximum speedups of above 200×, while achieving results of similar quality. AVAILABILITY AND IMPLEMENTATION: The source code is released under GNU General Public License version 2 and can be downloaded from http://mmb.pcb.ub.es/~cpons/Cell-Dock. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Carles Pons, Daniel Jiménez-González, Cecilia González-Alvarez, Harald Servat, Daniel Cabrera-Benitez, Xavier Aguilar, Juan Fernández-Recio
Bioinform.2
2008 Drug Design Issues on the Cell BE
Harald Servat, Cecilia González-Alvarez, Xavier Aguilar, Daniel Cabrera-Benitez, Daniel Jiménez-González
HiPEAC5
2007 Drug Design on the Cell BroadBand Engine
Harald Servat, Cecilia Gonzalez, Xavier Aguilar, Daniel Cabrera, Daniel Jiménez-González
PACT5
2007 Performance Analysis of Cell Broadband Engine for High Memory Bandwidth Applications
abstract
The cell broadband engine (CBE) is designed to be a general purpose platform exposing an enormous arithmetic performance due to its eight SIMD-only synergistic processor elements (SPEs), capable of achieving 134.4 GFLOPS (16.8 GFLOPS * 8) at 2.1 GHz, and a 64-bit power processor element (PPE). Each SPE has a 256Kb non-coherent local memory, and communicates to other SPEs and main memory through its DMA controller. CBE main memory is connected to all the CBE processor elements (PPE and SPEs) through the element interconnect bus (EIB), which has a 134.4 GB/s bandwidth performance peak at half the processor speed. Therefore, CBE platform is suitable to be used by applications using MPI and streaming programming models with a potential high performance peak. In this paper we focus on the communication part of those applications, and measure the actual memory bandwidth that each of the CBE processor components can sustain. We have measured the sustained bandwidth between PPE and memory, SPE and memory, two individual SPEs to determine if this bandwidth depends on their physical location, pairs of SPEs to achieve maximum bandwidth in nearly-ideal conditions, and in a cycle of SPEs representing a streaming kind of computation. Our results on a real machine show that following some strict programming rules, individual SPE to SPE communication almost achieves the peak bandwidth when using the DMA controllers to transfer memory chunks of at least 1024 Bytes. In addition, SPE to memory bandwidth should be considered in streaming programming. For instance, implementing two data streams using 4 SPEs each can be more efficient than having a single data stream using the 8 SPEs
Daniel Jiménez-González, Xavier Martorell, Alex Ramírez
ISPASS1
2004 Characterization of the data access behavior for TPC-C traces
abstract
In this paper, we look into the characteristics of the reference stream of TPC-C workloads from the buffer pool point of view. We analyze a trace coming from DB2 UDB version 8.1 fix pack 4 and compare it to a trace from DB2 UDB version 8.1 GA. We perform three types of analysis. A static analysis of the number of reads and writes for index and data pages. We conclude that index pages receive less references than data pages by are more frequently accessed individually. Then, we analyze how DB2 processes access those pages. Index pages have more references than data pages when accessed by more than one process. Finally, we understand the accesses along the life of a page. We conclude that there is a significant burstiness in the reference stream, where, each burst is caused by one process.
R. Bonilla-Lucas, Peter Plachta, Aamer Sachedina, Daniel Jiménez-González, Calisto Zuzarte, Josep Lluís Larriba-Pey
ISPASS4
2001 Fast parallel in-memory 64-bit sorting
abstract
Parallel in-memory 64-bit sorting is an important problem in Database Management Systems and other applications such as Internet Search Engines and Data Mining Tools.
Daniel Jiménez-González, Juan J. Navarro, Josep Lluís Larriba-Pey
ICS1
1999 Communication conscious radix sort
abstract
The exploitation of data locality in parallel computers is paramount to reduce the memory traffic and communication among processing nodes. We focus on the exploitation of locality by Parallel Radix sort. The original Parallel Radix sort has several communication steps in which one sorting key may have to visit several processing nodes. In response to this, we propose a reorganization of Radix sort that leads to a highly local version of the algorithm at a very low cost. As a key feature, our algorithm performs one only communication step, forcing keys to move at most once between processing nodes. Also, our algorithm reduces the amount of data communicated. Finally, the new algorithm achieves a good load balance which makes it insensitive to skewed data distributions. We call the new version of Parallel Radix sort that combines locality and load balance, Communication and Cache Conscious Radix sort (C 3 -Radix sort). Our results on 16 processors of the SGI O2000 show that C 3 -R...
Daniel Jiménez-González, Josep Lluís Larriba-Pey, Juan J. Navarro
International Conference on Supercomputing1