EDBT 2026 Demo / reviewers in the wild / expert
José Luis Bosque
dblp:29/3095 · also Jose Luis Bosque Orero
· DBLP profile ↗
62ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-7718-8449ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 57 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Spike Transmission in Large-Scale Brain Network Simulations
Mario Ibáñez 0001, Marvin Kaster, Borja Pérez 0001, Han Lu 0001, Fabian Czappa, Sandra Díaz-Pier, José Luis Bosque, Felix Wolf 0001, Thorsten Hater |
IPDPS | 7 |
| 2026 | REX: A remote execution model for continuos scalability in multi-chiplet-module GPUsabstractMonolithic GPU architectures face growing limitations due to power density, yield issues, and manufacturing complexity, motivating a shift toward multi-chiplet designs. While promising, these architectures struggle with workloads exhibiting irregular memory access patterns, where static data placement is often insufficient. Though data locality can help, it does not adapt well to dynamic access behaviour, leading to performance degradation. This paper introduces REX, a runtime mechanism that migrates threads to the chiplet where their data resides, adapting dynamically to the generated memory access patterns with a fine granularity. By relocating computation instead of data, REX improves locality and minimises remote memory accesses, which are especially costly in multi-chiplet environments. As a result, it reduces inter-chiplet traffic and scales efficiently with the number of chiplets. On irregular workloads, the solution demonstrates consistent performance gains, averaging a 13 % speedup, with improvements reaching up to 38 %. Moreover, its scalability with chiplet count is particularly noteworthy, delivering a 25 % average gain, and peaking at an impressive 84 % in the most favourable scenarios. Mario Ibáñez 0001, Borja Pérez 0001, José Luis Bosque |
Future Gener. Comput. Syst. | 3 |
| 2026 | HRB: A backfilling algorithm for heterogeneous clusters with job prioritizationabstractBackfilling is a widely used scheduling technique in High-Performance Computing (HPC) systems to improve resource utilization. However, traditional approaches like EASY Backfill were devised for mono-core homogeneous environments, without considering the implications of multi-core architectures or the individual characteristics of nodes in heterogeneous clusters. This article proposes two refinements of EASY called Heterogeneous Backfill (HB) and Heterogeneous Reordering Backfill (HRB). These algorithms adapt the backfilling strategy to heterogeneous multi-core environments by incorporating node properties into the scheduling process. The HB algorithm sorts nodes based on a given criterion, such as power consumption or performance, to improve resource allocation. The HRB algorithm extends this approach by incorporating job reordering criteria, allowing for more efficient backfilling decisions. An evaluation of these algorithms shows that they can significantly reduce energy consumption and improve scheduling efficiency in heterogeneous clusters. The results demonstrate that the proposed algorithms outperform traditional backfilling methods, such as EASY Backfill, in terms of energy consumption, waiting time or makespan. By embracing the heterogeneity of modern HPC systems, these algorithms enable more efficient resource utilization and contribute to the overall performance of large-scale computing environments. Jaime Palacios, Esteban Stafford, José Luis Bosque |
Future Gener. Comput. Syst. | 3 |
| 2026 | Bicameral+ Cache: re-assessing split vector and scalar cache designs for increased efficiencyabstractAbstract Addressing the growing impact of the memory wall is critical to sustain performance in modern vector architectures. This work introduces the Bicameral+ Cache, an enhanced version of the Bicameral Cache architecture, which separates scalar and vector memory accesses into distinct cache structures, optimized for their respective locality patterns. Bicameral+ Cache incorporates two key improvements: a transition from a fully associative to a set-associative organization in the vector cache, reducing implementation complexity while preserving performance, and a novel replacement policy based on a configurable write-back threshold (WBT), which improves memory traffic efficiency. Experimental results show speedups of up to 1.59 $$\times $$ × in dense workloads and 1.63 $$\times $$ × in sparse ones, with respect to a conventional cache, when using a 16-way set-associative Bicameral+ Cache configuration. These findings, combined with estimations of a sevenfold area reduction and energy savings of one order of magnitude, confirm the practicality and effectiveness of the proposed enhancements for vector processing systems, retaining the benefits of the original Bicameral Cache design at reduced complexity and implementation costs. Aitor Echevarría, Susana Rebolledo Ruiz, Borja Pérez 0001, José Luis Bosque, Peter Hsu |
J. Supercomput. | 4 |
| 2025 | Intelligent energy pairing scheduler (InEPS) for heterogeneous HPC clustersabstractAbstract In recent years, energy consumption has become a limiting factor in the evolution of high-performance computing (HPC) clusters in terms of environmental concern and maintenance cost. The computing power of these clusters is increasing, together with the demands of the workloads they execute. A key component in HPC systems is the workload manager, whose operation has a substantial impact on the performance and energy consumption of the clusters. Recent research has employed machine learning techniques to optimise the operation of this component. However, these attempts have focused on homogeneous clusters where all the cores are pooled together and considered equal, disregarding the fact that they are contained in nodes and that they can have different performances. This work presents an intelligent job scheduler based on deep reinforcement learning that focuses on reducing energy consumption of heterogeneous HPC clusters. To this aim it leverages information provided by the users as well as the power consumption specifications of the compute resources of the cluster. The scheduler is evaluated against a set of heuristic algorithms showing that it has potential to give similar results, even in the face of the extra complexity of the heterogeneous cluster. Esteban Stafford, José Luis Bosque |
J. Supercomput. | 3 |
| 2025 | CPU-GPU co-execution through the exploitation of hybrid technologies via SYCLabstractAbstract The performance and energy efficiency offered by heterogeneous systems are highly useful for modern C++ applications, but the technological variety demands adequate portability and programmability. Initiatives such as Intel oneAPI facilitate the exploitation of Intel CPUs and GPUs, but not NVIDIA GPUs, which are present in systems of all kinds and are necessarily leveraged by CUDA technology. Frequently, only GPUs are used, leaving the CPU for management tasks, with the consequent loss of energy and system utilization. In this work, the CoexecutorRuntime system design and API are extended to transparently integrate backends of diverse technologies, unifying offloading mechanisms under a consistent co-execution API and scheduling runtime. Moreover, CPU-GPU co-execution of hybrid technologies is enabled to ensure performance portability. Experimental results show performance improvements for all programs studied, achieving average efficiencies of 0.91 and speedups of 1.31 over using only the GPU. Raúl Nozal, José Luis Bosque |
J. Supercomput. | 2 |
| 2024 | Hardware support for balanced co-execution in heterogeneous processorsabstractHeterogeneous systems are the go-to solution in computing, ranging from HPC to mobile, due to their excellent performance and energy efficiency. However, using this kind of systems adequately poses challenges. Namely, each of the devices that comprise the system are often considered as independent entities that need to be managed and dispatched work to manually. This represents a significant burden on programming and often results in a fastest-device-only approach, in which compute intensive regions are offloaded to the fastest device available, while the rest of the system idles. This idling represents a waste of computing capabilities that could be leveraged if the workload was co-executed. Software solutions have been proposed to provide transparent co-execution, but they always trade abstraction and ease of use for performance. In general, a higher level of abstraction, which improves programmability, will generate overheads. This paper presents HCoD (Hardware co-execution Dispatcher), a design for a hardware dispatcher to enable transparent co-execution without the overheads in integrated heterogeneous SoCs. The dispatcher distributes the work associated to a single kernel among CPU cores and GPU compute units at runtime, while monitoring co-execution to balance the load and prevent a slow device from delaying computation. HCoD achieves an excellent balance among all the compute elements and improves performance by an average of 14%, by transparently leveraging the computing capabilities already available in the hardware. Borja Pérez 0001, José Luis Bosque |
CF | 2 |
| 2024 | The Bicameral Cache: a split cache for vector architecturesabstractThe Bicameral Cache is a cache organization proposal for a vector architecture that segregates data according to their access type, distinguishing scalar from vector references. Its aim is to avoid both types of references from interfering in each other’s data locality, with a special focus on prioritizing the performance on vector references. The proposed system incorporates an additional, non-polluting prefetching mechanism to help populate the long vector cache lines in advance to increase the hit rate by further exploiting the spatial locality on vector data. Its evaluation was conducted on the Cavatools simulator, comparing the performance to a standard conventional cache, over different typical vector benchmarks for several vector lengths. The results proved the proposed cache speeds up performance on stride- 1 vector benchmarks, while hardly impacting non-stride-1’s. In addition, the prefetching feature consistently provided an additional value. Susana Rebolledo Ruiz, Borja Pérez 0001, José Luis Bosque, Peter Hsu |
ICPADS | 3 |
| 2024 | Enhancing heterogeneous cluster efficiency through node-centric schedulingabstractAbstract This article delves into the critical realm of modern computer cluster management. It focuses on the effect that the increasing heterogeneity of the clusters has on the workload managers. The proposed schedulers consider node properties instead of job properties to make decisions, which is something not currently done by mainstream scheduling algorithms. In order to increase the knowledge in this topic, this paper proposes two novel algorithms whose main task is to choose the best compute nodes to schedule the incoming jobs. To this effect, they exclusively take into account the properties of the nodes, instead of the common trend of considering the properties of the jobs. The experimental results show that these algorithms outperform well-known heuristic algorithms found in the literature. Esteban Stafford, José Luis Bosque |
J. Supercomput. | 2 |
| 2023 | Parallelisation of decision-making techniques in aquaculture enterprisesabstractAbstract Nowadays, theArtificial Intelligent (AI)techniques are applied in enterprise software to solveBig DataandBusiness Intelligence (BI)problems. But most AI techniques are computationally excessive, and they become unfeasible for common business use. Therefore, specific high performance computing is needed to reduce the response time and make these software applications viable on an industrial environment. The main objective of this paper is to demonstrate the improvement of an aquaculture BI tool based in AI techniques, using parallel programming. This tool, called AquiAID, was created by the research group of Economic Management for the Sustainable Development of Primary Sector of the Universidad de Cantabria. The parallelisation reduces the computation time up to 60 times, and the energy efficiency by 600 times with respect to the sequential program. With these improvements, the software will improve the fish farming management in aquaculture industry. Mario Ibáñez 0001, Manuel Luna, José Luis Bosque, Ramón Beivide |
J. Supercomput. | 3 |
| 2023 | Mashing load balancing algorithm to boost hybrid kernels in molecular dynamics simulationsabstractAbstract The path to the efficient exploitation of molecular dynamics simulators is strongly driven by the increasingly intensive use of accelerators. However, they suffer performance portability issues, making it necessary both to achieve technological combinations that allow taking advantage of each programming model and device, and to define more effective load distribution strategies that consider the simulation conditions. In this work, a new load balancing algorithm is presented, together with a set of optimizations to support hybrid co-execution in a runtime system for heterogeneous computing. The new extended design enables the exploitation of custom kernels and acceleration technologies altogether, being encapsulated for the rest of the runtime and its scheduling system. With this support, Mash algorithm allows to simultaneously leverage different workload distribution strategies, benefiting from the most advantageous one per device and technology. Experiments show that these proposals achieve an efficiency close to 0.90 and an energy efficiency improvement around 1.80 over the original optimized version. Raúl Nozal, José Luis Bosque |
J. Supercomput. | 2 |
| 2021 | A Simulator for Intelligent Workload Managers in Heterogeneous ClustersabstractModern High Performance Computing (HPC) clusters often comprise a huge amount of computing resources of different capabilities, making them heterogeneous and difficult to manage. In addition, they must deal with a wide range of applications with different requirements. All this poses a great challenge to the workload managers that assign applications to resources. There are many new proposals to overcome this challenge, including some that employ Deep Reinforcement Learning (DRL) techniques. This paper proposes a novel simulation framework for the study of workload managers, that has been conceived to foster the study of workload managers based on DRL techniques. Its main features include the simulation of heterogeneous clusters based on multicore architectures, taking into account the contention in shared memory access and the energy consumption. A validation of the accuracy and performance of the simulator was made, compared with a real environment based on Slurm. This shows good accuracy of the results, with a relative error below 5% in makespan and 10% in energy consumption, and speedups up to 200. Adrián Herrera, Mario Ibáñez 0001, Esteban Stafford, José Luis Bosque |
CCGRID | 4 |
| 2021 | Exploiting Co-execution with OneAPI: Heterogeneity from a Modern Perspective
Raúl Nozal, José Luis Bosque |
Euro-Par | 2 |
| 2021 | Sigmoid: An auto-tuned load balancing algorithm for heterogeneous systemsabstractA challenge that heterogeneous system programmers face is leveraging the performance of all the devices that integrate the system. This paper presents Sigmoid, a new load balancing algorithm that efficiently co-executes a single OpenCL data-parallel kernel on all the devices of heterogeneous systems. Sigmoid splits the workload proportionally to the capabilities of the devices, drastically reducing response time and energy consumption. It is designed around several features; it is dynamic, adaptive, guided and effortless, as it does not require the user to give any parameter, adapting to the behaviour of each kernel at runtime. To evaluate Sigmoid's performance, it has been implemented in Maat, a system abstraction library. Experimental results with different kernel types show that Sigmoid exhibits excellent performance, reaching a utilization of 90%, together with energy savings up to 20%, always reducing programming effort compared to OpenCL, and facilitating the portability to other heterogeneous machines. Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide |
J. Parallel Distributed Comput. | 3 |
| 2021 | Performance and energy task migration model for heterogeneous clusters
Esteban Stafford, José Luis Bosque |
J. Supercomput. | 2 |
| 2020 | EngineCL: Usability and Performance in Heterogeneous Computing
Raúl Nozal, José Luis Bosque, Ramón Beivide |
Future Gener. Comput. Syst. | 2 |
| 2020 | Improving utilization of heterogeneous clusters
Esteban Stafford, José Luis Bosque |
J. Supercomput. | 2 |
| 2019 | Simulation with skeletons of applications using dimemasabstractLarge computer systems, like those in the TOP 500 ranking, comprise about hundreds of thousands cores. Simulating application execution in these systems is very complex and costly. This article explores the option of using application skeletons, together with an analytic simulator, to study the performance of these large systems. With this aim, the Dimemas simulator has been enhanced with the capability of simulating application skeletons. This enhancement allows simulating the skeleton of Lulesh, an application with 90k processes in a single day. In addition, it also generates traces, which is of great value to validate skeletons and simulations. Cristobal Camarero, Carmen Martínez 0001, José Luis Bosque |
CF | 3 |
| 2019 | Auto-tuned OpenCL kernel co-execution in OmpSs for heterogeneous systems
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé |
J. Parallel Distributed Comput. | 3 |
| 2019 | Cooperative CPU, GPU, and FPGA heterogeneous execution with EngineCL
Maria Angelica Davila Guzman, Raúl Nozal, Ruben Gran Tejero, María Villarroya-Gaudó, Darío Suárez Gracia, José Luis Bosque |
J. Supercomput. | 6 |
| 2019 | Load balancing in a heterogeneous world: CPU-Xeon Phi co-execution of data-parallel kernels
Raúl Nozal, Borja Pérez 0001, José Luis Bosque, Ramón Beivide |
J. Supercomput. | 3 |
| 2018 | Architectural Support for Task Dependence Management with Flexible Software SchedulingabstractThe growing complexity of multi-core architectures has motivated a wide range of software mechanisms to improve the orchestration of parallel executions. Task parallelism has become a very attractive approach thanks to its programmability, portability and potential for optimizations. However, with the expected increase in core counts, finer-grained tasking will be required to exploit the available parallelism, which will increase the overheads introduced by the runtime system. This work presents Task Dependence Manager (TDM), a hardware/software co-designed mechanism to mitigate runtime system overheads. TDM introduces a hardware unit, denoted Dependence Management Unit (DMU), and minimal ISA extensions that allow the runtime system to offload costly dependence tracking operations to the DMU and to still perform task scheduling in software. With lower hardware cost, TDM outperforms hardware-based solutions and enhances the flexibility, adaptability and composability of the system. Results show that TDM improves performance by 12.3% and reduces EDP by 20.4% on average with respect to a software runtime system. Compared to a runtime system fully implemented in hardware, TDM achieves an average speedup of 4.2% with 7.3x less area requirements and significant EDP reductions. In addition, five different software schedulers are evaluated with TDM, illustrating its flexibility and performance gains. Emilio Castillo, Lluc Alvarez, Miquel Moretó, Marc Casas, Enrique Vallejo 0001, José Luis Bosque, Ramón Beivide, Mateo Valero |
HPCA | 6 |
| 2017 | To Distribute or Not to Distribute: The Question of Load Balancing for Performance or Energy
Esteban Stafford, Borja Pérez 0001, José Luis Bosque, Ramón Beivide, Mateo Valero |
Euro-Par | 3 |
| 2017 | Extending OmpSs for OpenCL Kernel Co-Execution in Heterogeneous SystemsabstractHeterogeneous systems have a very high potential performance but present difficulties in their programming. OmpSs is a well known framework for task based parallel applications, which is an interesting tool to simplify the programming of these systems. However, it does not support the co-execution of a single OpenCL kernel instance on several compute devices. To overcome this limitation, this paper presents an extension of the OmpSs framework that solves two main objectives: the automatic division of datasets among several devices and the management of their memory address spaces. To adapt to different kinds of applications, the data division can be performed by the novel HGuided load balancing algorithm or by the well known Static and Dynamic. All this is accomplished with negligible impact on the programming. Experimental results reveal that there is always one load balancing algorithm that improves the performance and energy consumption of the system. Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé |
SBAC-PAD | 3 |
| 2017 | A scalable synthetic traffic model of Graph500 for computer networks analysisabstractSummary The Graph500 benchmark attempts to steer the design of High‐Performance Computing systems to maximize the performance under memory‐constricted application workloads. A realistic simulation of such benchmarks for architectural research is challenging due to size and detail limitations. By contrast, synthetic traffic workloads constitute one of the least resource‐consuming methods to evaluate the performance. In this work, we provide a simulation tool for network architects that need to evaluate the suitability of their interconnect for BigData applications. Our development is a low computation‐ and memory‐demanding synthetic traffic model that emulates the behavior of the Graph500 communications and is publicly available in an open‐source network simulator. The characterization of network traffic is inferred from a profile of several executions of the benchmark with different input parameters. We verify the validity of the equations in our model against an execution of the benchmark with a different set of parameters. Furthermore, we identify the impact of the node computation capabilities and network characteristics in the execution time of the model in a Dragonfly network. Pablo Fuentes 0001, Mariano Benito, Enrique Vallejo 0001, José Luis Bosque, Ramón Beivide, Andreea Anghel, Mitchell Gusat, Cyriel Minkenberg, Mateo Valero |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | A clustering-based knowledge discovery process for data centre infrastructure management
Diego García-Saiz, Marta E. Zorrilla, José Luis Bosque |
J. Supercomput. | 3 |
| 2017 | Energy efficiency of load balancing for data-parallel applications in heterogeneous systems
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide |
J. Supercomput. | 3 |
| 2016 | Synthetic Traffic Model of the Graph500 Communications
Pablo Fuentes 0001, Enrique Vallejo 0001, José Luis Bosque, Ramón Beivide, Andreea Anghel, Mitchell Gusat, Cyriel Minkenberg |
ICA3PP | 3 |
| 2016 | CATA: Criticality Aware Task Acceleration for Multicore ProcessorsabstractManaging criticality in task-based programming models opens a wide range of performance and power optimization opportunities in future manycore systems. Criticality aware task schedulers can benefit from these opportunities by scheduling tasks to the most appropriate cores. However, these schedulers may suffer from priority inversion and static binding problems that limit their expected improvements. Based on the observation that task criticality information can be exploited to drive hardware reconfigurations, we propose a Criticality Aware Task Acceleration (CATA) mechanism that dynamically adapts the computational power of a task depending on its criticality. As a result, CATA achieves significant improvements over a baseline static scheduler, reaching average improvements up to 18.4% in execution time and 30.1% in Energy-Delay Product (EDP) on a simulated 32-core system. The cost of reconfiguring hardware by means of a software-only solution rises with the number of cores due to lock contention and reconfiguration overhead. Therefore, novel architectural support is proposed to eliminate these overheads on future manycore systems. This architectural support minimally extends hardware structures already present in current processors, which allows further improvements in performance with negligible overhead. As a consequence, average improvements of up to 20.4% in execution time and 34.0% in EDP are obtained, outperforming state-of-the-art acceleration proposals not aware of task criticality. Emilio Castillo, Miquel Moretó, Marc Casas, Lluc Alvarez, Enrique Vallejo 0001, Kallia Chronaki, Rosa M. Badia, José Luis Bosque, Ramón Beivide, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
IPDPS | 8 |
| 2016 | Assessing the Suitability of King Topologies for Interconnection NetworksabstractIn the late years many different interconnection networks have been used with two main tendencies. One is characterized by the use of high-degree routers with long wires while the other uses routers of much smaller degree. The latter rely on two-dimensional mesh and torus topologies with shorter local links. This paper focuses on doubling the degree of common 2D meshes and tori while still preserving an attractive layout for VLSI design. By adding a set of diagonal links in one direction, diagonal networks are obtained. By adding a second set of links, networks of degree eight are built, named king networks. This research presents a comprehensive study of these networks which includes a topological analysis, the proposal of appropriate routing procedures and an empirical evaluation. King networks exhibit a number of attractive characteristics which translate to reduced execution times of parallel applications. For example, the execution times NPB suite are reduced up to a 30 percent. In addition, this work reveals other properties of king networks such as perfect partitioning that deserves further attention for its convenient exploitation in forthcoming high-performance parallel systems. Esteban Stafford, José Luis Bosque, Carmen Martínez 0001, Fernando Vallejo, Ramón Beivide, Cristobal Camarero, Emilio Castillo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Financial applications on multi-CPU and multi-GPU architectures
Emilio Castillo, Cristobal Camarero, Ana Borrego, José Luis Bosque |
J. Supercomput. | 4 |
| 2014 | Leveraging OmpSs to Exploit Hardware AcceleratorsabstractCUDA and OpenCL are the most widely used programming models to exploit hardware accelerators. Both programming models provide a C-based programming language to write accelerator kernels and a host API used to glue the host and kernel parts. Although this model is a clear improvement over a low-level and ad-hoc programming model for each hardware accelerator, it is still too complex and cumbersome for general adoption. For large and complex applications using several accelerators, the main problem becomes the explicit coordination and management of resources required between the host and the hardware accelerators that introduce a new family of issues (scheduling, data transfers, synchronization, ) that the programmer must take into account. In this paper, we propose a simple extension to OmpSs -- a data-flow programming model -- that dramatically simplifies the integration of accelerated code, in the form of CUDA or OpenCL kernels, into any C, C++ or Fortran application. Our proposal fully replaces the CUDA and OpenCL host APIs with a few pragmas, so we can leverage any kernel written in CUDA C or OpenCL C without any performance impact. Our compiler generates all the boilerplat code while our runtime system takes care of kernels scheduling, data transfers between host and accelerators and synchronizations between host and kernels parts. To evaluate our approach, we have ported several native CUDA and OpenCL applications to OmpSs by replacing all the CUDA or OpenCL API calls by a few number of pragmas. The OmpSs versions of these applications have competitive performance and scalability but with a significantly lower complexity than the original ones. Florentino Sainz, Sergi Mateo, Vicenç Beltran 0001, José Luis Bosque, Xavier Martorell, Eduard Ayguadé |
SBAC-PAD | 4 |
| 2013 | Advanced Switching Mechanisms for Forthcoming On-Chip NetworksabstractMany current VLSI on-chip multiprocessors and systems-on-chip employ point-to-point switched interconnection networks. Rings and 2D-meshes are among the most popular interconnection topologies for these increasingly important onchip networks. Nevertheless, rings cannot scale beyond dozens of nodes and meshes are asymmetric. Two of the key features of square 2D-tori are their scalability and symmetry. As higher scalability is demanded by the increasing number of cores (or specialized units) integrated on a chip and symmetry is critical for high-performance and load balancing, we concentrate on 2D-tori. However, most popular deadlock-free routing mechanisms are based on Dimension Order Routing (DOR) which breaks the torus symmetry when managing adversarial traffic patterns. This paper analyzes this problem and its consequences. After that, it proposes a new deadlock-free fully adaptive minimal routing, denoted as σDOR, that preserves tori symmetry under any load. It uses just two virtual channels to avoid DOR-induced asymmetry, the same as in previous competitive proposals. σDOR exhibits better behavior than any of previous solutions as it allows packets to dynamically adapt to local congestion. Experimental results evidence the superior performance of our mechanism, confirming the negative impact of DOR asymmetry. Emilio Castillo, Cristobal Camarero, Esteban Stafford, Fernando Vallejo, José Luis Bosque, Ramón Beivide |
DSD | 5 |
| 2013 | Analyzing scalability of parallel systems with unbalanced workload
José Luis Bosque, Oscar David Robles, Pablo Toharia, Luis Pastor |
J. Supercomput. | 1 |
| 2013 | A load index and load balancing algorithm for heterogeneous clusters
José Luis Bosque, Pablo Toharia, Oscar David Robles, Luis Pastor |
J. Supercomput. | 1 |
| 2013 | Scalable shot boundary detection
Pablo Toharia, Oscar David Robles, José Luis Bosque, Angel Rodríguez |
J. Supercomput. | 3 |
| 2012 | Static Multi-device Load Balancing for OpenCLabstractThis paper presents the Load Balancing for OpenCL (lbcl) library, devoted to automatically solve load balancing issues on both multi-platform and heterogeneous environments. Using this library, a single kernel can be executed on a set of heterogeneous devices, giving each device an amount of work proportional to its computing power. A wrapper has been developed so the library can balance the workload of an existing application not only without introducing any changes into its source code, but without any recompilation stage. Also a general OpenCL profiler has been developed to easily do a detailed profiling of the obtained results. Carlos S. de La Lama, Pablo Toharia, José Luis Bosque, Oscar David Robles |
ISPA | 3 |
| 2012 | Shot boundary detection using Zernike moments in multi-GPU multi-CPU architectures
Pablo Toharia, Oscar David Robles, Ricardo Suárez, José Luis Bosque, Luis Pastor |
J. Parallel Distributed Comput. | 4 |
| 2011 | Evaluating scalability in heterogeneous systems
José Luis Bosque, Oscar David Robles, Pablo Toharia, Luis Pastor |
J. Supercomput. | 1 |
| 2011 | Genetic Algorithm for Boolean minimization in an FPGA cluster
César Pedraza, Javier Castillo, José Ignacio Martínez, Pablo Huerta, José Luis Bosque, Javier Cano-Montero |
J. Supercomput. | 5 |
| 2010 | A First Approach to King Topologies for On-Chip Networks
Esteban Stafford, José Luis Bosque, Carmen Martínez 0001, Fernando Vallejo, Ramón Beivide, Cristobal Camarero |
Euro-Par (2) | 2 |
| 2010 | GCViR: grid content-based video retrieval with work allocation brokeringabstractAbstract Nowadays TV channels generate a large amount of video data each day. A very huge number of videos of news, shows, series, movies, and so on have to be stored with the aim of being accessed later on. Moreover, channels have a clear need of sharing videos so as to settle a real collaboration among them that minimizes the cost of information acquisition. These features demand a huge storage capacity and a sharing information environment. Both requirements can be solved by using grid computing. It provides both computing and storage capacities to store that great volume of data required, as well as the resource sharing capabilities for the cooperation of different TV channels. This paper presents a video retrieval system that covers these needs and suggests a work allocation (WA) broker to improve the performance of video accesses. An evaluation shows the feasibility and the scalability of this approach showing the benefits of the WA made to store and retrieve large video data. Copyright © 2009 John Wiley & Sons, Ltd. Pablo Toharia, Alberto Sánchez 0001, José Luis Bosque, Oscar David Robles |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | Study of neural net training methods in parallel and distributed architectures
Rafael Menéndez de Llano, José Luis Bosque |
Future Gener. Comput. Syst. | 2 |
| 2010 | A semantic collaborative awareness model to deal with resource sharing in grids
Manuel Salvadores, Pilar Herrero, José Luis Bosque, María S. Pérez 0001 |
Future Gener. Comput. Syst. | 3 |
| 2010 | Content-based image retrieval algorithm acceleration in a low-cost reconfigurable FPGA cluster
César Pedraza, Emilio Castillo, Javier Castillo, José Luis Bosque, José Ignacio Martínez, Oscar David Robles, Javier Cano-Montero, Pablo Huerta |
J. Syst. Archit. | 4 |
| 2009 | Hardware accelerated montecarlo financial simulation over low cost FPGA clusterabstractThe use of computational systems to help making the right investment decisions in financial markets is an open research field where multiple efforts have being carried out during the last few years. The ability of improving the assessment process and being faster than the rest of the players is one of the keys for the success on this competitive scenario. This paper explores different options to accelerate the computation of the option pricing problem (supercomputer, FPGA cluster or GPU) using the Montecarlo method to solve the Black-Scholes formula, and presents a quantitative study of their performance and scalability. Javier Castillo, José Luis Bosque, Emilio Castillo, Pablo Huerta, José Ignacio Martínez |
IPDPS | 2 |
| 2008 | SMILE: Scientific Parallel Multiprocessing based on Low-Cost Reconfigurable HardwareabstractThe SMILE project attempts to build efficient lowcost clusters based on FPGA boards using their reconfigurability capabilities. A real parallel application of Content-Based Information Retrieval over the SMILE cluster is presented. Using this application the SMILE cluster’s performance is evaluated and compared in terms of time and power consumption with traditional cluster architecture. Emilio Castillo, César Pedraza, Javier Castillo, Cristobal Camarero, José Luis Bosque, Rafael Menéndez de Llano, José Ignacio Martínez |
FCCM | 5 |
| 2008 | Cluster architecture based on low cost reconfigurable hardwareabstractThe SMILE project accelerates scientific and industrial applications by means of a cluster of low-cost FPGA boards. With this approach the intensive calculation tasks are accelerated using the FPGA logic, while the communication patterns of the applications remains unchanged by using a Message Passing Library over Linux. This paper explains the cluster architecture: the SMILE nodes and the developed high-speed communication network for the FPGA RocketIO interfaces. A SystemC model developed to simulate the cluster is also detailed. In order to show the potential of the SMILE proposal a Content-Based Information Retrieval parallel application has been developed and compared with a HP cluster architecture in terms of response time andpower consumption. César Pedraza, Emilio Castillo, Javier Castillo, Cristobal Camarero, José Luis Bosque, José Ignacio Martínez, Rafael Menéndez de Llano |
FPL | 5 |
| 2008 | A New CPU Availability Prediction Model for Time-Shared SystemsabstractThe success of different computing models, performance analysis and load balancing and algorithms depends on the processor availability information because there is a strong relationship between a process response time and the processor time available for its execution. Therefore, predicting the processor availability for a new process or task in a computer system is a basic problem that arises in in many important contexts. Unfortunately, making such predictions is not easy because of the dynamic nature of current computer systems and their workload, which can vary drastically in a short interval of time. This paper presents two new availability prediction models. The first, called SPAP (Static Process Assignment Prediction) model, is capable of predicting the CPU availability for a new task on a computer system having information about the tasks in its run queue. The second, called DYPAP (DYnamic Process Assignment Prediction) model, is an improvement of the SPAP model capable of making these predictions from real-time measurements provided by a monitoring tool, without any kind of information about the tasks in the run queue. Furthermore, the implementation of this monitoring tool for Linux workstations is presented. Marta Beltrán, Antonio Guzmán, José Luis Bosque |
IEEE Trans. Computers | 3 |
| 2007 | A Collaborative-Aware Task Balancing Delivery Model for Clusters
José Luis Bosque, Pilar Herrero, Manuel Salvadores, María S. Pérez 0001 |
GPC | 1 |
| 2006 | Parallel implementation of evolutionary strategies on heterogeneous clusters with load balancingabstractThis paper presents a load balancing algorithm for a parallel implementation of an evolutionary strategy on heterogeneous clusters. Evolutionary strategies can efficiently solve a diverse set of optimization problems. Due to cluster heterogeneity and in order to improve the speedup of the parallel implementation a load balancing algorithm has been implemented. This load balancing algorithm takes into account cluster heterogeneity and it is based on an optimal initial distribution. This initial distribution is determined based on the cluster nodes' computational powers that are dynamically measured in each slave node by an ad hoc load-benchmark. The implementation presents very satisfactory parallelization results, both in performance and scalability and super-linear speedup is reached for several tests configurations. Experimental results show excellent performance, increasing the improvements with the load balancing algorithm Juan Francisco Garamendi, José Luis Bosque |
IPDPS | 2 |
| 2006 | Video Shot Extraction on Parallel Architectures
Pablo Toharia, Oscar David Robles, José Luis Bosque, Angel Rodríguez |
ISPA | 3 |
| 2006 | Dealing with Heterogeneity in Load Balancing AlgorithmsabstractCluster heterogeneity increases the difficulty of balancing the load across the system nodes. Although the relationship between heterogeneity and load balancing is difficult to describe analytically, in this paper an exhaustive analysis of the effects of this system feature on load balancing algorithms performance is presented. Considering the performed analysis, there are two main challenges that need to be faced when dealing with cluster heterogeneity in load balancing algorithms: one related to the state measurement stage and another to the initiation rule. In this paper techniques to deal with heterogeneity in these two algorithm stages are proposed. Furthermore, suggestions to improve the performance of the rest of algorithm stages in heterogeneous environments are made too Marta Beltrán, Antonio Guzmán, José Luis Bosque |
ISPDC | 3 |
| 2006 | On Board: Sharing Resources in a Collaborative Grid-TV EnvironmentabstractThe TV environment is not so different from any other grid computing scenario, as different TV channels can need to share multimedia information and resources to achieve some specific purposes on time. In this paper, we present a web service specification to manage awareness in collaborative grid environments, WS-AMBLE. This specification has been designed by merging web services with multi-agents theories and principles to provide an autonomous, efficient and independent management of the amount of resources available in TV environment. WSAMBLE implementation makes easier the collaboration and cooperation among different TV channels. In this paper, we also present some experimental results carried out over a real and heterogeneous grid environment with the end of emphasizing the performance improvements that WS-AMBLE has in a Grid-TV environment. Pilar Herrero, José Luis Bosque, Manuel Salvadores, María S. Pérez 0001 |
Web Intelligence | 2 |
| 2006 | Parallel CBIR implementations with load balancing algorithms
José Luis Bosque, Oscar David Robles, Luis Pastor, Angel Rodríguez |
J. Parallel Distributed Comput. | 1 |
| 2006 | A Parallel Computational Model for Heterogeneous ClustersabstractThis paper addresses the maximal lifetime scheduling for sensor surveillance systems with K sensors to 1 target. Given a set of sensors and targets in an Euclidean plane, a sensor can watch only one target at a time and a target should be watched by k, kges1, sensors at any time. Our task is to schedule sensors to watch targets and pass data to the base station, such that the lifetime of the surveillance system is maximized, where the lifetime is the duration up to the time when there exists one target that cannot be watched by k sensors or data cannot be forwarded to the base station due to the depletion of energy of the sensor nodes. We propose an optimal solution to find the target watching schedule for sensors that achieves the maximal lifetime. Our solution consists of three steps: 1) computing the maximal lifetime of the surveillance system and a workload matrix by using linear programming techniques, 2) decomposing the workload matrix into a sequence of schedule matrices that can achieve the maximal lifetime, and 3) determining the sensor surveillance trees based on the above obtained schedule matrices, which specify the active sensors and the routes to pass sensed data to the base station. This is the first time in the literature that this scheduling problem of sensor surveillance systems has been formulated and the optimal solution has been found. We illustrate our optimal method by a numeric example and experiments in the end José Luis Bosque, Luis Pastor |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | Information policies for load balancing on heterogeneous systemsabstractDynamic load balancing decisions on cluster or grid computing systems are always based on some kind of knowledge about the system nodes state. Therefore, to balance the computational workload among the different system nodes, a local or global information exchange is needed. The information policy determines how and when this exchange is performed. This paper discusses and compares the three traditional information policies: periodic, on demand and event-driven, in order to conclude if there is a better approach for the current heterogeneous systems. All the comparisons are performed evaluating a new metric, called information efficiency, defined in this paper to quantify the benefits obtained with an information policy considering the network traffic generated by the information exchange and its effectiveness in the information updating. Marta Beltrán, José Luis Bosque |
CCGRID | 2 |
| 2005 | Initiating Load Balancing Operations
Marta Beltrán, José Luis Bosque, Antonio Guzmán |
Euro-Par | 2 |
| 2004 | Theoretical scalability analysis for heterogeneous clustersabstractScalability is one key concept for analyzing the performance of parallel systems, specially in clusters and Grid environments. Even though heterogeneous systems are becoming more common, there is no suitable definition of this concept for studying their behaviour. This paper presents an extension of the isoefficiency method that permits its application to heterogeneous environments as well. The heterogeneous isoefficiency function can predict the system scalability independently of the algorithm nature. This method is based only on the algorithm's overhead time and an adequate workload distribution. This paper also presents a set of theoretical scalability analyses, and finally some experiments where the method's usefulness and validity are shown. José Luis Bosque, L. P. Perez |
CCGRID | 1 |
| 2004 | HLogGP: a new parallel computational model for heterogeneous clustersabstractHeterogeneous clusters claim for new models and algorithms. In this paper a new parallel computational model is presented. The model, based on the LogGP model, has been extended to be able to deal with heterogeneous parallel systems. For that purpose, the LogGP's scalar parameters have been replaced by vector and matrix parameters to take into account the different node's features. The work presented here includes the parameterization of a real cluster which illustrates the impact of node heterogeneity over the model's parameters. Finally, the paper presents some experiments performed in a real heterogeneous cluster that can be used for assessing the method's validity, together with the main conclusions and future work. José Luis Bosque, L. P. Perez |
CCGRID | 1 |
| 2002 | Brain Activity Detection in Functional Magnetic Resonance Imaging on Heterogeneous ClusterabstractIn this paper we show a cluster-based solution for analyzing magnetic resonance imaging from the brain in order to obtain information about the brain activity in conscious and awake subjects. The huge amount of data to be analyzed makes sense the use of computer clusters. Furthermore, one heterogeneous cluster has been chosen because is a more realistic net environment in the most of the medical institutions. The complete method in the future would be a new approach to real time detection of brain activity. Juan Antonio Hernández Tamames, José Luis Bosque, J. Canto |
CCGRID | 2 |
| 2001 | An efficiency and scalability model for heterogeneous clustersabstractEfficiency and scalability are two key concepts for analyzing the performance of parallel systems. Even though heterogeneous systems are becoming more common, there are no suitable definitions of similar concepts for studying their behaviour. This paper first presents a definition of efficiency that can be applied equally to heterogeneous and homogeneous systems. From that definition, the paper also presents an extension of the isoefficiency method that permits its application to heterogeneous environments as well. Finally the paper presents some experiments where the methods' usefulness and validity are shown. Luis Pastor, José Luis Bosque |
CLUSTER | 2 |