Borja Pérez 0001

dblp:160/2194 · also Borja Pérez Pavón · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-3695-2906ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Scalable Spike Transmission in Large-Scale Brain Network Simulations
Mario Ibáñez 0001, Marvin Kaster, Borja Pérez 0001, Han Lu 0001, Fabian Czappa, Sandra Díaz-Pier, José Luis Bosque, Felix Wolf 0001, Thorsten Hater
IPDPS3
2026 REX: A remote execution model for continuos scalability in multi-chiplet-module GPUs
abstract
Monolithic GPU architectures face growing limitations due to power density, yield issues, and manufacturing complexity, motivating a shift toward multi-chiplet designs. While promising, these architectures struggle with workloads exhibiting irregular memory access patterns, where static data placement is often insufficient. Though data locality can help, it does not adapt well to dynamic access behaviour, leading to performance degradation. This paper introduces REX, a runtime mechanism that migrates threads to the chiplet where their data resides, adapting dynamically to the generated memory access patterns with a fine granularity. By relocating computation instead of data, REX improves locality and minimises remote memory accesses, which are especially costly in multi-chiplet environments. As a result, it reduces inter-chiplet traffic and scales efficiently with the number of chiplets. On irregular workloads, the solution demonstrates consistent performance gains, averaging a 13 % speedup, with improvements reaching up to 38 %. Moreover, its scalability with chiplet count is particularly noteworthy, delivering a 25 % average gain, and peaking at an impressive 84 % in the most favourable scenarios.
Mario Ibáñez 0001, Borja Pérez 0001, José Luis Bosque
Future Gener. Comput. Syst.2
2026 Bicameral+ Cache: re-assessing split vector and scalar cache designs for increased efficiency
abstract
Abstract Addressing the growing impact of the memory wall is critical to sustain performance in modern vector architectures. This work introduces the Bicameral+ Cache, an enhanced version of the Bicameral Cache architecture, which separates scalar and vector memory accesses into distinct cache structures, optimized for their respective locality patterns. Bicameral+ Cache incorporates two key improvements: a transition from a fully associative to a set-associative organization in the vector cache, reducing implementation complexity while preserving performance, and a novel replacement policy based on a configurable write-back threshold (WBT), which improves memory traffic efficiency. Experimental results show speedups of up to 1.59 $$\times $$ × in dense workloads and 1.63 $$\times $$ × in sparse ones, with respect to a conventional cache, when using a 16-way set-associative Bicameral+ Cache configuration. These findings, combined with estimations of a sevenfold area reduction and energy savings of one order of magnitude, confirm the practicality and effectiveness of the proposed enhancements for vector processing systems, retaining the benefits of the original Bicameral Cache design at reduced complexity and implementation costs.
Aitor Echevarría, Susana Rebolledo Ruiz, Borja Pérez 0001, José Luis Bosque, Peter Hsu
J. Supercomput.3
2024 Hardware support for balanced co-execution in heterogeneous processors
abstract
Heterogeneous systems are the go-to solution in computing, ranging from HPC to mobile, due to their excellent performance and energy efficiency. However, using this kind of systems adequately poses challenges. Namely, each of the devices that comprise the system are often considered as independent entities that need to be managed and dispatched work to manually. This represents a significant burden on programming and often results in a fastest-device-only approach, in which compute intensive regions are offloaded to the fastest device available, while the rest of the system idles. This idling represents a waste of computing capabilities that could be leveraged if the workload was co-executed. Software solutions have been proposed to provide transparent co-execution, but they always trade abstraction and ease of use for performance. In general, a higher level of abstraction, which improves programmability, will generate overheads. This paper presents HCoD (Hardware co-execution Dispatcher), a design for a hardware dispatcher to enable transparent co-execution without the overheads in integrated heterogeneous SoCs. The dispatcher distributes the work associated to a single kernel among CPU cores and GPU compute units at runtime, while monitoring co-execution to balance the load and prevent a slow device from delaying computation. HCoD achieves an excellent balance among all the compute elements and improves performance by an average of 14%, by transparently leveraging the computing capabilities already available in the hardware.
Borja Pérez 0001, José Luis Bosque
CF1
2024 The Bicameral Cache: a split cache for vector architectures
abstract
The Bicameral Cache is a cache organization proposal for a vector architecture that segregates data according to their access type, distinguishing scalar from vector references. Its aim is to avoid both types of references from interfering in each other’s data locality, with a special focus on prioritizing the performance on vector references. The proposed system incorporates an additional, non-polluting prefetching mechanism to help populate the long vector cache lines in advance to increase the hit rate by further exploiting the spatial locality on vector data. Its evaluation was conducted on the Cavatools simulator, comparing the performance to a standard conventional cache, over different typical vector benchmarks for several vector lengths. The results proved the proposed cache speeds up performance on stride- 1 vector benchmarks, while hardly impacting non-stride-1’s. In addition, the prefetching feature consistently provided an additional value.
Susana Rebolledo Ruiz, Borja Pérez 0001, José Luis Bosque, Peter Hsu
ICPADS2
2021 Sigmoid: An auto-tuned load balancing algorithm for heterogeneous systems
abstract
A challenge that heterogeneous system programmers face is leveraging the performance of all the devices that integrate the system. This paper presents Sigmoid, a new load balancing algorithm that efficiently co-executes a single OpenCL data-parallel kernel on all the devices of heterogeneous systems. Sigmoid splits the workload proportionally to the capabilities of the devices, drastically reducing response time and energy consumption. It is designed around several features; it is dynamic, adaptive, guided and effortless, as it does not require the user to give any parameter, adapting to the behaviour of each kernel at runtime. To evaluate Sigmoid's performance, it has been implemented in Maat, a system abstraction library. Experimental results with different kernel types show that Sigmoid exhibits excellent performance, reaching a utilization of 90%, together with energy savings up to 20%, always reducing programming effort compared to OpenCL, and facilitating the portability to other heterogeneous machines.
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide
J. Parallel Distributed Comput.1
2019 Auto-tuned OpenCL kernel co-execution in OmpSs for heterogeneous systems
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé
J. Parallel Distributed Comput.1
2019 Load balancing in a heterogeneous world: CPU-Xeon Phi co-execution of data-parallel kernels
Raúl Nozal, Borja Pérez 0001, José Luis Bosque, Ramón Beivide
J. Supercomput.2
2017 To Distribute or Not to Distribute: The Question of Load Balancing for Performance or Energy
Esteban Stafford, Borja Pérez 0001, José Luis Bosque, Ramón Beivide, Mateo Valero
Euro-Par2
2017 Extending OmpSs for OpenCL Kernel Co-Execution in Heterogeneous Systems
abstract
Heterogeneous systems have a very high potential performance but present difficulties in their programming. OmpSs is a well known framework for task based parallel applications, which is an interesting tool to simplify the programming of these systems. However, it does not support the co-execution of a single OpenCL kernel instance on several compute devices. To overcome this limitation, this paper presents an extension of the OmpSs framework that solves two main objectives: the automatic division of datasets among several devices and the management of their memory address spaces. To adapt to different kinds of applications, the data division can be performed by the novel HGuided load balancing algorithm or by the well known Static and Dynamic. All this is accomplished with negligible impact on the programming. Experimental results reveal that there is always one load balancing algorithm that improves the performance and energy consumption of the system.
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé
SBAC-PAD1
2017 Energy efficiency of load balancing for data-parallel applications in heterogeneous systems
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide
J. Supercomput.1