VLDB 2026 Research / reviewers in the wild / expert
Jaume Bosch
dblp:180/6658
· DBLP profile ↗
8ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-4040-3416ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 62% Processor architecture and microarchitecture · 17% Reconfigurable computing and FPGAs · 12% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
task-based programming |
1.3 | 3 | 2021 | OmpSs@FPGA Framework for High Performance FPGA Computing · IEEE Trans. Computers 2021 Breaking master-slave model between host and FPGAs · PPoPP 2020 A Hardware Runtime for Task-Based Programming Models · IEEE Trans. Parallel Distributed Syst. 2019 |
Parallel and multicore computing
parallel programming models |
0.6 | 2 | 2021 | OmpSs@FPGA Framework for High Performance FPGA Computing · IEEE Trans. Computers 2021 Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core Processor · MICRO 2019 |
Parallel and multicore computing › parallel programming models
dataflow programming |
0.5 | 1 | 2021 | OmpSs@FPGA Framework for High Performance FPGA Computing · IEEE Trans. Computers 2021 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA offloading |
0.4 | 1 | 2020 | Breaking master-slave model between host and FPGAs · PPoPP 2020 |
Processor architecture and microarchitecture
chip multiprocessor |
0.4 | 1 | 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core Processor · MICRO 2019 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.4 | 1 | 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core Processor · MICRO 2019 |
Processor architecture and microarchitecture › special-purpose processor
tightly-coupled accelerator |
0.4 | 1 | 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core Processor · MICRO 2019 |
Reconfigurable computing and FPGAs
FPGA SoC |
0.1 | 1 | 2019 | A Hardware Runtime for Task-Based Programming Models · IEEE Trans. Parallel Distributed Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
task-based programming · 0.4OpenMP tasking · 0.4task dependence analysis · 0.4runtime dependence inference · 0.4heterogeneous task scheduling · 0.4hardware task scheduling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | OmpSs@FPGA Framework for High Performance FPGA ComputingabstractThis article presents the new features of the OmpSs@FPGA framework. OmpSs is a data-flow programming model that supports task nesting and dependencies to target asynchronous parallelism and heterogeneity. OmpSs@FPGA is the extension of the programming model addressed specifically to FPGAs. OmpSs environment is built on top of Mercurium source to source compiler and Nanos++ runtime system. To address FPGA specifics Mercurium compiler implements several FPGA related features as local variable caching, wide memory accesses or accelerator replication. In addition, part of the Nanos++ runtime has been ported to hardware. Driven by the compiler this new hardware runtime adds new features to FPGA codes, such as task creation and dependence management, providing both performance increases and ease of programming. To demonstrate these new capabilities, different high performance benchmarks have been evaluated over different FPGA platforms using the OmpSs programming model. The results demonstrate that programs that use the OmpSs programming model achieve very competitive performance with low to moderate porting effort compared to other FPGA implementations. Juan Miguel De Haro Ruiz, Jaume Bosch, Antonio Filgueras, Miquel Vidal, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
IEEE Trans. Computers | 2 |
| 2020 | Breaking master-slave model between host and FPGAsabstractThis paper proposes to enhance current task-based programming models by breaking their current master-slave approach between the main processor and its hardware accelerators. As a proof-of-concept, it presents an extension of the [email protected] toolchain that allows the tasks offloaded into the FPGA to create and synchronize nested tasks on their own without involving the host. Those FPGA spawned tasks may target the host to execute code not suitable for the FPGA, like system calls or I/O operations; or target other kernel accelerators inside the same FPGA. In addition to the programmability benefits of this new feature, the proposed system presents significant performance improvements and a better productivity over the classical master-slave approach. Jaume Bosch, Miquel Vidal, Antonio Filgueras, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé |
PPoPP | 1 |
| 2020 | Asynchronous runtime with distributed manager for task-based programming models
Jaume Bosch, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé |
Parallel Comput. | 1 |
| 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core ProcessorabstractTask Parallelism is a parallel programming model that provides code annotation constructs to outline tasks and describe how their pointer parameters are accessed so that they might be executed in parallel, and asynchronously, by a runtime capable of inferring and honoring their data dependence relationships. It is supported by several parallelization frameworks, as OpenMP and StarSs. Lucas Morais, Vitor Silva, Alfredo Goldman, Carlos Álvarez 0001, Jaume Bosch, Michael Frank 0008, Guido Araujo |
MICRO | 5 |
| 2019 | A Hardware Runtime for Task-Based Programming ModelsabstractTask-based programming models such as OpenMP 5.0 and OmpSs are simple to use and powerful enough to exploit task parallelism of applications over multicore, manycore and heterogeneous systems. However, their software-only runtimes introduce relevant overhead when targeting fine-grained tasks, resulting in performance losses. To overcome this drawback, we present a hardware runtime Picos++ that accelerates critical runtime functions such as task dependence analysis, nested task support, and heterogeneous task scheduling. As a proof-of-concept, the Picos++ hardware runtime has been integrated with a compiler infrastructure that supports parallel task-based programming models. A FPGA SoC running Linux OS has been used to implement the hardware accelerated part of Picos++, integrated with a heterogeneous system composed of 4 symmetric multiprocessor (SMP) cores and several hardware functional accelerators (HwAccs) for task execution. Results show significant improvements on energy and performance compared to state-of-the-art parallel software-only runtimes. With Picos++, applications can achieve up to 7.6x speedup and save up to 90 percent of energy, when using 4 threads and up to 4 HwAccs, and even reach a speedup of 16x over the software alternative when using 12 HwAccs and small tasks. Xubin Tan, Jaume Bosch, Carlos Álvarez 0001, Daniel Jiménez-González, Eduard Ayguadé, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Application Acceleration on FPGAs with OmpSs@FPGAabstractOmpSs@FPGA is the flavor of OmpSs that allows offloading application functionality to FPGAs. Similarly to OpenMP, it is based on compiler directives. While the OpenMP specification also includes support for heterogeneous execution, we use OmpSs and OmpSs@FPGA as prototype implementation to develop new ideas for OpenMP. OmpSs@FPGA implements the tasking model with runtime support to automatically exploit all SMP and FPGA resources available in the execution platform. In this paper, we present the OmpSs@FPGA ecosystem, based on the Mercurium compiler and the Nanos++ runtime system. We show how the applications are transformed to run on the SMP cores and the FPGA. The application kernels defined as tasks to be accelerated, using the OmpSs directives are: 1) transformed by the compiler into kernels connected with the proper synchronization and communication ports, 2) extracted to intermediate files, 3) compiled through the FPGA vendor HLS tool, and 4) used to configure the FPGA. Our Nanos++ runtime system schedules the application tasks on the platform, being able to use the SMP cores and the FPGA accelerators at the same time. We present the evaluation of the OmpSs@FPGA environment with the Matrix Multiplication, Cholesky and N-Body benchmarks, showing the internal details of the execution, and the performance obtained on a Zynq Ultrascale+ MPSoC (up to 128x). The source code uses OmpSs@FPGA annotations and different Vivado HLS optimization directives are applied for acceleration. Jaume Bosch, Xubin Tan, Antonio Filgueras, Miquel Vidal, Marc Mateu, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
FPT | 1 |
| 2017 | General Purpose Task-Dependence Management Hardware for Task-Based Dataflow Programming ModelsabstractTask-based programming models such as OpenMP, IntelTBB and OmpSs offer the possibility of expressing dependences among tasks to drive their execution at runtime. Managing these dependences introduces noticeable overheads when targeting fine-grained tasks, diminishing the potential speedups or even introducing performance losses. To overcome this drawback, we present a general purpose hardware accelerator, Picos++, to manage the inter-task dependences efficiently in both time and energy. Our design also includes a novel nested task support. To this end, a new hardware/software co-design is presented to overcome the fact that nested tasks with dependences could result in system deadlocks due to the limited amount of resources in hardware task dependence managers. In this paper we describe a detailed implementation of this design and evaluate a parallel task-based programming model using Picos++ in a Linux embedded system with two ARM Cortex-A9 and a FPGA. The scalability and energy consumption of the real system implemented have been studied and compared against a software runtime. Even in a system limited to 2 threads, using Picos++ results in more than 1.8x speedup and 40% of energy savings in the most demanding parallelizations of real benchmarks. As a matter of fact, a hardware task dependence manager should be able to achieve much higher speedup and provide more energy savings with more threads. Xubin Tan, Jaume Bosch, Miquel Vidal, Carlos Álvarez 0001, Daniel Jiménez-González, Eduard Ayguadé, Mateo Valero |
IPDPS | 2 |
| 2016 | Performance analysis of a hardware accelerator of dependence management for task-based dataflow programming modelsabstractAlong with the popularity of multicore and manycore, task-based dataflow programming models obtain great attention for being able to extract high parallelism from applications without exposing the complexity to programmers. One of these pioneers is the OpenMP Superscalar (OmpSs). By implementing dynamic task dependence analysis, dataflow scheduling and out-of-order execution in runtime, OmpSs achieves high performance using coarse and medium granularity tasks. In theory, for the same application, the more parallel tasks can be exposed, the higher possible speedup can be achieved. Yet this factor is limited by task granularity, up to a point where the runtime overhead outweighs the performance increase and slows down the application. To overcome this handicap, Picos was proposed to support task-based dataflow programming models like OmpSs as a fast hardware accelerator for fine-grained task and dependence management, and a simulator was developed to perform design space exploration. This paper presents the very first functional hardware prototype inspired by Picos. An embedded system based on a Zynq 7000 All-Programmable SoC is developed to study its capabilities and possible bottlenecks. Initial scalability and hardware consumption studies of different Picos designs are performed to find the one with the highest performance and lowest hardware cost. A further thorough performance study is employed on both the prototype with the most balanced configuration and the OmpSs software-only alternative. Results show that our OmpSs runtime hardware support significantly outperforms the software-only implementation currently available in the runtime system for fine-grained tasks. Xubin Tan, Jaume Bosch, Daniel Jiménez-González, Carlos Álvarez 0001, Eduard Ayguadé, Mateo Valero |
ISPASS | 2 |