EDBT 2026 Demo / reviewers in the wild / expert
Xavier Teruel
dblp:84/2461
· DBLP profile ↗
9ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0001-5181-7545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models › task parallelism
OpenMP tasking |
0.1 | 1 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.1 | 1 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 |
Runtime systems and virtual machines
parallel runtime systems |
0.0 | 1 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 |
Methods — techniques the papers use, named apart from their topics
task-based parallelism · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
Ruimin Shi, Maya B. Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Bo Peng |
Euro-Par (1) | 4 |
| 2020 | Evaluating Worksharing Tasks on Distributed EnvironmentsabstractHybrid programming is a promising approach to exploit clusters of multicore systems. Our focus is on the combination of MPI and tasking. This hybrid approach combines the low-latency and high throughput of MPI with the flexibility of tasking models and their inherent ability to handle load imbalance. However, combining tasking with standard MPI implementations can be a challenge. The Task-Aware MPI library (TAMPI) eases the development of applications combining tasking with MPI. TAMPI enables developers to overlap computation and communication phases by relying on the tasking data-flow execution model. Using this approach, the original computation that was distributed in many different MPI ranks is grouped together in fewer MPI ranks, and split into several tasks per rank. Nevertheless, programmers must be careful with task granularity. Too fine-grained tasks introduce too much overhead, while too coarse-grained tasks lead to lack of parallelism. An adequate granularity may not always exist, especially in distributed environments where the same amount of work is distributed among many more cores. Worksharing tasks are a special kind of tasks, recently proposed, that internally leverage worksharing techniques. By doing so, a single worksharing task may run in several cores concurrently. Nonetheless, the task management costs remain the same than a regular task. In this work, we study the combination of worksharing tasks and TAMPI on distributed environments using two well known mini-apps: HPCCG and LULESH. Our results show significant improvements using worksharing tasks compared to regular tasks, and to other state-of-the-art alternatives such as OpenMP worksharing. Marcos Maronas, Xavier Teruel, J. Mark Bull, Eduard Ayguadé, Vicenç Beltran 0001 |
CLUSTER | 2 |
| 2020 | HDOT - An approach towards productive programming of hybrid applications
Jan Ciesko, Pedro J. Martínez-Ferrer, Raúl Peñacoba Veigas, Xavier Teruel, Vicenç Beltran 0001 |
J. Parallel Distributed Comput. | 4 |
| 2019 | Auto-tuned OpenCL kernel co-execution in OmpSs for heterogeneous systems
Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé |
J. Parallel Distributed Comput. | 6 |
| 2019 | Integrating blocking and non-blocking MPI primitives with task-based programming models
Kevin Sala, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Vicenç Beltran 0001, Jesús Labarta |
Parallel Comput. | 2 |
| 2018 | Improving the Interoperability between MPI and Task-Based Programming ModelsabstractIn this paper we propose an API to pause and resume task execution depending on external events. We leverage this generic API to improve the interoperability between MPI synchronous communication primitives and tasks. When an MPI operation blocks, the task running is paused so that the runtime system can schedule a new task on the core that became idle. Once the MPI operation is completed, the paused task is put again on the runtime system's ready queue. We expose our proposal through a new MPI threading level which we implement through two approaches. Kevin Sala, Jorge Bellón, Pau Farré, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Daniel J. Holmes, Vicenç Beltran 0001, Jesús Labarta |
EuroMPI | 4 |
| 2017 | Extending OmpSs for OpenCL Kernel Co-Execution in Heterogeneous SystemsabstractHeterogeneous systems have a very high potential performance but present difficulties in their programming. OmpSs is a well known framework for task based parallel applications, which is an interesting tool to simplify the programming of these systems. However, it does not support the co-execution of a single OpenCL kernel instance on several compute devices. To overcome this limitation, this paper presents an extension of the OmpSs framework that solves two main objectives: the automatic division of datasets among several devices and the management of their memory address spaces. To adapt to different kinds of applications, the data division can be performed by the novel HGuided load balancing algorithm or by the well known Static and Dynamic. All this is accomplished with negligible impact on the programming. Experimental results reveal that there is always one load balancing algorithm that improves the performance and energy consumption of the system. Borja Pérez 0001, Esteban Stafford, José Luis Bosque, Ramón Beivide, Sergi Mateo, Xavier Teruel, Xavier Martorell, Eduard Ayguadé |
SBAC-PAD | 6 |
| 2009 | Barcelona OpenMP Tasks Suite: A Set of Benchmarks Targeting the Exploitation of Task Parallelism in OpenMPabstractTraditional parallel applications have exploited regular parallelism, based on parallel loops. Only a few applications exploit sections parallelism. With the release of the new OpenMP specification (3.0), this programming model supports tasking. Parallel tasks allow the exploitation of irregular parallelism, but there is a lack of benchmarks exploiting tasks in OpenMP. With the current (and projected) multicore architectures that offer many more alternatives to execute parallel applications than traditional SMP machines, this kind of parallelism is increasingly important. And so, the need to have some set of benchmarks to evaluate it. In this paper, we motivate the need of having such a benchmarks suite, for irregular and/or recursive task parallelism. We present our proposal, the Barcelona OpenMP Tasks Suite (BOTS), with a set of applications exploiting regular and irregular parallelism, based on tasks. We present an overall evaluation of the BOTS benchmarks in an Altix system and we discuss some of the different experiments that can be done with the different compilation and runtime alternatives of the benchmarks. Alejandro Duran, Xavier Teruel, Roger Ferrer, Xavier Martorell, Eduard Ayguadé |
ICPP | 2 |
| 2009 | The Design of OpenMP TasksabstractOpenMP has been very successful in exploiting structured parallelism in applications. With increasing application complexity, there is a growing need for addressing irregular parallelism in the presence of complicated control structures. This is evident in various efforts by the industry and research communities to provide a solution to this challenging problem. One of the primary goals of OpenMP 3.0 is to define a standard dialect to express and efficiently exploit unstructured parallelism. This paper presents the design of the OpenMP tasking model by members of the OpenMP 3.0 tasking sub-committee which was formed for this purpose. The paper summarizes the efforts of the sub-committee (spanning over two years) in designing, evaluating and seamlessly integrating the tasking model into the OpenMP specification. In this paper, we present the design goals and key features of the tasking model, including a rich set of examples and an in-depth discussion of the rationale behind various design choices. We compare a prototype implementation of the tasking model with existing models, and evaluate it on a wide range of applications. The comparison shows that the OpenMP tasking model provides expressiveness, flexibility, and huge potential for performance and scalability. Eduard Ayguadé, Nawal Copty, Alejandro Duran, Jay P. Hoeflinger, Federico Massaioli, Xavier Teruel, Priya Unnikrishnan, Guansong Zhang |
IEEE Trans. Parallel Distributed Syst. | 7 |