EDBT 2026 Demo / reviewers in the wild / expert
Vicenç Beltran 0001
dblp:08/2158 · also Vicenç Beltran Querol
· DBLP profile ↗
49ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-3580-9630ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 6 first-author · 15 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process WorkloadsabstractThe convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that place significant stress on traditional OS schedulers. When oversubscribed (there are more ready threads than cores), OS schedulers rely on periodic preemptions to multiplex cores, often introducing interference that may degrade performance. In this paper, we present: (1) The User-space Scheduling Framework (USF), a novel seamless process scheduling framework completely implemented in user-space. USF enables users to implement their own process scheduling algorithms without requiring special permissions. We evaluate USF with its default cooperative policy, (2) SCHED_COOP, designed to reduce interference by switching threads only upon blocking. This approach mitigates well-known issues such as Lock-Holder Preemption (LHP), Lock-Waiter Preemption (LWP), and scalability collapse. We implement USF and SCHED_COOP by extending the GNU C library with the nOS-V runtime, enabling seamless coordination across multiple runtimes (e.g., OpenMP) without requiring invasive application changes. Evaluations show gains up to 2.4x in oversubscribed multi-process scenarios, including nested BLAS workloads, multi-process PyTorch inference with LLaMA-3, and Molecular Dynamics (MD) simulations. Aleix Roca, Vicenç Beltran 0001 |
PPoPP | 2 |
| 2026 | A task-based data-flow methodology for programming heterogeneous systems with multiple accelerator APIs
Aleix Boné, Alejandro Aguirre 0005, David Álvarez 0006, Pedro J. Martínez-Ferrer, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 5 |
| 2025 | Enhancing OmpSs-2 Suspendable Tasks by Combining Operating System and User-Level Threads with C++ CoroutinesabstractThis paper explores three methods for implementing suspendable tasks within task-based programming models: OS threads (pthreads), User-Level Threads (ULTs), and C++ coroutines. We enhance the OmpSs-2 programming model, originally supporting suspendable tasks via pthreads, to also accommodate ULTs and C++ coroutines. This unified approach facilitates a comprehensive comparative analysis using various benchmarks that includes recursive fork-join and data-flow parallelization strategies. Additionally, we contrast these suspension methods with the Cilk and OpenMP task-based programming models, which, despite their efficiency, lack support for suspendable tasks. Key contributions of this study include the novel integration of C++20 coroutines into the OmpSs-2 programming model, which can be combined with pthreads or ULTs. Furthermore, we introduce a new Linux kernel syscall that accelerates pthread context switches by an order of magnitude, thus narrowing the performance gap between ULTs and pthreads in context-switch times from two to one order of magnitude. C++ Coroutines are the ideal solution for scenarios where many tasks are simultaneously in a suspended state, or the frequency of task suspension and resumption is high because they have the smallest memory footprint and minor contextswitch overhead. However, they are limited to C++ programs, do not support TLS, and only allow task suspension at top-level functions. Still, pthreads and ULTs can bring remarkable benefits where C++ Coroutines cannot. We conclude that combining the strengths of coroutines with pthreads or ULTs brings productivity and performance benefits for programming models. Arnau Cinca, Aleix Roca, Kevin Sala, Raúl Peñacoba Veigas, David Álvarez 0006, Vicenç Beltran 0001 |
IPDPS | 6 |
| 2025 | Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
Pedro J. Martínez-Ferrer, Albert-Jan Nicholas Yzelman, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | Leveraging iterative applications to improve the scalability of task-based programming models on distributed systemsabstractDistributed tasking models such as OmpSs-2@Cluster, StarPU-MPI, and PaRSEC express HPC applications as task graphs with explicit dependencies. The single task graph unifies the representation of parallelism across CPU cores, accelerators, and distributed-memory nodes, offering higher programmer productivity compared to traditional MPI + X. Most task-based models construct the task graph sequentially, which provides a clear and familiar programming model, simplifying code development, maintenance, and porting. However, this design introduces a bottleneck in task creation and dependency management, limiting performance and scalability. As a result, unless the tasks are very coarse-grained, current distributed sequential tasking models cannot match the performance of MPI + X. Many scientific applications, however, are iterative in nature, constructing the same directed acyclic task graph at each timestep. We exploit this structure to eliminate the sequential bottleneck and control message overhead in a sequentially-constructed distributed tasking model, while preserving its simplicity and productivity. Our approach builts on the recently proposed taskiter directive for OpenMP and OmpSs-2, allowing a single iteration to be expressed as a cyclic graph. The runtime partitions the cyclic graph across nodes, precomputes the MPI transfers, and then executes the loop body at low overhead. By integrating the MPI communications directly into the application’s task graph, our approach naturally overlaps computation and communication, in some cases exposing dramatically more parallelism than fork–join MPI + OpenMP. We define the programming model and describe the full runtime implementation, and integrate our proposal into OmpSs-2@Cluster. We evaluate it using five benchmarks on up to 128 nodes of the MareNostrum 5 supercomputer. For applications with fork–join parallelism, our approach has performance similar to fork–join MPI + OpenMP, making it a viable productive alternative, unlike the existing OmpSs-2@Cluster model, which is up to 7.7 times slower than MPI + OpenMP. For a 2D Gauss–Seidel stencil computation, our approach enables 3D wavefront computation, giving performance up to 22 times faster than fork–join MPI + OpenMP and on-a-par with state-of-the-art TAMPI + OmpSs-2. All software, comprising the compiler, runtime, and benchmarks, is released open source. 1 Omar Shaaban Ibrahim ali, Juliette Fournis d'Albiat, Isabel Piedrahita, Vicenç Beltran 0001, Xavier Martorell, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | nOS-V: Co-Executing HPC Applications Using System-Wide Task SchedulingabstractFuture Exascale systems will feature massive parallelism, many-core processors and heterogeneous architectures. In this scenario, it is increasingly difficult for HPC applications to fully and efficiently utilize the resources in system nodes. Moreover, the increased parallelism exacerbates the effects of existing inefficiencies in current applications. Research has shown that co-scheduling applications to share system nodes instead of executing each application exclusively can increase resource utilization and efficiency. Nevertheless, the current oversubscription and co-location techniques to share nodes have several drawbacks which limit their applicability and make them very application-dependent.This paper presents co-execution through system-wide scheduling. Co-execution is a novel fine-grained technique to execute multiple HPC applications simultaneously on the same node, outperforming current state-of-the-art approaches. We implement this technique in nOS-V, a lightweight tasking library that supports co-execution through system-wide task scheduling. Moreover, nOS-V can be easily integrated with existing programming models, requiring no changes to user applications. We showcase how co-execution with nOS-V significantly reduces schedule makespan for several applications on different scenarios, outperforming prior node-sharing techniques. David Álvarez 0006, Kevin Sala, Vicenç Beltran 0001 |
IPDPS | 3 |
| 2023 | Assessing Saiph, a task-based DSL for high-performance computational fluid dynamics
Sandra Macià, Pedro J. Martínez-Ferrer, Eduard Ayguadé, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 4 |
| 2023 | Improving the performance of classical linear algebra iterative methods via hybrid parallelism
Pedro J. Martínez-Ferrer, Tufan Arslan, Vicenç Beltran 0001 |
J. Parallel Distributed Comput. | 3 |
| 2023 | Mitigating the NUMA effect on task-based runtime systems
Marcos Maronas, Antoni C. Navarro, Eduard Ayguadé, Vicenç Beltran 0001 |
J. Supercomput. | 4 |
| 2022 | OmpSs-2@Cluster: Distributed Memory Execution of Nested OpenMP-style Tasks
Jimmy Aguilar Mena, Omar Shaaban, Vicenç Beltran 0001, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
Euro-Par | 3 |
| 2022 | Seamless optimization of the GEMM kernel for task-based programming modelsabstractThe general matrix-matrix multiplication (GEMM) kernel is a fundamental building block of many scientific applications. Many libraries such as Intel MKL and BLIS provide highly optimized sequential and parallel versions of this kernel. The parallel implementations of the GEMM kernel rely on the well-known fork-join execution model to exploit multi-core systems efficiently. However, these implementations are not well suited for task-based applications as they break the data-flow execution model. In this paper, we present a task-based implementation of the GEMM kernel that can be seamlessly leveraged by task-based applications while providing better performance than the fork-join version. Our implementation leverages several advanced features of the OmpSs-2 programming model and a new heuristic to select the best parallelization strategy and blocking parameters based on the matrix and hardware characteristics. When evaluating the performance and energy consumption on two modern multi-core systems, we show that our implementations provide significant performance improvements over an optimized OpenMP fork-join implementation, and can beat vendor implementations of the GEMM (e.g., Intel MKL and AMD AOCL). We also demonstrate that a real application can leverage our optimized task-based implementation to enhance performance. Arthur Francisco Lorenzon, Sandro Matheus V. N. Marques, Antoni C. Navarro, Vicenç Beltran 0001 |
ICS | 4 |
| 2022 | Automatic aggregation of subtask accesses for nested OpenMP-style tasksabstractTask-based programming is a high performance and productive model to express parallelism. Tasks encapsulate work to be executed across multiple cores or offloaded to GPUs, FPGAs, other accelerators or other nodes. In order to maintain parallelism and afford maximum freedom to the scheduler, the task dependency graph should be created in parallel and well in advance of task execution. A key limitation with OpenMP and OmpSs-2 tasking is that a task cannot be created until all its accesses and its descendents' accesses are known. Current approaches to work around this limitation either stop task creation and execution using a taskwait or they substitute “fake” accesses known as sentinels. This paper proposes the auto clause, which indicates that the task may create subtasks that access unspecified memory regions or it may allocate and return memory at addresses that are of course not yet known. Unlike approaches using taskwaits, there is no interruption to the concurrent creation and execution of tasks, maintaining parallelism and the scheduler's ability to optimize load balance and data locality. Unlike existing approaches using sentinels, all tasks can be given a precise specification of their own data accesses, so that a single mechanism is used to control task ordering, program data transfers on distributed memory and optimize data locality, e.g. on NUMA systems. The auto clause also provides an incremental path to develop programs with nested tasks, by removing the need for every parent task to have a complete specification of the accesses of its descendent tasks. This is redundant information that can be time consuming and error-prone to describe. We present a straightforward runtime implementation that achieves a 1.4 times speedup for n-body with OmpSs-2@Cluster task offloading to 32 nodes and <4% slowdown for three benchmarks with task offloading to 8 nodes. All code is open source. Omar Shaaban, Jimmy Aguilar Mena, Vicenç Beltran 0001, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
SBAC-PAD | 3 |
| 2022 | A Native Tensor-Vector Multiplication Algorithm for High Performance ComputingabstractTensor computations are important mathematical operations for applications that rely on multidimensional data. The tensor–vector multiplication (TVM) is the most memory-bound tensor contraction in this class of operations. This article proposes an open-source TVM algorithm which is much simpler and efficient than previous approaches, making it suitable for integration in the most popular BLAS libraries available today. Our algorithm has been written from scratch and features unit-stride memory accesses, cache awareness, mode obliviousness, full vectorization and multi-threading as well as NUMA awareness for non-hierarchically stored dense tensors. Numerical experiments are carried out on tensors up to order 10 and various compilers and hardware architectures equipped with traditional DDR and high bandwidth memory (HBM). For large tensors the average performance of the TVM ranges between 62% and 76% of the theoretical bandwidth for NUMA systems with DDR memory and remains independent of the contraction mode. On NUMA systems with HBM the TVM exhibits some mode dependency but manages to reach performance figures close to peak values. Finally, the higher-order power method is benchmarked with the proposed TVM kernel and delivers on average between 58% and 69% of the theoretical bandwidth for large tensors. Pedro J. Martínez-Ferrer, Albert-Jan Nicholas Yzelman, Vicenç Beltran 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Combining One-Sided Communications with Task-Based Programming ModelsabstractHybrid programming combining task-based and message-passing models is an increasingly popular technique to exploit multi-core clusters. The Task-Aware MPI (TAMPI) library integrates both models enabling the safe overlap of computation and communication tasks using two-sided MPI communications. Two-sided primitives combine data transfers with implicit synchronizations, but one-sided models usually offer more efficient data transfers decoupling synchronizations. MPI offers four distinct one-sided synchronization modes, while GASPI is a PGAS API providing one-sided operations with remote notifications for fine inter-process synchronizations.In this paper, we study the challenges of integrating MPI and GASPI one-sided operations with the OpenMP and OmpSs-2 tasking models. We propose and implement several extensions to the GASPI and OmpSs-2 programming models, which are leveraged by a new library called Task-Aware GASPI (TAGASPI). The TAGASPI library allows the efficient and safe use of one-sided operations with remote notifications inside tasks. Both TAGASPI and TAMPI transparently manage communications issued by tasks and allow these to overlap with computation tasks naturally, following a data-flow model. These libraries are complementary and can be mixed in the same application.Our experience porting several mini-apps to this hybrid model shows that TAGASPI helps leverage one-sided communications with similar complexity to pure and hybrid two-sided MPI approaches. We show that our hybrid one-sided approach outperforms the pure MPI strategies, but it also surpasses the TAMPI’s performance when stressing communication phases, e.g., increasing the communication parallelism and reducing the communication tasks’ sizes. Kevin Sala, Sandra Macià, Vicenç Beltran 0001 |
CLUSTER | 3 |
| 2021 | Combining Dynamic Concurrency Throttling with Voltage and Frequency Scaling on Task-based Programming ModelsabstractBeing on the verge of exascale performance has shifted the prioritization of performance in applications to the inclusion of power-performance efficiency as a primary objective in the High Performance Computing (HPC) community. Simultaneously, this has surfaced hardware and software efforts that employ techniques such as dynamic voltage and frequency scaling (DVFS) for core and uncore units or dynamic concurrency throttling (DCT) to exploit hardware resources efficiently, by saving energy while maintaining performance. These techniques are complementary, so they can be used together. However, employing them is not a straightforward task, as they have to be adjusted based on the workload, and it is even more complex to combine them properly. Thus, these techniques should be applied transparently by a runtime system, without relying on application developers. In this paper, we extend a task-based runtime system with an infrastructure that categorizes workloads based on their computational profile – memory-bounded, compute-bounded, or balanced. This categorization is done in an on-line manner and with a negligible overhead. With this additional information, we enhance the CPU-manager and scheduler of OmpSs-2, a task-based parallel programming model, to automatically combine DVFS and DCT techniques based on workloads. Moreover, we show that our heuristics transparently improve energy efficiency on average by 15% with no significant performance loss and either equal or surpass the energy efficiency of the best static configuration available. Antoni Navarro Muñoz, Arthur Francisco Lorenzon, Eduard Ayguadé, Vicenç Beltran 0001 |
ICPP | 4 |
| 2021 | Advanced synchronization techniques for task-based runtime systemsabstractTask-based programming models like OmpSs-2 and OpenMP provide a flexible data-flow execution model to exploit dynamic, irregular and nested parallelism. Providing an efficient implementation that scales well with small granularity tasks remains a challenge, and bottlenecks can manifest in several runtime components. In this paper, we analyze the limiting factors in the scalability of a task-based runtime system and propose individual solutions for each of the challenges, including a wait-free dependency system and a novel scalable scheduler design based on delegation. We evaluate how the optimizations impact the overall performance of the runtime, both individually and in combination. We also compare the resulting runtime against state of the art OpenMP implementations, showing equivalent or better performance, especially for fine-grained tasks. David Álvarez 0006, Kevin Sala, Marcos Maronas, Aleix Roca, Vicenç Beltran 0001 |
PPoPP | 5 |
| 2020 | Evaluating Worksharing Tasks on Distributed EnvironmentsabstractHybrid programming is a promising approach to exploit clusters of multicore systems. Our focus is on the combination of MPI and tasking. This hybrid approach combines the low-latency and high throughput of MPI with the flexibility of tasking models and their inherent ability to handle load imbalance. However, combining tasking with standard MPI implementations can be a challenge. The Task-Aware MPI library (TAMPI) eases the development of applications combining tasking with MPI. TAMPI enables developers to overlap computation and communication phases by relying on the tasking data-flow execution model. Using this approach, the original computation that was distributed in many different MPI ranks is grouped together in fewer MPI ranks, and split into several tasks per rank. Nevertheless, programmers must be careful with task granularity. Too fine-grained tasks introduce too much overhead, while too coarse-grained tasks lead to lack of parallelism. An adequate granularity may not always exist, especially in distributed environments where the same amount of work is distributed among many more cores. Worksharing tasks are a special kind of tasks, recently proposed, that internally leverage worksharing techniques. By doing so, a single worksharing task may run in several cores concurrently. Nonetheless, the task management costs remain the same than a regular task. In this work, we study the combination of worksharing tasks and TAMPI on distributed environments using two well known mini-apps: HPCCG and LULESH. Our results show significant improvements using worksharing tasks compared to regular tasks, and to other state-of-the-art alternatives such as OpenMP worksharing. Marcos Maronas, Xavier Teruel, J. Mark Bull, Eduard Ayguadé, Vicenç Beltran 0001 |
CLUSTER | 5 |
| 2020 | Towards Data-Flow Parallelization for Adaptive Mesh Refinement ApplicationsabstractAdaptive Mesh Refinement (AMR) is a prevalent method used by distributed-memory simulation applications to adapt the accuracy of their solutions depending on the turbulent conditions in each of their domain regions. These applications are usually dynamic since their domain areas are refined or coarsened in various refinement stages during their execution. Thus, they periodically redistribute their workloads among processes to avoid load imbalance. Although the defacto standard for scientific computing in distributed environments is MPI, in recent years, pure MPI applications are being ported to hybrid ones, attempting to cope with modern multi-core systems. Recently, the Task-Aware MPI library was proposed to efficiently integrate MPI communications and tasking models, providing also the transparent management of communications issued by tasks. In this paper, we demonstrate the benefits of porting AMR applications to data-flow programming models leveraging that novel hybrid approach. We exploit most of the application parallelism by taskifying all stages, allowing their natural overlap. We employ these techniques on the miniAMR proxy application, which mimics the refinement, load balancing, communication, and computation patterns of general AMR applications. We evaluate how this approach reduces the time in its computation and communication phases while achieving better programmability than other conventional hybrid techniques. Kevin Sala, Alejandro Rico, Vicenç Beltran 0001 |
CLUSTER | 3 |
| 2020 | A Toolchain to Verify the Parallelization of OmpSs-2 Applications
Simone Economo, Sara Royuela, Eduard Ayguadé, Vicenç Beltran 0001 |
Euro-Par | 4 |
| 2020 | Enhancing Resource Management Through Prediction-Based Policies
Antoni C. Navarro, Arthur Francisco Lorenzon, Eduard Ayguadé, Vicenç Beltran 0001 |
Euro-Par | 4 |
| 2020 | Extending the OpenCHK Model with advanced checkpoint features
Marcos Maronas, Sergi Mateo, Kai Keller, Leonardo Arturo Bautista-Gomez, Eduard Ayguadé, Vicenç Beltran 0001 |
Future Gener. Comput. Syst. | 6 |
| 2020 | HDOT - An approach towards productive programming of hybrid applications
Jan Ciesko, Pedro J. Martínez-Ferrer, Raúl Peñacoba Veigas, Xavier Teruel, Vicenç Beltran 0001 |
J. Parallel Distributed Comput. | 5 |
| 2020 | Iteration-fusing conjugate gradient for sparse linear systems with MPI + OmpSs
Maria Barreda, José Ignacio Aliaga, Vicenç Beltran 0001, Marc Casas |
J. Supercomput. | 3 |
| 2019 | Worksharing Tasks: An Efficient Way to Exploit Irregular and Fine-Grained Loop ParallelismabstractShared memory programming models usually provide worksharing and task constructs. The former relies on the efficient fork-join execution model to exploit structured parallelism; while the latter relies on fine-grained synchronization among tasks and a flexible data-flow execution model to exploit dynamic, irregular, and nested parallelism. On applications that show both structured and unstructured parallelism, both worksharing and task constructs can be combined. However, it is difficult to mix both execution models without penalizing the data-flow execution model. Hence, on many applications structured parallelism is also exploited using tasks to leverage the full benefits of a pure data-flow execution model. However, task creation and management might introduce a non-negligible overhead that prevents the efficient exploitation of fine-grained structured parallelism, especially on many-core processors. In this work, we propose worksharing tasks. These are tasks that internally leverage worksharing techniques to exploit fine-grained structured loop-based parallelism. The evaluation shows promising results on several benchmarks and platforms. Marcos Maronas, Kevin Sala, Sergi Mateo, Eduard Ayguadé, Vicenç Beltran 0001 |
HiPC | 5 |
| 2019 | A Linux Kernel Scheduler Extension for Multi-core SystemsabstractThe Linux kernel is mostly designed for multi-programed environments, but high-performance applications have other requirements. Such applications are run standalone, and usually rely on runtime systems to distribute the application's workload on worker threads, one per core. However, due to current OSes limitations, it is not feasible to track whether workers are actually running or blocked due to, for instance, a requested resource. For I/O intensive applications, this leads to a significant performance degradation given that the core of a blocked thread becomes idle until it is able to run again. In this paper, we present the proof-of-concept of a Linux kernel extension denoted User-Monitored Threads (UMT) which tackles this problem. Our extension allows a user-space process to be notified of when the selected threads become blocked or unblocked, making it possible for a runtime to schedule additional work on the idle core. We implemented the extension on the Linux Kernel 5.1 and adapted the Nanos6 runtime of the OmpSs-2 programming model to take advantage of it. The whole prototype was tested on two applications which, on the tested hardware and the appropriate conditions, reported speedups of almost 2x. Aleix Roca, Samuel Rodríguez, Albert Segura, Kevin Marquet, Vicenç Beltran 0001 |
HiPC | 5 |
| 2019 | Integrating blocking and non-blocking MPI primitives with task-based programming models
Kevin Sala, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Vicenç Beltran 0001, Jesús Labarta |
Parallel Comput. | 5 |
| 2018 | Variable Batched DGEMMabstractMany scientific applications are in need to solve a high number of small-size independent problems. These individual problems do not provide enough parallelism and then, these must be computed as a batch. Today, vendors such as Intel and NVIDIA are developing their own suite of batch routines. Although most of the works focus on computing batches of fixed size, in real applications we can not assume a uniform size for all set of problems. We explore and analyze different strategies based on parallel for, task and taskloop OpenMP pragmas. Although these strategies are straightforward from a programmer's point of view, they have a different impact on performance. We also analyze a new prototype provided by Intel (MKL), which deals with batch operations (cblas_dgemm_batch). We propose a new approach called grouping. It basically groups a set of problems until filling a limit in terms of memory occupancy or number of operations. In this way, groups composed by different number of problems are distributed on cores, achieving a more balanced distribution in terms of computational cost. This strategy is able to be up to 6× faster than the Intel (MKL) batch routine. Pedro Valero-Lara, Ivan Martínez-Pérez, Sergi Mateo, Raül Sirvent, Vicenç Beltran 0001, Xavier Martorell, Jesús Labarta |
PDP | 5 |
| 2018 | Improving the Interoperability between MPI and Task-Based Programming ModelsabstractIn this paper we propose an API to pause and resume task execution depending on external events. We leverage this generic API to improve the interoperability between MPI synchronous communication primitives and tasks. When an MPI operation blocks, the task running is paused so that the runtime system can schedule a new task on the core that became idle. Once the MPI operation is completed, the paused task is put again on the runtime system's ready queue. We expose our proposal through a new MPI threading level which we implement through two approaches. Kevin Sala, Jorge Bellón, Pau Farré, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Daniel J. Holmes, Vicenç Beltran 0001, Jesús Labarta |
EuroMPI | 8 |
| 2018 | On the adequacy of lightweight thread approaches for high-level parallel programming models
Adrián Castelló 0001, Rafael Mayo 0002, Kevin Sala, Vicenç Beltran 0001, Pavan Balaji, Antonio J. Peña |
Future Gener. Comput. Syst. | 4 |
| 2018 | DMR API: Improving cluster productivity by turning applications into malleable
Sergio Iserte, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicenç Beltran 0001, Antonio J. Peña |
Parallel Comput. | 4 |
| 2017 | Supporting automatic recovery in offloaded distributed programming models through MPI-3 techniquesabstractIn this paper we describe the design of fault tolerance capabilities for general-purpose offload semantics, based on the OmpSs programming model. Using ParaStation MPI, a production MPI-3.1 implementation, we explore the features that, being standard compliant, an MPI stack must support to provide the necessary fault tolerance guarantees, based on MPI's dynamic process management. Our results, including synthetic benchmarks and applications, reveal low runtime overhead and efficient recovery, demonstrating that the existing MPI standard provided us with sufficient mechanisms to implement an effective and efficient fault-tolerant solution. Antonio J. Peña, Vicenç Beltran 0001, Carsten Clauss, Thomas Moschny |
ICS | 2 |
| 2017 | Improving the Integration of Task Nesting and Dependencies in OpenMPabstractThe tasking model of OpenMP 4.0 supports both nesting and the definition of dependences between sibling tasks. A natural way to parallelize many codes with tasks is to first taskify the high-level functions and then to further refine these tasks with additional subtasks. However, this top-down approach has some drawbacks since combining nesting with dependencies usually requires additional measures to enforce the correct coordination of dependencies across nesting levels. For instance, most non-leaf tasks need to include a taskwait at the end of their code. While these measures enforce the correct order of execution, as a side effect, they also limit the discovery of parallelism. In this paper we extend the OpenMP tasking model to improve the integration of nesting and dependencies. Our proposal builds on both formulas, nesting and dependencies, and benefits from their individual strengths. On one hand, it encourages a top-down approach to parallelizing codes that also enables the parallel instantiation of tasks. On the other hand, it allows the runtime to control dependencies at a fine grain that until now was only possible using a single domain of dependencies. Our proposal is realized through additions to the OpenMP task directive that ensure backward compatibility with current codes. We have implemented a new runtime with these extensions and used it to evaluate the impact on several benchmarks. Our initial findings show that our extensions improve performance in three areas. First, they expose more parallelism. Second, they uncover dependencies across nesting levels, which allows the runtime to make better scheduling decisions. And third, they allow the parallel instantiation of tasks with dependencies between them. Josep M. Pérez, Vicenç Beltran 0001, Jesús Labarta, Eduard Ayguadé |
IPDPS | 2 |
| 2015 | Collective Offload for Heterogeneous ClustersabstractExascale performance requires a level of energy efficiency only achievable with specialized hardware. Hence, for building a general purpose HPC system with Exascale performance different types of processors, memory technologies and interconnection networks will be necessary. Heterogeneous hardware is already present on some top supercomputer systems that are composed of different compute nodes, which at the same time, contain different types of processors and memories. Moreover, heterogeneous hardware is much harder to manage and exploit than homogeneous hardware, further increasing the complexity of applications that run on HPC systems. Most HPC applications use MPI to implement a rigid Single Program Multiple Data (SPMD) execution model that no longer fits the heterogeneous nature of the underlying hardware. However, MPI provides a powerful and flexible MPI_Comm_spawn API call that was designed to exploit heterogeneous hardware dynamically but at the expense of higher complexity, hindering a wider adoption of this API. In this paper, we have extended the OmpSs programming model to offload MPI kernels dynamically, replacing the low-level and more error-prone MPI_Comm_ spawn call with high-level and easier to use OmpSs pragmas. The evaluation shows that our proposal simplifies the dynamic offload of MPI kernels while keeping competitive performance and scaling to a high number of nodes. Florentino Sainz, Jorge Bellón, Vicenç Beltran 0001, Jesús Labarta |
HiPC | 3 |
| 2014 | Leveraging OmpSs to Exploit Hardware AcceleratorsabstractCUDA and OpenCL are the most widely used programming models to exploit hardware accelerators. Both programming models provide a C-based programming language to write accelerator kernels and a host API used to glue the host and kernel parts. Although this model is a clear improvement over a low-level and ad-hoc programming model for each hardware accelerator, it is still too complex and cumbersome for general adoption. For large and complex applications using several accelerators, the main problem becomes the explicit coordination and management of resources required between the host and the hardware accelerators that introduce a new family of issues (scheduling, data transfers, synchronization, ) that the programmer must take into account. In this paper, we propose a simple extension to OmpSs -- a data-flow programming model -- that dramatically simplifies the integration of accelerated code, in the form of CUDA or OpenCL kernels, into any C, C++ or Fortran application. Our proposal fully replaces the CUDA and OpenCL host APIs with a few pragmas, so we can leverage any kernel written in CUDA C or OpenCL C without any performance impact. Our compiler generates all the boilerplat code while our runtime system takes care of kernels scheduling, data transfers between host and accelerators and synchronizations between host and kernels parts. To evaluate our approach, we have ported several native CUDA and OpenCL applications to OmpSs by replacing all the CUDA or OpenCL API calls by a few number of pragmas. The OmpSs versions of these applications have competitive performance and scalability but with a significantly lower complexity than the original ones. Florentino Sainz, Sergi Mateo, Vicenç Beltran 0001, José Luis Bosque, Xavier Martorell, Eduard Ayguadé |
SBAC-PAD | 3 |
| 2012 | Optimizing resource utilization with software-based temporal multi-threading (stmt)abstractCompute and memory access units are two of the most important resources to appropriately manage in current and future multi-/many-core architectures. Memory bandwidth and computational capacity need to be exploited in a combined way to achieve the best system performance. Coarse-grain multi-threading, also known as temporal multi-threading (TMT), is a well known technique that improves overall resource utilization by time-multiplexing the execution of a reduced number of hardware threads that are switched in case of a high-latency event, such as a memory miss. Hence, the processor does not stall on memory misses and the number of in-fly memory operations is increased, improving the overall processor resource utilization. In this paper, we propose a software-based implementation of TMT that supports and unbounded number of threads and enables a flexible combination of multiple computational kernels. Our TMT implementation is based on micro-threads that combine fast cooperative and preemptive context switches to overcome some intrinsic limitations of current TMT hardware implementations, such as the reduced and fixed number of hardware threads available. Our proposal is demonstrated with an implementation on the CelllB.E. which is evaluated using heterogeneous mixes of memory-/CPU-bound kernels. Experimental results show how the proposed technique reduce the execution time of several benchmarks by up to 78%. Vicenç Beltran 0001, Eduard Ayguadé |
HiPC | 1 |
| 2012 | Energy accounting for shared virtualized environments under DVFS using PMC-based power models
Ramon Bertran Monfort, Yolanda Becerra 0001, David Carrera 0001, Vicenç Beltran 0001, Marc González 0001, Xavier Martorell, Nacho Navarro, Jordi Torres, Eduard Ayguadé |
Future Gener. Comput. Syst. | 4 |
| 2010 | Analysis of Task Offloading for Accelerators
Roger Ferrer, Vicenç Beltran 0001, Marc González 0001, Xavier Martorell, Eduard Ayguadé |
HiPEAC | 2 |
| 2010 | A CellBE-based HPC Application for the Analysis of Vulnerabilities in Cryptographic Hash FunctionsabstractAfter some recent breaks presented in the technical literature, it has become of paramount importance to gain a deeper understanding of the robustness and weaknesses of cryptographic hash functions. In particular, in the light of the recent attacks to the MD5 hash function, SHA-1 remains currently the only function that can be used in practice, since it is the only alternative to MD5 in many security standards. This work presents a study of vulnerabilities in the SHA family, namely the SHA-0 and SHA-1 hash functions, based on a high-performance computing application run on the MariCel cluster available at the Barcelona Supercomputing Center. The effectiveness of the different optimizations and search strategies that have been used is validated by a comprehensive set of quantitative evaluations, presented in the paper. Most importantly, at the conclusion of our study, we were able to identify an actual collision for a 71-round version of SHA-1, the first ever found so far. Alessandro Cilardo, Luigi Esposito, Antonio Veniero, Antonino Mazzeo, Vicenç Beltran 0001, Eduard Ayguadé |
HPCC | 5 |
| 2010 | Performance Management of Accelerated MapReduce Workloads in Heterogeneous ClustersabstractNext generation data centers will be composed of thousands of hybrid systems in an attempt to increase overall cluster performance and to minimize energy consumption. New programming models, such as MapReduce, specifically designed to make the most of very large infrastructures will be leveraged to develop massively distributed services. At the same time, data centers will bring an unprecedented degree of workload consolidation, hosting in the same infrastructure distributed services from many different users. In this paper we present our advancements in leveraging the Adaptive MapReduce Scheduler to meet user defined high level performance goals while transparently and efficiently exploiting the capabilities of hybrid systems. While the Adaptive Scheduler was already able to dynamically allocate resources to co-located MapReduce jobs based on their completion time goals, it was completely unaware of specific hardware capabilities. In our work we describe the changes introduced in the Adaptive Scheduler to enable it with hardware awareness and with the ability to co-schedule accelerable and non-accelerable jobs on the same heterogeneous MapReduce cluster, making the most of the underlying hybrid systems. The developed prototype is tested in a cluster of Cell/BE blades and relies on the use of accelerated and non-accelerated versions of the MapReduce tasks of different deployed applications to dynamically select the best version to run on each node. Decisions are made after workload composition and jobs' completion time goals. Results show that the augmented Adaptive Scheduler provides dynamic resource allocation across jobs, hardware affinity when possible, and is even able to spread jobs' tasks across accelerated and non-accelerated nodes in order to meet performance goals in extreme conditions. To our knowledge this is the first MapReduce scheduler and prototype that is able to manage high-level performance goals even in presence of hybrid systems and accelerable jobs. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 4 |
| 2009 | CellMT: A cooperative multithreading library for the Cell/B.EabstractThe Cell BE processor has proved that heterogeneous multi-core systems can provide a huge computational power with high efficiency for a wide range of applications. The simple design of the computational units and the use of small managed local memories is the key to achieve high efficiency and performance at the same time. However, this simple and efficient hardware design comes at the price of higher code complexity. The code written to run in this kind of processors must deal with several issues such as code vectorization, loop unrolling or the explicit management of local memories. Some of these issues such as vectorization or loop unrolling can be partially solved by the compiler, but the overlapping of data transfer and computation times must be manually addressed by the programmer with techniques such as double buffering that increase the code complexity. In this paper we present a user level threading library called CellMT that effectively hide memory latencies. The concurrent execution of several threads inside each SPU naturally overlaps computation and data transfer times without increasing the code complexity. To prove the suitability and feasibility of our multi-threaded library, we perform an exhaustive performance evaluation with a synthetic benchmark and a real application. The experimental results show that the multithreaded approach can outperform a hand-coded double buffering scheme, with speedups from 0.96x to 3.2x, while maintaining the complexity of a naive buffering scheme. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
HiPC | 1 |
| 2009 | Speeding Up Distributed MapReduce Applications Using Hardware AcceleratorsabstractIn an attempt to increase the performance/cost ratio, large compute clusters are becoming heterogeneous at multiple levels: from asymmetric processors, to different system architectures, operating systems and networks. Exploiting the intrinsic multi-level parallelism present in such a complex execution environment has become a challenging task using traditional parallel and distributed programming models. As a result, an increasing need for novel approaches to exploiting parallelism has arisen in these environments. MapReduce is a data-driven programming model originally proposed by Google back in 2004 as a flexible alternative to the existing models, specially devoted to hiding the complexity of both developing and running massively distributed applications in large compute clusters. In some recent works, the MapReduce model has been also used to exploit parallelism in other non-distributed environments, such as multi-cores, heterogeneous processors and GPUs. In this paper we introduce a novel approach for exploiting the heterogeneity of a Cell BE cluster linking an existing MapReduce runtime implementation for distributed clusters and one runtime to exploit the parallelism of the Cell BE nodes. The novel contribution of this work is the design and evaluation of a MapReduce execution environment that effectively exploits the parallelism existing at both the Cell BE cluster level and the heterogeneous processors level. Yolanda Becerra 0001, Vicenç Beltran 0001, David Carrera 0001, Marc González 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 2 |
| 2008 | Improving Web Server Performance Through Main Memory CompressionabstractCurrent web servers are highly multithreaded applications whose scalability benefits from the current multi-core/multiprocessor trend. However, some workloads cannot capitalize on this because their performance is limited by the available memory and/or the disk bandwidth, which prevents the server from taking advantage of the computing resources provided by the system. To solve this situation we propose the use of main memory compression techniques to increment the available memory and mitigate the disk band-width problem, allowing the web server to improve its use of CPU system resources. In this paper we implement to the Linux OS a full SMP capable main memory compression subsystem to increase the performance of a web server running the SPEC web 2005 benchmark. Although main memory compression is not a new technique perse, its use in a multicore environment running heavily multithreaded applications like a webserver introduces new challenges in the technique, such as scalability issues and the trade-off between the compressed memory size and the computational power required to achieve it. Finally, the evaluation of our implementaiton shows promising results such as a 30% web server throughput improvement and a 70% reduction in the disk bandwidth usage. Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPADS | 1 |
| 2008 | Understanding tuning complexity in multithreaded and hybrid web serversabstractAdequately setting up a multi-threaded Web server is a challenging task because its performance is determined by a combination of configurable Web server parameters and unsteady external factors like the workload type, workload intensity and machine resources available. Usually administrators set up Web server parameters like the keep-alive timeout and number of worker threads based on their experience and judgment, expecting that this configuration will perform well for the guessed uncontrollable factors. The nontrivial interaction between the configuration parameters of a multi-threaded Web server makes it a hard task to properly tune it for a given workload, but the burst nature of the Internet quickly change the uncontrollable factors and make it impossible to obtain an optimal configuration that will always perform well. In this paper we show the complexity of optimally configuring a multi-threaded Web server for different workloads with an exhaustive study of the interactions between the keep-alive timeout and the number of worker threads for a wide range of workloads. We also analyze the Hybrid Web server architecture (multi-threaded and event-driven) as a feasible solution to simplify Web server tuning and obtain the best performance for a wide range of workloads that can dynamically change in intensity and type. Finally, we compare the performance of the optimally tuned multithreaded Web server and the hybrid Web server with different workloads to validate our assertions. We conclude from our study that the hybrid architecture clearly outperforms the multi-threaded one, not only in terms of performance, but also in terms of its tuning complexity and its adaptability over different workload types. In fact, from the obtained results, we expect that the hybrid architecture is well suited to simplify the self configuration of complex application servers. Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
IPDPS | 1 |
| 2008 | Reducing wasted resources to help achieve green data centersabstractIn this paper we introduce a new approach to the consolidation strategy of a data center that allows an important reduction in the amount of active nodes required to process a heterogeneous workload without degrading the offered service level. This article reflects and demonstrates that consolidation of dynamic workloads does not end with virtualization. If energy-efficiency is pursued, the workloads can be consolidated even more using two techniques, memory compression and request discrimination, which were separately studied and validated in previous work and are now to be combined in a joint effort. We evaluate the approach using a representative workload scenario composed of numerical applications and a real workload obtained from a top national travel website. Our results indicate that an important improvement can be achieved using 20% less servers to do the same work. We believe that this serves as an illustrative example of a new way of management: tailoring the resources to meet high level energy efficiency goals. Jordi Torres, David Carrera 0001, Kevin Hogan, Ricard Gavaldà, Vicenç Beltran 0001, Nicolás Poggi |
IPDPS | 5 |
| 2008 | Dynamic CPU provisioning for self-managed secure web applications in SMP hosting platforms
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 3 |
| 2007 | Designing an overload control strategy for secure e-commerce applications
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 3 |
| 2005 | A Hybrid Web Server Architecture for Secure e-Business Web Applications
Vicenç Beltran 0001, David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé |
HPCC | 1 |
| 2005 | Session-Based Adaptive Overload Control for Secure Dynamic Web ApplicationsabstractAs dynamic Web content and security capabilities are becoming popular in current Web sites, the performance demand on application servers that host the sites is increasing, leading sometimes these servers to overload. As a result, response times may grow to unacceptable levels and the server may saturate or even crash. In this paper we present a session-based adaptive overload control mechanism based on SSL (secure socket layer) connections differentiation and admission control. The SSL connections differentiation is a key factor because the cost of establishing a new SSL connection is much greater than establishing a resumed SSL connection (it reuses an existing SSL session on server). Considering this big difference, we have implemented an admission control algorithm that prioritizes the resumed SSL connections to maximize performance on session-based environments and limits dynamically the number of new SSL connections accepted depending on the available resources and the current number of connections in the system to avoid server overload. In order to allow the differentiation of resumed SSL connections from new SSL connections we propose a possible extension of the Java Secure Sockets Extension (JSSE) API. Our evaluation on Tomcat server demonstrates the benefit of our proposal for preventing server overload. Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 3 |
| 2004 | Evaluating the Scalability of Java Event-Driven Web ServersabstractThe two major strategies used to construct high-performance Web servers are thread pools and event-driven architectures. The Java platform is commonly used in Web environments but up to the moment it did not provide any standard API to implement event-driven architectures efficiently. The new 1.4 release of the J2SE introduces the NIO (New I/O) API to help in the development of event-driven I/O intensive applications. We evaluate the scalability that this API provides to the Java platform in the field of Web servers, bringing together the majorly used commercial server (Apache) and one experimental server developed using the NIO API. We study the scalability of the NIO-based server as well as of its rival in a number of different scenarios, including uniprocessor, multiprocessor, bandwidth-bounded and CPU-bounded environments. The study concludes that the NIO API can be successfully used to create event-driven Java servers that can scale as well as the best of the commercial native-compiled Web server, at a fraction of its complexity and using only one or two worker threads. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 1 |