EDBT 2026 Demo / reviewers in the wild / expert
Jorda Polo
dblp:13/8505 · also Jordà Polo
· DBLP profile ↗
17ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0001-5422-7890ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Faster Kernels Do Not Mean Faster Applications in Exascale GPU Systems
Mariana Toledo Costa, Antigoni Georgiadou, James B. White, Woong Shin, Bruno Villasenor Alvarez, Jorda Polo, Karl W. Schulz, Philippe Olivier Alexandre Navaux, O. E. Bronson Messer, Arthur Francisco Lorenzon |
Euro-Par (1) | 6 |
| 2025 | Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNNabstractWe present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier. Massimiliano Lupo Pasini, Jong Choi 0001, Kshitij Mehta, David M. Rogers 0001, Jonghyun Bae, Khaled Z. Ibrahim, Ashwin M. Aji, Karl W. Schulz, Jorda Polo, Prasanna Balaprakash |
J. Supercomput. | 10 |
| 2023 | Performance characterization of video analytics workloads in heterogeneous edge infrastructuresabstractSummary Powered by deep learning, video analytic applications process millions of camera feeds in real‐time to extract meaningful information from their surroundings. And this number grows by the minute. To avoid saturating the backhaul network and provide lower latencies, a distributed and heterogeneous edge cloud is postulated as a key enabler for widespread video analytics. This article provides a complete characterization of end‐to‐end video analytics across a set of hardware platforms and different neural network architectures. Each platform is selected to fill a different gap in a distributed, shared, and heterogeneous infrastructure. Moreover, we analyze how performance scales on each of these platforms with respect to the amount of resources dedicated to video analytics. Finally, we extract the key conclusions of the characterization to build an experimental model to estimate performance and cost of end‐to‐end video analytics in different edge scenarios. Our experiments show that managing video analytics workloads efficiently requires awareness of both, the platforms in which these are executed, and the full end‐to‐end pipeline. To the best of our knowledge, this is the first work that provides a complete characterization of end‐to‐end video analytics in heterogeneous edge platforms. Daniel Rivas-Barragan, Francesc Guim 0001, Jorda Polo, David Carrera 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Towards automatic model specialization for edge video analytics
Daniel Rivas-Barragan, Francesc Guim 0001, Jorda Polo, Pubudu Madhawa Silva, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 3 |
| 2022 | Burst-Aware Predictive Autoscaling for Containerized MicroservicesabstractAutoscaling methods are used for cloud-hosted applications to dynamically scale the allocated resources for guaranteeing Quality-of-Service (QoS). The public-facing application serves dynamic workloads, which contain bursts and pose challenges for autoscaling methods to ensure application performance. Existing State-of-the-art autoscaling methods are burst-oblivious to determine and provision the appropriate resources. For dynamic workloads, it is hard to detect and handle bursts online for maintaining application performance. In this article, we propose a novel burst-aware autoscaling method which detects burst in dynamic workloads using workload forecasting, resource prediction, and scaling decision making while minimizing response time service-level objectives (SLO) violations. We evaluated our approach through a trace-driven simulation, using multiple synthetic and realistic bursty workloads for containerized microservices, improving performance when comparing against existing state-of-the-art autoscaling methods. Such experiments show an increase of$\times $1.09 in total processed requests, a reduction of$\times $5.17 for SLO violations, and an increase of$\times $0.767 cost as compared to the baseline method. Muhammad Abdullah 0004, Waheed Iqbal, Josep Lluís Berral, Jorda Polo, David Carrera 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Performance Evaluation of Data-Centric Workloads in Serverless EnvironmentsabstractServerless computing is a cloud-based execution paradigm that allows provisioning resources on-demand, freeing developers from infrastructure management and operational concerns. It typically involves deploying workloads as stateless functions that take no resources when not in use, and is meant to scale transparently. To make serverless effective, providers impose limits on a per-function level, such as maximum duration, fixed amount of memory, and no persistent local storage. These constraints make it challenging for data-intensive workloads to take advantage of serverless because they lead to sharing significant amounts of data through remote storage. In this paper, we build a performance model for serverless workloads that considers how data is shared between functions, including the amount of data and the underlying technology that is being used. The model's accuracy is assessed by running a real workload in a cluster using Knative, a state-of-the-art serverless environment, showing a relative error of 5.52%. With the proposed model, we evaluate the performance of data-intensive workloads in serverless, analyzing parallelism, scalability, resource requirements, and scheduling policies. We also explore possible solutions for the data-sharing problem, like using local memory and storage. Our results show that the performance of data-intensive workloads in serverless can be up to 4.32= faster depending on how these are deployed. Anna Maria Nestorov, Jorda Polo, Claudia Misale, David Carrera 0001, Alaa Youssef |
CLOUD | 2 |
| 2020 | A Hardware/Software Co-Design of K-mer Counting Using a CAPI-Enabled FPGAabstractAdvances in Next Generation Sequencing (NGS) technologies have caused the proliferation of genomic applications to detect DNA mutations and guide personalized medicine. These applications have an enormous computational cost due to the large amount of genomic data they process. Although leveraging FPGAs can improve the processing time of such amount of data, the limited memory capacity of FPGAs often restricts the potential gains. To overcome this limitation, IBM CAPI (Coherent Accelerator Processor Interface) supported platforms provide FPGAs with direct access to the CPU memory. This paper proposes a hardware/software co-design for k-mer counting, one of the most time-consuming phases of genomic applications. The proposed co-design targets CAPI-enabled FPGAs and is integrated into SMUFIN, a state-of-the-art reference-free method for finding DNA mutations. Results show that the proposed co-design outperforms the CPU-only design by a factor of 2.14×, it consumes 2.93× less energy, and it requires 1.57× less memory. Abbas Haghi, Lluc Alvarez, Jorda Polo, Dionysios Diamantopoulos, Christoph Hagleitner, Miquel Moretó |
FPL | 3 |
| 2020 | A highly parameterizable framework for Conditional Restricted Boltzmann Machine based workloads accelerated with FPGAs and OpenCLabstractConditional Restricted Boltzmann Machine (CRBM) is a promising candidate for a multidimensional system modeling that can learn a probability distribution over a set of data. It is a specific type of an artificial neural network with one input (visible) and one output (hidden) layer. Recently published works demonstrate that CRBM is a suitable mechanism for modeling multidimensional time series such as human motion, workload characterization, city traffic analysis. The process of learning and inference of these systems relies on linear algebra functions like matrix–matrix multiplication, and for higher data sets, they are very compute-intensive. In this paper, we present a configurable framework for CRBM based workloads for arbitrary large models. We show how to accelerate the learning process of CRBM with FPGAs and OpenCL, and we conduct an extensive scalability study for different model sizes and system configurations. We show significant improvement in performance/Watt for large models and batch sizes (from 1.51x up to 5.71x depending on the host configuration) when we use FPGA and OpenCL for the acceleration, and limited benefits for small models comparing to the state-of-the-art CPU solution. Zoran Jaksic, Nicola Cadenelli, David Buchaca Prats, Jorda Polo, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 4 |
| 2019 | Considerations in using OpenCL on GPUs and FPGAs for throughput-oriented genomics workloadsabstractThe recent upsurge in the available amount of health data and the advances in next-generation sequencing are setting the ground for the long-awaited precision medicine. To process this deluge of data, bioinformatics workloads are becoming more complex and more computationally demanding. For this reasons they have been extended to support different computing architectures, such as GPUs and FPGAs, to leverage the form of parallelism typical of each of such architectures. The paper describes how a genomic workload such as k-mer frequency counting that takes advantage of a GPU can be offloaded to one or even more FPGAs. Moreover, it performs a comprehensive analysis of the FPGA acceleration comparing its performance to a non-accelerated configuration and when using a GPU. Lastly, the paper focuses on how, when using accelerators with a throughput-oriented workload, one should also take into consideration both kernel execution time and how well each accelerator board overlaps kernels and PCIe transferred. Results show that acceleration with two FPGAs can improve both time- and energy-to-solution for the entire accelerated part by a factor of 1.32x. Per contra, acceleration with one GPU delivers an improvement of 1.77x in time-to-solution but of a lower 1.49x in energy-to-solution due to persistently higher power consumption. The paper also evaluates how future FPGA boards with components (i.e., off-chip memory and PCIe) on par with those of the GPU board could provide an energy-efficient alternative to GPUs. Nicola Cadenelli, Zoran Jaksic, Jorda Polo, David Carrera 0001 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Topology-aware GPU scheduling for learning workloads in cloud environmentsabstractRecent advances in hardware, such as systems with multiple GPUs and their availability in the cloud, are enabling deep learning in various domains including health care, autonomous vehicles, and Internet of Things. Multi-GPU systems exhibit complex connectivity among GPUs and between GPUs and CPUs. Workload schedulers must consider hardware topology and workload communication requirements in order to allocate CPU and GPU resources for optimal execution time and improved utilization in shared cloud environments. Marcelo Amaral, Jorda Polo, David Carrera 0001, Seetharami R. Seelam, Malgorzata Steinder |
SC | 2 |
| 2015 | Performance Evaluation of Microservices Architectures Using ContainersabstractMicro services architecture has started a new trend for application development for a number of reasons: (1) to reduce complexity by using tiny services, (2) to scale, remove and deploy parts of the system easily, (3) to improve flexibility to use different frameworks and tools, (4) to increase the overall scalability, and (5) to improve the resilience of the system. Containers have empowered the usage of micro services architectures by being lightweight, providing fast start-up times, and having a low overhead. Containers can be used to develop applications based on monolithic architectures where the whole system runs inside a single container or inside a micro services architecture where one or few processes run inside the containers. Two models can be used to implement a micro services architecture using containers: master-slave, or nested-container. The goal of this work is to compare the performance of CPU and network running benchmarks in the two aforementioned models of micro services architecture hence provide a benchmark analysis guidance for system designers. Marcelo Amaral, Jorda Polo, David Carrera 0001, Iqbal Mohomed, Merve Unuvar, Malgorzata Steinder |
NCA | 2 |
| 2014 | Adaptive MapReduce Scheduling in Shared EnvironmentsabstractIn this paper we present a MapReduce task scheduler for shared environments in which MapReduce is executed along with other resource-consuming workloads, such as transactional applications. All workloads may potentially share the same data store, some of them consuming data for analytics purposes while others acting as data generators. This kind of scenario is becoming increasingly important in data centers where improved resource utilization can be achieved through workload consolidation, and is specially challenging due to the interaction between workloads of different nature that compete for limited resources. The proposed scheduler aims to improve resource utilization across machines while observing completion time goals. Unlike other MapReduce schedulers, our approach also takes into account the resource demands for non-MapReduce workloads, and assumes that the amount of resources made available to the MapReduce applications is variable over time. As shown in our experiments, our proposal improves the management of MapReduce jobs in the presence of variable resource availability, increasing the accuracy of the estimations made by the scheduler, thus improving completion time goals without an impact on the fairness of the scheduler. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Malgorzata Steinder |
CCGRID | 1 |
| 2013 | Enabling Distributed Key-Value Stores with Low Latency-Impact Snapshot SupportabstractCurrent distributed key-value stores generally provide greater scalability at the expense of weaker consistency and isolation. However, additional isolation support is becoming increasingly important in the environments in which these stores are deployed, where different kinds of applications with different needs are executed, from transactional workloads to data analytics. While fully-fledged ACID support may not be feasible, it is still possible to take advantage of the design of these data stores, which often include the notion of multiversion concurrency control, to enable them with additional features at a much lower performance cost and maintaining its scalability and availability. In this paper we explore the effects that additional consistency guarantees and isolation capabilities may have on a state of the art key-value store: Apache Cassandra. We propose and implement a new multiversioned isolation level that provides stronger guarantees without compromising Cassandra's scalability and availability. As shown in our experiments, our version of Cassandra allows Snapshot Isolation-like transactions, preserving the overall performance and scalability of the system. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Mike Spreitzer, Malgorzata Steinder |
NCA | 1 |
| 2013 | Deadline-Based MapReduce Workload ManagementabstractThis paper presents a scheduling technique for multi-job MapReduce workloads that is able to dynamically build performance models of the executing workloads, and then use these models for scheduling purposes. This ability is leveraged to adaptively manage workload performance while observing and taking advantage of the particulars of the execution environment of modern data analytics applications, such as hardware heterogeneity and distributed storage. The technique targets a highly dynamic environment in which new jobs can be submitted at any time, and in which MapReduce workloads share physical resources with other workloads. Thus the actual amount of resources available for applications can vary over time. Beyond the formulation of the problem and the description of the algorithm and technique, a working prototype (called Adaptive Scheduler) has been implemented. Using the prototype and medium-sized clusters (of the order of tens of nodes), the following aspects have been studied separately: the scheduler's ability to meet high-level performance goals guided only by user-defined completion time goals; the scheduler's ability to favor data-locality in the scheduling algorithm; and the scheduler's ability to deal with hardware heterogeneity, which introduces hardware affinity and relative performance characterization for those applications that can benefit from executing on specialized processors. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2011 | Resource-Aware Adaptive Scheduling for MapReduce Clusters
Jorda Polo, Claris Castillo, David Carrera 0001, Yolanda Becerra 0001, Ian Whalley, Malgorzata Steinder, Jordi Torres, Eduard Ayguadé |
Middleware | 1 |
| 2010 | Performance Management of Accelerated MapReduce Workloads in Heterogeneous ClustersabstractNext generation data centers will be composed of thousands of hybrid systems in an attempt to increase overall cluster performance and to minimize energy consumption. New programming models, such as MapReduce, specifically designed to make the most of very large infrastructures will be leveraged to develop massively distributed services. At the same time, data centers will bring an unprecedented degree of workload consolidation, hosting in the same infrastructure distributed services from many different users. In this paper we present our advancements in leveraging the Adaptive MapReduce Scheduler to meet user defined high level performance goals while transparently and efficiently exploiting the capabilities of hybrid systems. While the Adaptive Scheduler was already able to dynamically allocate resources to co-located MapReduce jobs based on their completion time goals, it was completely unaware of specific hardware capabilities. In our work we describe the changes introduced in the Adaptive Scheduler to enable it with hardware awareness and with the ability to co-schedule accelerable and non-accelerable jobs on the same heterogeneous MapReduce cluster, making the most of the underlying hybrid systems. The developed prototype is tested in a cluster of Cell/BE blades and relies on the use of accelerated and non-accelerated versions of the MapReduce tasks of different deployed applications to dynamically select the best version to run on each node. Decisions are made after workload composition and jobs' completion time goals. Results show that the augmented Adaptive Scheduler provides dynamic resource allocation across jobs, hardware affinity when possible, and is even able to spread jobs' tasks across accelerated and non-accelerated nodes in order to meet performance goals in extreme conditions. To our knowledge this is the first MapReduce scheduler and prototype that is able to manage high-level performance goals even in presence of hybrid systems and accelerable jobs. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 1 |
| 2010 | Performance-driven task co-scheduling for MapReduce environmentsabstractMapReduce is a data-driven programming model proposed by Google in 2004 which is especially well suited for distributed data analytics applications. We consider the management of MapReduce applications in an environment where multiple applications share the same physical resources. Such sharing is in line with recent trends in data center management which aim to consolidate workloads in order to achieve cost and energy savings. In a shared environment, it is necessary to predict and manage the performance of workloads given a set of performance goals defined for them. In this paper, we address this problem by introducing a new task scheduler for a MapReduce framework that allows performance-driven management of MapReduce tasks. The proposed task scheduler dynamically predicts the performance of concurrent MapReduce jobs and adjusts the resource allocation for the jobs. It allows applications to meet their performance objectives without over-provisioning of physical resources. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Malgorzata Steinder, Ian Whalley |
NOMS | 1 |