VLDB 2026 Research / reviewers in the wild / expert
Josep Lluís Berral
dblp:56/5336 · also Josep Lluis Berral, Josep Lluis Berral-Garcia
· DBLP profile ↗
38ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-3037-3580ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 4 since 2021Computer networks · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloudSkin: AI-Based Learning Plane for Autonomic Management in the Cloud-Edge Continuum
Peini Liu, Joan Oliveras Torra, Marc Palacín, Ramon Nou, Josep Lluís Berral, Jordi Guitart |
COMPSAC | 5 |
| 2026 | FRIDA: Free-rider detection using privacy attacksabstractFederated learning is increasingly popular as it enables multiple parties with limited datasets and resources to train a machine learning model collaboratively. However, similar to other collaborative systems, federated learning is vulnerable to free-riders — participants who benefit from the global model without contributing. Free-riders compromise the integrity of the learning process and slow down the convergence of the global model, resulting in increased costs for honest participants. To address this challenge, we propose FRIDA: f ree- ri der d etection using privacy a ttacks. Instead of focusing on implicit effects of free-riding, FRIDA utilizes membership and property inference attacks to directly infer evidence of genuine client training. Our extensive evaluation demonstrates that FRIDA is effective across a wide range of scenarios. Pol G. Recasens, Ádám Horváth, Alberto Gutierrez-Torre, Jordi Torres, Josep Lluís Berral, Balazs Pejo |
J. Inf. Secur. Appl. | 5 |
| 2025 | Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM InferenceabstractLarge language models have been widely adopted across different tasks, but their auto-regressive generation nature often leads to inefficient resource utilization during inference. While batching is commonly used to increase throughput, per-formance gains plateau beyond a certain batch size, especially with smaller models, a phenomenon that existing literature typically explains as a shift to the compute-bound regime. In this paper, through an in-depth GPU-level analysis, we reveal that large-batch inference remains memory-bound, with most GPU compute capabilities underutilized due to DRAM bandwidth saturation as the primary bottleneck. To address this, we propose a Batching Configuration Advisor (BCA) that optimizes memory allocation, reducing GPU memory requirements with minimal impact on throughput. The freed memory and underutilized GPU compute capabilities can then be leveraged by concurrent workloads. Specifically, we use model replication to improve serving throughput and GPU utilization. Our findings challenge conventional assumptions about LLM inference, offering new in-sights and practical strategies for improving resource utilization, particularly for smaller language models. Pol G. Recasens, Ferran Agullo, Chen Wang 0039, Olivier Tardieu, Jordi Torres, Josep Lluís Berral |
CLOUD | 8 |
| 2025 | Integrating Reliability into Intent-Driven Orchestration on Kubernetes-based Edge NodesabstractPlacing workloads across the Edge and Cloud can be performed through containerization. However, conditions in Edge nodes are very different from Cloud, as environmental factors such as temperature, humidity, voltage provisioning or dust, heavily affect execution performance and device health. Reliability is an important factor when deciding to deploy such load onto Edge devices, indicating their capability to achieve a desired Quality of Service (QoS). For this, research on Edge computing must focus on how to monitor, estimate and manage devices, in a distributed, autonomous and reliable manner. Our current efforts on performance analysis for devices under "wild conditions" are moving towards integrating reliability into orchestrator systems. Here we present our vision and roadmap for expanding Cloud orchestration towards the Edge with technologies capable of providing knowledge about node environmental conditions and mitigate its impact. Through Out-Of-Band telemetry, we can retrieve temperature and power consumption variables from node components, indicating its health and estimating its reliability given external stress factors. In particular, using intent-based orchestration for containerized platforms, such factors can be used for enforcing reliability as a key-performance indicator. The current work in progress focuses on industrial and commercial scenarios, e.g., Edge computing for urban mobility, with road-side nodes performing AI-based Video-Analytics, exposed to uncontrolled weather conditions. The principal objective is to achieve an Edge network orchestration that takes into account node health in an automatic manner for reducing operational costs such as energy consumption, device repair and replacement, while maintaining QoS in the Edge. Josep Lluís Berral, David Aguilera-Luzón, Peini Liu, Ramon Nou, Maria A. Serrano, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Mario José Diván |
ICNP | 1 |
| 2025 | Mobility Usecase: Intelligent Service Migration in Cloud-Edge ContinuumabstractAs industries increasingly embrace digital and intelligent transformation, enterprises face significant challenges in the containerization upgrades for their Artificial Intelligence(AI) applications and dynamic service migration and management in Cloud-Edge continuum(CEC). This paper presents CloudSkin, an innovative platform designed to realize streamlined, seamless and intelligent service migration in Cloud-Edge Continuum by integrating advanced containerization techniques and AI-driven orchestration capabilities. Our approach, leveraging intelligent algorithms for service migration, can seamlessly transit services between cloud and edge environments, ensuring optimised resource allocation and reducing service latency to assure quality of service(QoS). CloudSkin has been enabled in a Mobility Usecase, empowering Cellnex businesses undergoing digital transformation to achieve higher operational efficiency. The experimental results show that compared to traditional reactive service migration, using intelligent proactive service migration can provide better migration detection, improving F1-Score up to 23.5%, and reducing 28.9% the service running time where the service latency violates SLA. Peini Liu, Joan Oliveras Torra, Marc Palacín, Michail Dalgitsis, Maria A. Serrano, Eftychia G. Datsika, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 10 |
| 2025 | Enhancing the output of time series forecasting algorithms for cloud resource provisioningabstractForecasting the resource consumption of workloads is a frequent approach in the cloud provisioning field. Ideally, such predictions allow obtaining a more accurate scheduling and management of resources in a computing cluster. However, the current approaches fail to properly forecast the future consumption in areas where sudden increases of consumption are present, i.e ., spikes. Even, commonly employed metrics lack the ability to properly evaluate sharp behaviours in the traces. This may generate resource starvation problems in the running workloads and decreases the Quality of Service (QoS) provided to external users. To address this issue, we propose two strategies that modify the outputs of forecasting algorithms without changing the algorithms’ internals. The new outputs considerably enhance the prediction of sudden increases, duplicating the F1 score metric in average for all tested algorithms. This improvement in the handling of spikes comes with an increased over-provision of resources. Nevertheless, the proposed strategies give the user an easy way to control this trade-off between predicting spikes and the amount of over-provision. The user can decide which is the right balance that better fits the requirements of its specific scenario. Furthermore, we propose a new evaluation methodology that better assesses the behaviour of forecasting algorithms in cloud traces, especially focused on the performance around increases of consumption, and we give insights on the reasons behind the predictions of the algorithms with the application of explainability techniques. The code repository of this work can be accessed through GitHub at this link https://github.com/FerranAgulloLopez/ResourceForecasting . • Tackling the forecasting of workload resource consumption for cloud provisioning. • The forecasts can improve the sharing of resources between co-allocated workloads. • The current approaches lie far behind when predicting increases of consumption. • The work proposes new strategies and a new evaluation to enhance the forecasts. • Explainability techniques are used to understand the predictions in cloud time series. Ferran Agullo, Alberto Gutierrez-Torre, Jordi Torres, Josep Lluís Berral |
Future Gener. Comput. Syst. | 4 |
| 2024 | On the Relation Between Open Project-Based Learning in Undergraduate Computer Science Education and Contemporary Technological TrendsabstractAmidst rapid and constant technological change, keeping IT higher education curricula up to date is becoming increasingly challenging. For a decade, the course”Project on Information Technologies” within the undergraduate computer science studies offered by a major technical university has been pursuing an effort of continuous curriculum adaptation and active learning practices based on an open project-based learning (PBL) methodology. This article researches the implications of employing an open-ended PBL approach, with a specific focus on the alignment of acquired skills with the trends in the professional landscape. Our analysis has identified a strong correlation between the technologies utilized in the projects and the contemporary technological trends in the areas of global technological focus, programming languages, server-side technologies, database management systems, and DevOps-related tools. For the study, we analyzed empirical data gathered from over 100 projects involving more than 400 students who enrolled in this reference course, belonging to the last year of the IT higher education programme in the last 10 years. The results suggest the open project-based design as a teaching means in the student’s study plan for fostering the student the learning of prevailing current practical technologies. Rubén Tous, Felix Freitag, Josep Lluís Berral |
CSEDU (2) | 3 |
| 2024 | SECURED for Health: Scaling Up Privacy to Enable the Integration of the European Health Data SpaceabstractIn this paper, we present the SECURED project11Funded in part by the European Union (EU), Grant Agreement no. 10109571. Views and opinions expressed are those of the authors and do not necessarily reflect those of the EU or the Health and Digital Executive Agency. Neither the EU nor the granting authority are responsible for them., aimed at improving privacy-preserving processing of data in the health domain. The technologies developed in the project will be demonstrated in four health-related use cases and with the involvement of SME's selected through an open funding call. Francesco Regazzoni 0001, Gergely Ács, Albert Zoltan Aszalos, Christos Avgerinos, Nikolaos Bakalos, Josep Lluís Berral, Joppe W. Bos, Marco Brohet, Andrés G. Castillo, Gareth T. Davies, Stefanos Florescu, Pierre-Elisée Flory, Alberto Gutierrez-Torre, Evangelos Haleplidis, Alice Héliou, Sotiris Ioannidis, Alexander El-Kady, Katarzyna Kapusta, Konstantina Karagianni, Pieter Kruizinga, Kyrian Maat, Zoltán Ádám Mann, Kalliopi Mastoraki, SeoJeong Moon, Maja Nisevic, Balazs Pejo, Kostas Papagiannopoulos, Vassilis Paliouras, Paolo Palmieri 0001, Francesca Palumbo, Juan Carlos Pérez Baun, Péter Pollner, Eduard Porta-Pardo, Luca Pulina, Muhammad Ali Siddiqi, Daniela Spajic, Christos Strydis, George Tasopoulos, Vincent Thouvenot, Christos Tselios, Apostolos P. Fournaris |
DATE | 6 |
| 2024 | HealthMesh: An Architectural Framework for Federated Healthcare Data Management
Aniol Bisquert, Achraf Hmimou, Josep Lluís Berral, Alberto Gutierrez-Torre, Oscar Romero 0001 |
DOLAP | 3 |
| 2024 | Data-Connector: An Agent-Based Framework for Autonomous ML-Based Smart Management in Cloud-Edge ContinuumabstractMachine Learning (ML) is becoming pervasive and integrated into different kinds of intelligent applications, and the collaborative Cloud-Edge continuum has been introduced as an emerging trend to support their adoption into use cases. However, managing these ML applications in the CloudEdge continuum is challenging due to the ML application's dynamic resource usage with different user loads and Cloud and Edge's dynamic resource availability. We envision machine learning methods that can be used for smart management in this dynamic environment, but how to deploy and utilize them for the adaptation scenario in Cloud-Edge continuum is unknown. This paper proposes an agent-based framework to enable autonomous smart management mechanisms that can be broadly enabled in diverse adaptation scenarios. The agent acts as a data-connector11https://github.com/bsc-scanflow/data-connector, connecting different sources of data, utilizing ML models for decision-making and triggering adaptations in Cloud-edge platforms. The case study shows the feasibility of our proposed data-connector for smart migration of an ML workload in the Cloud-edge continuum. The result shows that the smart migration-enabled Cloud-edge scenario has 11.9% ML application prediction time better than the Cloud scenario without migration. Moreover, with minimal customization, the data connector agent can be adapted for more use cases. Peini Liu, Joan Oliveras Torra, Marc Palacín, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 5 |
| 2024 | Dexter: A Performance-Cost Efficient Resource Allocation Manager for Serverless Data AnalyticsabstractLeveraging serverless platforms for the efficient execution of distributed data analytics frameworks, such as Apache Spark [3], has gained substantial interest since early 2022. The elasticity, free-of-management, and on-demand scalability of serverless have motivated the effort in deploying distributed data analytics applications to serverless platforms. However, effectively auto-scaling resources for such complex workloads so that we can fully benefit from the resource elasticity of serverless remains challenging. Mis-configuration can result in severe performance and cost issues arising from resource under- and over-provisioning. Anna Maria Nestorov, Diego Marron, Alberto Gutierrez-Torre, Chen Wang 0039, Claudia Misale, Alaa Youssef, David Carrera 0001, Josep Lluís Berral |
Middleware | 8 |
| 2024 | Time-Quality Tradeoff of MuseHash Query Processing Performance
Maria Pegia, Ferran Agullo, Anastasia Moumtzidou, Alberto Gutierrez-Torre, Björn Þór Jónsson 0001, Josep Lluís Berral, Ilias Gialampoukidis, Stefanos Vrochidis, Ioannis Kompatsiaris |
MMM (3) | 6 |
| 2023 | MScheduler: Leveraging Spot Instances for High-Performance Reservoir Simulation in the CloudabstractPetroleum reservoir simulation uses computer models to predict fluid flow in porous media, aiding to forecast oil production. Engineers execute numerous simulations with different geological realizations to refine the accuracy of the model. These experiments require considerable computational resources, which are not always available within the on-premises infrastructure. Commercial public cloud platforms can offer many advantages, such as virtually unlimited scalability and pay-per-use pricing. This paper introduces MScheduler, a meta scheduler framework for reservoir simulations at Petrobras, a Brazilian energy company. It efficiently executes jobs in the cloud, utilizing spot Virtual Machines (VMs) to reduce costs and ensure job completion even with VM termination. Contributions include a novel methodology for reservoir simulation checkpointing, a cost-based scheduler, and an analysis of the strategy using real production jobs from Petrobras. Felipe Albuquerque Portella, Paulo J. B. Estrela, Renzo Q. Malini, Luan Teylo, Josep Lluís Berral, Lúcia M. A. Drummond |
CloudCom | 5 |
| 2023 | VITAMIN-V: Virtual Environment and Tool-Boxing for Trustworthy Development of RISC-V Based Cloud ServicesabstractVITAMIN-V is a 2023–2025 Horizon Europe project that aims to develop a complete RISC-V open-source software stack for cloud services with comparable performance to the cloud-dominant x86 counterpart and a powerful virtual execution environment for software development, validation, verification, and testing that considers the relevant RISC-VISA extensions for cloud deployment. VITAMIN-V will specifically support the RISC-V extensions for virtualization, cryptography, and vec-torization in three virtual environments: QEMU, gem5, and cloud FPGA prototype platforms. The project will focus on European Processor Initiative (EPI) based RISC-V designs and accelerators. VITAMIN-V will also support the ISA extensions by adding the compiler and toolchain support. Furthermore, it will develop novel software validation, verification, and testing approaches to ensure software trustworthiness. To enable the execution of complete cloud stacks, VITAMIN-V will port all necessary machine-dependent modules in relevant open-source cloud software distributions, focusing on three cloud setups. Finally, VITAMIN-V will demonstrate and benchmark these three cloud setups using relevant AI, big-data, and serverless applications. VITAMIN-V aims to match the software performance of its x86 equivalent while contributing to RISC-V open-source virtual environments, software validation, and cloud software suites. Ramon Canal, Cristiano Pegoraro Chenet, Aggelos Arelakis, José-María Arnau, Josep Lluís Berral, Aaron Call, Stefano Di Carlo, Juan José Costa, Dimitris Gizopoulos, Vasileios Karakostas, Francesco Lubrano, Konstantinos Nikas, Yiannis Nikolakopoulos, Beatriz Otero, George Papadimitriou 0001, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos, Daniel Raho, Alvise Rigo, Eva Rodríguez, Alessandro Savino 0001, Alberto Scionti, Nikolaos Tampouratzis, Alex Torregrosa |
DSD | 5 |
| 2022 | Towards automatic model specialization for edge video analytics
Daniel Rivas-Barragan, Francesc Guim 0001, Jorda Polo, Pubudu Madhawa Silva, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 5 |
| 2022 | Automatic Distributed Deep Learning Using Resource-Constrained Edge DevicesabstractProcessing data generated at high volume and speed from the Internet of Things, smart cities, domotic, intelligent surveillance, and e-healthcare systems require efficient data processing and analytics services at the Edge to reduce the latency and response time of the applications. The fog computing edge infrastructure consists of devices with limited computing, memory, and bandwidth resources, which challenge the construction of predictive analytics solutions that require resource-intensive tasks for training machine learning models. In this work, we focus on the development of predictive analytics for urban traffic. Our solution is based on deep learning techniques localized in the Edge, where computing devices have very limited computational resources. We present an innovative method for efficiently training the gated recurrent-units (GRUs) across available resource-constrained CPU and GPU Edge devices. Our solution employs distributed GRU model learning and dynamically stops the training process to utilize the low-power and resource-constrained Edge devices while ensuring good estimation accuracy effectively. The proposed solution was extensively evaluated using low-powered ARM-based devices, including Raspberry Pi v3 and the low-powered GPU-enabled device NVIDIA Jetson Nano, and also compared them with Single-CPU Intel Xeon machines. For the evaluation experiments, we used real-world Floating Car Data. The experiments show that the proposed solution delivers excellent prediction accuracy and computational performance on the Edge when compared to the baseline methods. Alberto Gutierrez-Torre, Kiyana Bahadori, Shuja-ur-Rehman Baig, Waheed Iqbal, Tullio Vardanega, Josep Lluís Berral, David Carrera 0001 |
IEEE Internet Things J. | 6 |
| 2022 | Burst-Aware Predictive Autoscaling for Containerized MicroservicesabstractAutoscaling methods are used for cloud-hosted applications to dynamically scale the allocated resources for guaranteeing Quality-of-Service (QoS). The public-facing application serves dynamic workloads, which contain bursts and pose challenges for autoscaling methods to ensure application performance. Existing State-of-the-art autoscaling methods are burst-oblivious to determine and provision the appropriate resources. For dynamic workloads, it is hard to detect and handle bursts online for maintaining application performance. In this article, we propose a novel burst-aware autoscaling method which detects burst in dynamic workloads using workload forecasting, resource prediction, and scaling decision making while minimizing response time service-level objectives (SLO) violations. We evaluated our approach through a trace-driven simulation, using multiple synthetic and realistic bursty workloads for containerized microservices, improving performance when comparing against existing state-of-the-art autoscaling methods. Such experiments show an increase of$\times $1.09 in total processed requests, a reduction of$\times $5.17 for SLO violations, and an increase of$\times $0.767 cost as compared to the baseline method. Muhammad Abdullah 0004, Waheed Iqbal, Josep Lluís Berral, Jorda Polo, David Carrera 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Theta-Scan: Leveraging Behavior-Driven Forecasting for Vertical Auto-Scaling in Container CloudabstractDetection of behavior patterns on resource usage in containerized Cloud applications is necessary for proper resource provisioning. Applications can use CPU/Memory with repetitive patterns, following a trend over time independently. By identifying such patterns, resource forecasting models can be fit better, reducing over/under-provisioning via fewer resizing operations. Here we present ThetaScan, a time-series analysis method for vertical auto-scaling of containers in the Cloud, based on the detection of stationarity/trending and periodicity on resource consumption. Our method leverages the Theta Forecaster algorithm with deseasonalization that, in our provisioning scenario, only requires the estimated periodicity for resource consumption as principal hyper-parameter. Commonly used behavior detection methods require manual hyper-parameter tuning, making them infeasible for automation. Besides, it can be used at multi-scales (minute/hour/day), detecting hourly and daily patterns to improve resource usage prediction. Experiments show that we can detect behaviors in resource consumption that common methods miss, without requiring extensive manual tuning. We can reduce the resizing triggers compared to fixed-size scheduling around ~ 10% – 15%, reduce over-provisioning of CPU and Memory through periodic-based provisioning. Also a ~ 60% on multiscale resource forecasting for traces showing periodicity at different levels in respect to single-scale. Josep Lluís Berral, David Buchaca Prats, Claudia Herron, Chen Wang 0039, Alaa Youssef |
CLOUD | 1 |
| 2020 | Proactive Container Auto-scaling for Cloud Native Machine Learning ServicesabstractUnderstanding the resource usage behaviors of the ever-increasing machine learning workloads are critical to cloud providers offering Machine Learning (ML) services. Capable of auto-scaling resources for customer workloads can significantly improve resource utilization, thus greatly reducing the cost. Here we leverage the AI4DL framework [1] to characterize workload and discover resource consumption phases. We advance the existing technology to an incremental phase discovery method that applies to more general types of ML workload for both training and inference. We use a time-window MultiLayer Perceptron (MLP) to predict phases in containers with different types of workload. Then, we propose a predictive vertical auto-scaling policy to resize the container dynamically according to phase predictions. We evaluate our predictive auto-scaling policies on 561 long-running containers with multiple types of ML workloads. The predictive policy can reduce up to 38% of allocated CPU compared to the default resource provisioning policies by developers. By comparing our predictive policies with commonly used reactive auto-scaling policies, we find that they can accurately predict sudden phase transitions (with an F1-score of 0.92) and significantly reduce the number of out-of-memory errors (350 vs. 20). Besides, we show that the predictive auto-scaling policy maintains the number of resizing operations close to the best reactive policies. David Buchaca Prats, Josep Lluís Berral, Chen Wang 0039, Alaa Youssef |
CLOUD | 2 |
| 2020 | Improving maritime traffic emission estimations on missing data with CRBMs
Alberto Gutierrez-Torre, Josep Lluís Berral, David Buchaca Prats, Marc Guevara, Albert Soret, David Carrera 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Adaptive sliding windows for improved estimation of data center resource utilizationabstractAccurate prediction of data center resource utilization is required for capacity planning, job scheduling, energy saving, workload placement, and load balancing to utilize the resources efficiently. However, accurately predicting those resources is challenging due to dynamic workloads, heterogeneous infrastructures, and multi-tenant co-hosted applications. Existing prediction methods use fixed size observation windows which cannot produce accurate results because of not being adaptively adjusted to capture local trends in the most recent data. Therefore, those methods train on large fixed sliding windows using an irrelevant large number of observations yielding to inaccurate estimations or fall for inaccuracy due to degradation of estimations with short windows on quick changing trends. In this paper we propose a deep learning-based adaptive window size selection method, dynamically limiting the sliding window size to capture the trend for the latest resource utilization, then build an estimation model for each trend period. We evaluate the proposed method against multiple baseline and state-of-the-art methods, using real data-center workload data sets. The experimental evaluation shows that the proposed solution outperforms those state-of-the-art approaches and yields 16 to 54% improved prediction accuracy compared to the baseline methods. Shuja-ur-Rehman Baig, Waheed Iqbal, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 3 |
| 2020 | A highly parameterizable framework for Conditional Restricted Boltzmann Machine based workloads accelerated with FPGAs and OpenCLabstractConditional Restricted Boltzmann Machine (CRBM) is a promising candidate for a multidimensional system modeling that can learn a probability distribution over a set of data. It is a specific type of an artificial neural network with one input (visible) and one output (hidden) layer. Recently published works demonstrate that CRBM is a suitable mechanism for modeling multidimensional time series such as human motion, workload characterization, city traffic analysis. The process of learning and inference of these systems relies on linear algebra functions like matrix–matrix multiplication, and for higher data sets, they are very compute-intensive. In this paper, we present a configurable framework for CRBM based workloads for arbitrary large models. We show how to accelerate the learning process of CRBM with FPGAs and OpenCL, and we conduct an extensive scalability study for different model sizes and system configurations. We show significant improvement in performance/Watt for large models and batch sizes (from 1.51x up to 5.71x depending on the host configuration) when we use FPGA and OpenCL for the acceleration, and limited benefits for small models comparing to the state-of-the-art CPU solution. Zoran Jaksic, Nicola Cadenelli, David Buchaca Prats, Jorda Polo, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 5 |
| 2020 | Sequence-to-sequence models for workload interference prediction on batch processing datacenters
David Buchaca Prats, Joan Marcual, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 3 |
| 2020 | You Only Run Once: Spark Auto-Tuning From a Single RunabstractTuning configurations of Spark jobs is not a trivial task. State-of-the-art auto-tuning systems are based on iteratively running workloads with different configurations. During the optimization process, the relevant features are explored to find good solutions. Many optimizers enhance the time-to-solution using black-box optimization algorithms that do not take into account any information from the Spark workloads. In this article, we present a new method for tuning configurations that uses information from one run of a Spark workload. To achieve good performance, we mine the SparkEventLog that is generated by the Spark engine. This log file contains a large amount of information from the executed application. We use this information to enhance a performance model with low-level features from the workload to be optimized. These features include Spark Actions, Transformations, and Task metrics. This process allows us to obtain application-specific workload information. With this information our system can predict sensible Spark configurations for unseen jobs, given that it has been trained with reasonable coverage of Spark applications. Experiments show that the presented system correctly produces good configurations, while achieving up to 80% speedup with respect to the default Spark configuration, and up to 12x speedup of the time-to-solution with respect to a standard Bayesian Optimization procedure. David Buchaca Prats, Felipe Albuquerque Portella, Carlos H. A. Costa, Josep Lluís Berral |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2019 | Adaptive Prediction Models for Data Center Resources Utilization EstimationabstractAccurate estimation of data center resource utilization is a challenging task due to multi-tenant co-hosted applications having dynamic and time-varying workloads. Accurate estimation of future resources utilization helps in better job scheduling, workload placement, capacity planning, proactive auto-scaling, and load balancing. The inaccurate estimation leads to either under or over-provisioning of data center resources. Most existing estimation methods are based on a single model that often does not appropriately estimate different workload scenarios. To address these problems, we propose a novel method to adaptively and automatically identify the most appropriate model to accurately estimate data center resources utilization. The proposed approach trains a classifier based on statistical features of historical resources usage to decide the appropriate prediction model to use for given resource utilization observations collected during a specific time interval. We evaluated our approach on real datasets and compared the results with multiple baseline methods. The experimental evaluation shows that the proposed approach outperforms the state-of-the-art approaches and delivers 6% to 27% improved resource utilization estimation accuracy compared to baseline methods. Shuja-ur-Rehman Baig, Waheed Iqbal, Josep Lluís Berral, Abdelkarim Erradi, David Carrera 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Constant-Time Sliding Window Framework with Reduced Memory Footprint and Efficient Bulk EvictionsabstractThe fast evolution of data analytics platforms has resulted in an increasing demand for real-time data stream processing. From Internet of Things applications to the monitoring of telemetry generated in large data centers, a common demand for currently emerging scenarios is the need to process vast amounts of data with low latencies, generally performing the analysis process as close to the data source as possible. Stream processing platforms are required to be malleable and absorb spikes generated by fluctuations of data generation rates. Data is usually produced as time series that have to be aggregated using multiple operators, being sliding windows one of the most common abstractions used to process data in real-time. To satisfy the above-mentioned demands, efficient stream processing techniques that aggregate data with minimal computational cost need to be developed. In this paper we present the Monoid Tree Aggregator general sliding window aggregation framework, which seamlessly combines the following features: amortized$O(1)$time complexity and a worst-case of$O(\log {n})$between insertions; it provides both a window aggregation mechanism and a window slide policy that are user programmable; the enforcement of the window sliding policy exhibits amortized$O(1)$computational cost for single evictions and supports bulk evictions with cost$O(\log {n})$; and it requires a local memory space of$O(\log {n})$. The framework can compute aggregations over multiple data dimensions, and has been designed to support decoupling computation and data storage through the use of distributedKey-Value Storesto keep window elements and partial aggregations. Álvaro Villalba, Josep Lluís Berral, David Carrera 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | A resilient and distributed near real-time traffic forecasting application for Fog computing environmentsabstractIn this paper we propose an architecture for a city-wide traffic modeling and prediction service based on the Fog Computing paradigm. The work assumes an scenario in which a number of distributed antennas receive data generated by vehicles across the city. In the Fog nodes data is collected, processed in local and intermediate nodes, and finally forwarded to a central Cloud location for further analysis. We propose a combination of a data distribution algorithm, resilient to back-haul connectivity issues, and a traffic modeling approach based on deep learning techniques to provide distributed traffic forecasting capabilities. In our experiments, we leverage real traffic logs from one week of Floating Car Data (FCD) generated in the city of Barcelona by a road-assistance service fleet comprising thousands of vehicles. FCD was processed across several simulated conditions, ranging from scenarios in which no connectivity failures occurred in the Fog nodes, to situations with long and frequent connectivity outage periods. For each scenario, the resilience and accuracy of both the data distribution algorithm, and the learning methods were analyzed. Results show that the data distribution process running in the Fog nodes is resilient to back-haul connectivity issues and is able to deliver data to the Cloud location even in presence of severe connectivity problems. Additionally, the proposed traffic modeling and forecasting method exhibits better behavior when run distributed in the Fog instead of centralized in the Cloud, especially when connectivity issues occur that force data to be delivered out of order to the Cloud. Juan Luis Pérez 0003, Alberto Gutierrez-Torre, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | Automatic Generation of Workload Profiles Using Unsupervised Learning PipelinesabstractThe complexity of resource usage and power consumption on cloud-based applications makes the understanding of application behavior through expert examination difficult. The difficulty increases when applications are seen as “black boxes,” where only external monitoring can be retrieved. Furthermore, given the different amount of scenarios and applications, automation is required. Here, we examine and model application behavior by finding behavior phases. We use conditional restricted Boltzmann machines (CRBMs) to model time-series containing resources traces measurements like CPU, memory, and IO. CRBMs can be used to map a given historic window of trace behavior into a single vector. This low dimensional and time-aware vector can be passed through clustering methods, from simplistic ones like k-means to more complex ones like those based on hidden Markov models. We use these methods to find phases of similar behavior in the workloads. Our experimental evaluation shows that the proposed method is able to identify different phases of resource consumption across different workloads. We show that the distinct phases contain specific resource patterns that distinguish them. David Buchaca Prats, Josep Lluís Berral, David Carrera 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2016 | The state of SQL-on-Hadoop in the cloudabstractManaged Hadoop in the cloud, especially SQL-on-Hadoop, has been gaining attention recently. On Platform-as-a-Service (PaaS), analytical services like Hive and Spark come pre-configured for general-purpose and ready to use. Thus, giving companies a quick entry and on-demand deployment of ready SQL-like solutions for their big data needs. This study evaluates cloud services from an end-user perspective, comparing providers including: Microsoft Azure, Amazon Web Services, Google Cloud, and Rackspace. The study focuses on performance, readiness, scalability, and cost-effectiveness of the different solutions at entry/test level clusters sizes. Results are based on over 15,000 Hive queries derived from the industry standard TPC-H benchmark. The study is framed within the ALOJA research project, which features an open source benchmarking and analysis platform that has been recently extended to support SQL-on-Hadoop engines. The ALOJA Project aims to lower the total cost of ownership (TCO) of big data deployments and study their performance characteristics for optimization. The study benchmarks cloud providers across a diverse range instance types, and uses input data scales from 1GB to 1TB, in order to survey the popular entry-level PaaS SQL-on-Hadoop solutions, thereby establishing a common results-base upon which subsequent research can be carried out by the project. Initial results already show the main performance trends to both hardware and software configuration, pricing, similarities and architectural differences of the evaluated PaaS solutions. Whereas some providers focus on decoupling storage and computing resources while offering network-based elastic storage, others choose to keep the local processing model from Hadoop for high performance, but reducing flexibility. Results also show the importance of application-level tuning and how keeping up-to-date hardware and software stacks can influence performance even more than replicating the on-premises model in the cloud. Nicolás Poggi, Josep Lluís Berral, Thomas Fenech, David Carrera 0001, José A. Blakeley, Umar Farooq Minhas, Nikola Vujic |
IEEE BigData | 2 |
| 2015 | From performance profiling to predictive analytics while evaluating hadoop cost-efficiency in ALOJAabstractDuring the past years the exponential growth of data, its generation speed, and its expected consumption rate presents one of the most important challenges in IT both for industry and research. For these reasons, the ALOJA research project was created by BSC and Microsoft as an open initiative to increase cost-efficiency and the general understanding of Big Data systems via automation and learning. The development of the project over its first year, has resulted in a open source benchmarking platform used to produce the largest public repository of Big Data results1, featuring over 42,000 job execution details. ALOJA also includes web-based analytic tools to evaluate and gather insights about cost-performance of benchmarked systems. The tools offer means to extract knowledge that can lead to optimize configuration and deployment options in the Cloud i.e., selecting the most cost-effective VMs and cluster sizes. This article describes the evolution of the project focus and research lines, for a period of over a year while continuously benchmarking systems for Big Data. As well discusses the motivation - both technical and market-based - of such changes. It also presents the main results from the evaluation of different OS and Hadoop configurations, covering over 100 hardware deployments. During this time, ALOJA's initial target has shifted from a previous low-level profiling of Hadoop runtime with HPC tools, passing through extensive benchmarking and evaluation of a large body of results via aggregation, to currently leveraging Predictive Analytics (PA) techniques. The ongoing efforts in PA show promising results to automatically model the behavior of systems i.e., predicting job execution times with high accuracy or to reduce the number of benchmark runs needed. As well as for Knowledge Discovery (KD) to find relations among software and hardware components. Techniques that jointly support foresighting cost-effectiveness of new defined systems, reducing benchmarking time and costs. Nicolás Poggi, Josep Lluís Berral, David Carrera 0001, Aaron Call, Fabrizio Gagliardi, Rob Reinauer, Nikola Vujic, Daron Green, José A. Blakeley |
IEEE BigData | 2 |
| 2015 | ALOJA-ML: A Framework for Automating Characterization and Knowledge Discovery in Hadoop DeploymentsabstractThis article presents ALOJA-Machine Learning (ALOJA-ML) an extension to the ALOJA project that uses machine learning techniques to interpret Hadoop benchmark performance data and performance tuning; here we detail the approach, efficacy of the model and initial results. The ALOJA-ML project is the latest phase of a long-term collaboration between BSC and Microsoft, to automate the characterization of cost-effectiveness on Big Data deployments, focusing on Hadoop. Hadoop presents a complex execution environment, where costs and performance depends on a large number of software (SW) configurations and on multiple hardware (HW) deployment choices. Recently the ALOJA project presented an open, vendor-neutral repository, featuring over 16.000 Hadoop executions. These results are accompanied by a test bed and tools to deploy and evaluate the cost-effectiveness of the different hardware configurations, parameter tunings, and Cloud services. Despite early success within ALOJA from expert-guided benchmarking, it became clear that a genuinely comprehensive study requires automation of modeling procedures to allow a systematic analysis of large and resource-constrained search spaces. ALOJA-ML provides such an automated system allowing knowledge discovery by modeling Hadoop executions from observed benchmarks across a broad set of configuration parameters. The resulting empirically-derived performance models can be used to forecast execution behavior of various workloads; they allow a-priori prediction of the execution times for new configurations and HW choices and they offer a route to model-based anomaly detection. In addition, these models can guide the benchmarking exploration efficiently, by automatically prioritizing candidate future benchmark tests. Insights from ALOJA-ML's models can be used to reduce the operational time on clusters, speed-up the data acquisition and knowledge discovery process, and importantly, reduce running costs. In addition to learning from the methodology presented in this work, the community can benefit in general from ALOJA data-sets, framework, and derived insights to improve the design and deployment of Big Data applications. Josep Lluís Berral, Nicolás Poggi, David Carrera 0001, Aaron Call, Rob Reinauer, Daron Green |
KDD | 1 |
| 2014 | Building Green Cloud Services at Low CostabstractInterest in powering data enters at least partially using on-site renewable sources, e.g. solar or wind, has been growing. In fact, researchers have studied distributed services comprising networks of such "green" data centers, and load distribution approaches that "follow the renewables" to maximize their use. However, prior works have not considered where to site such a network for efficient production of renewable energy, while minimizing both data center and renewable plant building costs. Moreover, researchers have not built real load management systems for follow-the-renewables services. Thus, in this paper, we propose a framework, optimization problem, and solution approach for sitting and provisioning green data centers for a follow-the-renewables HPC cloud service. We illustrate the location selection tradeoffs by quantifying the minimum cost of achieving different amounts of renewable energy. Finally, we design and implement a system capable of migrating virtual machines across the green data centers to follow the renewables. Among other interesting results, we demonstrate that one can build green HPC cloud services at a relatively low additional cost compared to existing services. Josep Lluís Berral, Íñigo Goiri, Thu D. Nguyen, Ricard Gavaldà, Jordi Torres, Ricardo Bianchini |
ICDCS | 1 |
| 2014 | Elastic operations in federated datacenters for performance and cost optimization
Luis Velasco 0001, Adrian Asensio, Josep Lluís Berral, Edoardo Bonetto, Francesco Musumeci 0001, Víctor López 0001 |
Comput. Commun. | 3 |
| 2013 | Power-Aware Multi-data Center Management Using Machine LearningabstractThe cloud relies upon multi-data center (multi-DC) infrastructures distributed along the world, where people and enterprises pay for resources to offer their web-services to worldwide clients. Intelligent management is required to automate and manage these infrastructures, as the amount of resources and data to manage exceeds the capacities of human operators. Also, it must take into account the cost of running the resources (energy) and the quality of service towards web-services and clients. (De-)consolidation and priming proximity to clients become two main strategies to allocate resources and properly place these web-services in the multi-DC network. Here we present a mathematical model to describe the scheduling problem given web-services and hosts across a multi-DC system, enhancing the decision makers with models for the system behavior obtained using machine learning. After running the system on real DC infrastructures we see that the model drives web-services to the best locations given quality of service, energy consumption, and client proximity, also (de-)consolidating according to the resources required for each web-service given its load. Josep Lluís Berral, Ricard Gavaldà, Jordi Torres |
ICPP | 1 |
| 2012 | Energy-efficient and multifaceted resource management for profit-driven virtualized data centers
Íñigo Goiri, Josep Lluís Berral, Josep Oriol Fitó, Ferran Julià, Ramon Nou, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
Future Gener. Comput. Syst. | 2 |
| 2010 | Energy-Aware Scheduling in Virtualized DatacentersabstractThe reduction of energy consumption in large-scale datacenters is being accomplished through an extensive use of virtualization, which enables the consolidation of multiple workloads in a smaller number of machines. Nevertheless, virtualization also incurs some additional overheads (e.g. virtual machine creation and migration) that can influence what is the best consolidated configuration, and thus, they must be taken into account. In this paper, we present a dynamic job scheduling policy for power-aware resource allocation in a virtualized datacenter. Our policy tries to consolidate workloads from separate machines into a smaller number of nodes, while fulfilling the amount of hardware resources needed to preserve the quality of service of each job. This allows turning off the spare servers, thus reducing the overall datacenter power consumption. As a novelty, this policy incorporates all the virtualization overheads in the decision process. In addition, our policy is prepared to consider other important parameters for a datacenter, such as reliability or dynamic SLA enforcement, in a synergistic way with power consumption. The introduced policy is evaluated comparing it against common policies in a simulated environment that accurately models HPC jobs execution in a virtualized datacenter including power consumption modeling and obtains a power consumption reduction of 15% with respect to typical policies. Íñigo Goiri, Ferran Julià, Ramon Nou, Josep Lluís Berral, Jordi Guitart, Jordi Torres |
CLUSTER | 4 |
| 2010 | Adaptive on-line software aging prediction based on machine learningabstractThe growing complexity of software systems is resulting in an increasing number of software faults. According to the literature, software faults are becoming one of the main sources of unplanned system outages, and have an important impact on company benefits and image. For this reason, a lot of techniques (such as clustering, fail-over techniques, or server redundancy) have been proposed to avoid software failures, and yet they still happen. Many software failures are those due to the software aging phenomena. In this work, we present a detailed evaluation of our chosen machine learning prediction algorithm (M5P) in front of dynamic and non-deterministic software aging. We have tested our prediction model on a three-tier web J2EE application achieving acceptable prediction accuracy against complex scenarios with small training data sets. Furthermore, we have found an interesting approach to help to determine the root cause failure: The model generated by machine learning algorithms. Javier Alonso 0001, Jordi Torres, Josep Lluís Berral, Ricard Gavaldà |
DSN | 3 |
| 2009 | Self-adaptive utility-based web session management
Nicolás Poggi, Toni Moreno, Josep Lluís Berral, Ricard Gavaldà, Jordi Torres |
Comput. Networks | 3 |