VLDB 2026 Research / reviewers in the wild / expert
David Carrera 0001
dblp:13/6290
· DBLP profile ↗
58ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0003-4898-3424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 2 first-author · 2 since 2021Computer networks · 10 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6Software engineering, systems software and programming languages · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Dexter: A Performance-Cost Efficient Resource Allocation Manager for Serverless Data AnalyticsabstractLeveraging serverless platforms for the efficient execution of distributed data analytics frameworks, such as Apache Spark [3], has gained substantial interest since early 2022. The elasticity, free-of-management, and on-demand scalability of serverless have motivated the effort in deploying distributed data analytics applications to serverless platforms. However, effectively auto-scaling resources for such complex workloads so that we can fully benefit from the resource elasticity of serverless remains challenging. Mis-configuration can result in severe performance and cost issues arising from resource under- and over-provisioning. Anna Maria Nestorov, Diego Marron, Alberto Gutierrez-Torre, Chen Wang 0039, Claudia Misale, Alaa Youssef, David Carrera 0001, Josep Lluís Berral |
Middleware | 7 |
| 2023 | Performance characterization of video analytics workloads in heterogeneous edge infrastructuresabstractSummary Powered by deep learning, video analytic applications process millions of camera feeds in real‐time to extract meaningful information from their surroundings. And this number grows by the minute. To avoid saturating the backhaul network and provide lower latencies, a distributed and heterogeneous edge cloud is postulated as a key enabler for widespread video analytics. This article provides a complete characterization of end‐to‐end video analytics across a set of hardware platforms and different neural network architectures. Each platform is selected to fill a different gap in a distributed, shared, and heterogeneous infrastructure. Moreover, we analyze how performance scales on each of these platforms with respect to the amount of resources dedicated to video analytics. Finally, we extract the key conclusions of the characterization to build an experimental model to estimate performance and cost of end‐to‐end video analytics in different edge scenarios. Our experiments show that managing video analytics workloads efficiently requires awareness of both, the platforms in which these are executed, and the full end‐to‐end pipeline. To the best of our knowledge, this is the first work that provides a complete characterization of end‐to‐end video analytics in heterogeneous edge platforms. Daniel Rivas-Barragan, Francesc Guim 0001, Jorda Polo, David Carrera 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Towards automatic model specialization for edge video analytics
Daniel Rivas-Barragan, Francesc Guim 0001, Jorda Polo, Pubudu Madhawa Silva, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 6 |
| 2022 | Automatic Distributed Deep Learning Using Resource-Constrained Edge DevicesabstractProcessing data generated at high volume and speed from the Internet of Things, smart cities, domotic, intelligent surveillance, and e-healthcare systems require efficient data processing and analytics services at the Edge to reduce the latency and response time of the applications. The fog computing edge infrastructure consists of devices with limited computing, memory, and bandwidth resources, which challenge the construction of predictive analytics solutions that require resource-intensive tasks for training machine learning models. In this work, we focus on the development of predictive analytics for urban traffic. Our solution is based on deep learning techniques localized in the Edge, where computing devices have very limited computational resources. We present an innovative method for efficiently training the gated recurrent-units (GRUs) across available resource-constrained CPU and GPU Edge devices. Our solution employs distributed GRU model learning and dynamically stops the training process to utilize the low-power and resource-constrained Edge devices while ensuring good estimation accuracy effectively. The proposed solution was extensively evaluated using low-powered ARM-based devices, including Raspberry Pi v3 and the low-powered GPU-enabled device NVIDIA Jetson Nano, and also compared them with Single-CPU Intel Xeon machines. For the evaluation experiments, we used real-world Floating Car Data. The experiments show that the proposed solution delivers excellent prediction accuracy and computational performance on the Edge when compared to the baseline methods. Alberto Gutierrez-Torre, Kiyana Bahadori, Shuja-ur-Rehman Baig, Waheed Iqbal, Tullio Vardanega, Josep Lluís Berral, David Carrera 0001 |
IEEE Internet Things J. | 7 |
| 2022 | Burst-Aware Predictive Autoscaling for Containerized MicroservicesabstractAutoscaling methods are used for cloud-hosted applications to dynamically scale the allocated resources for guaranteeing Quality-of-Service (QoS). The public-facing application serves dynamic workloads, which contain bursts and pose challenges for autoscaling methods to ensure application performance. Existing State-of-the-art autoscaling methods are burst-oblivious to determine and provision the appropriate resources. For dynamic workloads, it is hard to detect and handle bursts online for maintaining application performance. In this article, we propose a novel burst-aware autoscaling method which detects burst in dynamic workloads using workload forecasting, resource prediction, and scaling decision making while minimizing response time service-level objectives (SLO) violations. We evaluated our approach through a trace-driven simulation, using multiple synthetic and realistic bursty workloads for containerized microservices, improving performance when comparing against existing state-of-the-art autoscaling methods. Such experiments show an increase of$\times $1.09 in total processed requests, a reduction of$\times $5.17 for SLO violations, and an increase of$\times $0.767 cost as compared to the baseline method. Muhammad Abdullah 0004, Waheed Iqbal, Josep Lluís Berral, Jorda Polo, David Carrera 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Performance Evaluation of Data-Centric Workloads in Serverless EnvironmentsabstractServerless computing is a cloud-based execution paradigm that allows provisioning resources on-demand, freeing developers from infrastructure management and operational concerns. It typically involves deploying workloads as stateless functions that take no resources when not in use, and is meant to scale transparently. To make serverless effective, providers impose limits on a per-function level, such as maximum duration, fixed amount of memory, and no persistent local storage. These constraints make it challenging for data-intensive workloads to take advantage of serverless because they lead to sharing significant amounts of data through remote storage. In this paper, we build a performance model for serverless workloads that considers how data is shared between functions, including the amount of data and the underlying technology that is being used. The model's accuracy is assessed by running a real workload in a cluster using Knative, a state-of-the-art serverless environment, showing a relative error of 5.52%. With the proposed model, we evaluate the performance of data-intensive workloads in serverless, analyzing parallelism, scalability, resource requirements, and scheduling policies. We also explore possible solutions for the data-sharing problem, like using local memory and storage. Our results show that the performance of data-intensive workloads in serverless can be up to 4.32= faster depending on how these are deployed. Anna Maria Nestorov, Jorda Polo, Claudia Misale, David Carrera 0001, Alaa Youssef |
CLOUD | 4 |
| 2021 | Guest Editorial: Special Issue on Data Analytics and Machine Learning for Network and Service Management - Part IIabstractNetwork and Service analytics can harness the immense stream of operational data from clouds, to services, to social and communication networks. In the era of big data and connected devices of all varieties, analytics and machine learning have found ways to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using machine learning, artificial intelligence and data analytics to improve operations and management of information technology services, systems and networks. Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2020 | Improving maritime traffic emission estimations on missing data with CRBMs
Alberto Gutierrez-Torre, Josep Lluís Berral, David Buchaca Prats, Marc Guevara, Albert Soret, David Carrera 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2020 | Adaptive sliding windows for improved estimation of data center resource utilizationabstractAccurate prediction of data center resource utilization is required for capacity planning, job scheduling, energy saving, workload placement, and load balancing to utilize the resources efficiently. However, accurately predicting those resources is challenging due to dynamic workloads, heterogeneous infrastructures, and multi-tenant co-hosted applications. Existing prediction methods use fixed size observation windows which cannot produce accurate results because of not being adaptively adjusted to capture local trends in the most recent data. Therefore, those methods train on large fixed sliding windows using an irrelevant large number of observations yielding to inaccurate estimations or fall for inaccuracy due to degradation of estimations with short windows on quick changing trends. In this paper we propose a deep learning-based adaptive window size selection method, dynamically limiting the sliding window size to capture the trend for the latest resource utilization, then build an estimation model for each trend period. We evaluate the proposed method against multiple baseline and state-of-the-art methods, using real data-center workload data sets. The experimental evaluation shows that the proposed solution outperforms those state-of-the-art approaches and yields 16 to 54% improved prediction accuracy compared to the baseline methods. Shuja-ur-Rehman Baig, Waheed Iqbal, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 4 |
| 2020 | A highly parameterizable framework for Conditional Restricted Boltzmann Machine based workloads accelerated with FPGAs and OpenCLabstractConditional Restricted Boltzmann Machine (CRBM) is a promising candidate for a multidimensional system modeling that can learn a probability distribution over a set of data. It is a specific type of an artificial neural network with one input (visible) and one output (hidden) layer. Recently published works demonstrate that CRBM is a suitable mechanism for modeling multidimensional time series such as human motion, workload characterization, city traffic analysis. The process of learning and inference of these systems relies on linear algebra functions like matrix–matrix multiplication, and for higher data sets, they are very compute-intensive. In this paper, we present a configurable framework for CRBM based workloads for arbitrary large models. We show how to accelerate the learning process of CRBM with FPGAs and OpenCL, and we conduct an extensive scalability study for different model sizes and system configurations. We show significant improvement in performance/Watt for large models and batch sizes (from 1.51x up to 5.71x depending on the host configuration) when we use FPGA and OpenCL for the acceleration, and limited benefits for small models comparing to the state-of-the-art CPU solution. Zoran Jaksic, Nicola Cadenelli, David Buchaca Prats, Jorda Polo, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 6 |
| 2020 | Sequence-to-sequence models for workload interference prediction on batch processing datacenters
David Buchaca Prats, Joan Marcual, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 4 |
| 2020 | Guest Editorial: Special Section on Data Analytics and Machine Learning for Network and Service Management-Part I
Nur Zincir-Heywood, Giuliano Casale, David Carrera 0001, Lydia Y. Chen, Amogh Dhamdhere, Takeru Inoue, Hanan Lutfiyya, Taghrid Samak |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Constant- Time Approximate Sliding Window Framework with Error ControlabstractStream Processing is a crucial element for the Edge Computing paradigm, in which large amount of devices generate data at the edge of the network. This data needs to be aggregated and processed on-the-move across different layers before reaching the Cloud. Therefore, defining Stream Processing services that adapt to different levels of resource availability is of paramount importance. In this context, Stream Processing frameworks need to combine efficient algorithms with low computational complexity to manage sliding windows, with the ability to adjust resource demands for different deployment scenarios, from very low capacity edge devices to virtually unlimited Cloud platforms. The Approximate Computing paradigm provides improved performance and adaptive resource demands in data analytics, at the price of introducing some level of inaccuracy that can be calculated. In this paper we present the Approximate and Amortized Monoid Tree Aggregator (A2MTA). It is, to our knowledge, the first general purpose sliding window programable framework that combines constant-time aggregations with error bounded approximate computing techniques. It is very suitable for adverse stream processing environments, such as resource scarce multi-tenant edge computing. The framework can compute aggregations over multiple data dimensions, setting error bounds on any of them, and has been designed to support decoupling computation and data storage through the use of distributed Key-Value Stores to keep window elements and partial aggregations. Álvaro Villalba, David Carrera 0001 |
ISORC | 2 |
| 2019 | Considerations in using OpenCL on GPUs and FPGAs for throughput-oriented genomics workloadsabstractThe recent upsurge in the available amount of health data and the advances in next-generation sequencing are setting the ground for the long-awaited precision medicine. To process this deluge of data, bioinformatics workloads are becoming more complex and more computationally demanding. For this reasons they have been extended to support different computing architectures, such as GPUs and FPGAs, to leverage the form of parallelism typical of each of such architectures. The paper describes how a genomic workload such as k-mer frequency counting that takes advantage of a GPU can be offloaded to one or even more FPGAs. Moreover, it performs a comprehensive analysis of the FPGA acceleration comparing its performance to a non-accelerated configuration and when using a GPU. Lastly, the paper focuses on how, when using accelerators with a throughput-oriented workload, one should also take into consideration both kernel execution time and how well each accelerator board overlaps kernels and PCIe transferred. Results show that acceleration with two FPGAs can improve both time- and energy-to-solution for the entire accelerated part by a factor of 1.32x. Per contra, acceleration with one GPU delivers an improvement of 1.77x in time-to-solution but of a lower 1.49x in energy-to-solution due to persistently higher power consumption. The paper also evaluates how future FPGA boards with components (i.e., off-chip memory and PCIe) on par with those of the GPU board could provide an energy-efficient alternative to GPUs. Nicola Cadenelli, Zoran Jaksic, Jorda Polo, David Carrera 0001 |
Future Gener. Comput. Syst. | 4 |
| 2019 | Adaptive Prediction Models for Data Center Resources Utilization EstimationabstractAccurate estimation of data center resource utilization is a challenging task due to multi-tenant co-hosted applications having dynamic and time-varying workloads. Accurate estimation of future resources utilization helps in better job scheduling, workload placement, capacity planning, proactive auto-scaling, and load balancing. The inaccurate estimation leads to either under or over-provisioning of data center resources. Most existing estimation methods are based on a single model that often does not appropriately estimate different workload scenarios. To address these problems, we propose a novel method to adaptively and automatically identify the most appropriate model to accurately estimate data center resources utilization. The proposed approach trains a classifier based on statistical features of historical resources usage to decide the appropriate prediction model to use for given resource utilization observations collected during a specific time interval. We evaluated our approach on real datasets and compared the results with multiple baseline methods. The experimental evaluation shows that the proposed approach outperforms the state-of-the-art approaches and delivers 6% to 27% improved resource utilization estimation accuracy compared to baseline methods. Shuja-ur-Rehman Baig, Waheed Iqbal, Josep Lluís Berral, Abdelkarim Erradi, David Carrera 0001 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2019 | Guest Editorial: Special Issue on Novel Techniques in Big Data Analytics for ManagementabstractCloud and network analytics can harness the immense stream of operational data from clouds and networks, and can perform analytics processing to improve reliability, configuration, performance, fault and security management. In particular, we see a growing trend towards using statistical analysis, Artificial Intelligence (AI) and machine learning to improve operations and management of IT systems and networks. David Carrera 0001, Giuliano Casale, Takeru Inoue, Hanan Lutfiyya, Nur Zincir-Heywood |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2019 | Constant-Time Sliding Window Framework with Reduced Memory Footprint and Efficient Bulk EvictionsabstractThe fast evolution of data analytics platforms has resulted in an increasing demand for real-time data stream processing. From Internet of Things applications to the monitoring of telemetry generated in large data centers, a common demand for currently emerging scenarios is the need to process vast amounts of data with low latencies, generally performing the analysis process as close to the data source as possible. Stream processing platforms are required to be malleable and absorb spikes generated by fluctuations of data generation rates. Data is usually produced as time series that have to be aggregated using multiple operators, being sliding windows one of the most common abstractions used to process data in real-time. To satisfy the above-mentioned demands, efficient stream processing techniques that aggregate data with minimal computational cost need to be developed. In this paper we present the Monoid Tree Aggregator general sliding window aggregation framework, which seamlessly combines the following features: amortized$O(1)$time complexity and a worst-case of$O(\log {n})$between insertions; it provides both a window aggregation mechanism and a window slide policy that are user programmable; the enforcement of the window sliding policy exhibits amortized$O(1)$computational cost for single evictions and supports bulk evictions with cost$O(\log {n})$; and it requires a local memory space of$O(\log {n})$. The framework can compute aggregations over multiple data dimensions, and has been designed to support decoupling computation and data storage through the use of distributedKey-Value Storesto keep window elements and partial aggregations. Álvaro Villalba, Josep Lluís Berral, David Carrera 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Next Stop "NoOps": Enabling Cross-System Diagnostics Through Graph-Based Composition of Logs and MetricsabstractPerforming diagnostics in IT systems is an increasingly complicated task, and it is not doable in satisfactory time by even the most skillful operators. Systems and their architecture change very rapidly in response to business and user demand. Many organizations see value in the maintenance and management model of NoOps that stands for No Operations. One of the implementations of this model is a system that is maintained automatically without any human intervention. The path to NoOps involves not only precise and fast diagnostics but also reusing as much knowledge as possible after the system is reconfigured or changed. The biggest challenge is to leverage knowledge on one IT system and reuse this knowledge for diagnostics of another, different system. We propose a framework of weighted graphs which can transfer knowledge, and perform high-quality diagnostics of IT systems. We encode all possible data in a graph representation of a system state and automatically calculate weights of these graphs. Then, thanks to the evaluation of similarity between graphs, we transfer knowledge about failures from one system to another and use it for diagnostics. We successfully evaluate the proposed approach on Spark, Hadoop, Kafka and Cassandra systems. Michal Zasadzinski, Marc Solé, Álvaro Brandón, Victor Muntés-Mulero, David Carrera 0001 |
CLUSTER | 5 |
| 2018 | Early Termination of Failed HPC Jobs Through Machine and Deep Learning
Michal Zasadzinski, Victor Muntés-Mulero, Marc Solé, David Carrera 0001, Thomas Ludwig 0002 |
Euro-Par | 4 |
| 2018 | A resilient and distributed near real-time traffic forecasting application for Fog computing environmentsabstractIn this paper we propose an architecture for a city-wide traffic modeling and prediction service based on the Fog Computing paradigm. The work assumes an scenario in which a number of distributed antennas receive data generated by vehicles across the city. In the Fog nodes data is collected, processed in local and intermediate nodes, and finally forwarded to a central Cloud location for further analysis. We propose a combination of a data distribution algorithm, resilient to back-haul connectivity issues, and a traffic modeling approach based on deep learning techniques to provide distributed traffic forecasting capabilities. In our experiments, we leverage real traffic logs from one week of Floating Car Data (FCD) generated in the city of Barcelona by a road-assistance service fleet comprising thousands of vehicles. FCD was processed across several simulated conditions, ranging from scenarios in which no connectivity failures occurred in the Fog nodes, to situations with long and frequent connectivity outage periods. For each scenario, the resilience and accuracy of both the data distribution algorithm, and the learning methods were analyzed. Results show that the data distribution process running in the Fog nodes is resilient to back-haul connectivity issues and is able to deliver data to the Cloud location even in presence of severe connectivity problems. Additionally, the proposed traffic modeling and forecasting method exhibits better behavior when run distributed in the Fog instead of centralized in the Cloud, especially when connectivity issues occur that force data to be delivered out of order to the Cloud. Juan Luis Pérez 0003, Alberto Gutierrez-Torre, Josep Lluís Berral, David Carrera 0001 |
Future Gener. Comput. Syst. | 4 |
| 2018 | Automatic Generation of Workload Profiles Using Unsupervised Learning PipelinesabstractThe complexity of resource usage and power consumption on cloud-based applications makes the understanding of application behavior through expert examination difficult. The difficulty increases when applications are seen as “black boxes,” where only external monitoring can be retrieved. Furthermore, given the different amount of scenarios and applications, automation is required. Here, we examine and model application behavior by finding behavior phases. We use conditional restricted Boltzmann machines (CRBMs) to model time-series containing resources traces measurements like CPU, memory, and IO. CRBMs can be used to map a given historic window of trace behavior into a single vector. This low dimensional and time-aware vector can be passed through clustering methods, from simplistic ones like k-means to more complex ones like those based on hidden Markov models. We use these methods to find phases of similar behavior in the workloads. Our experimental evaluation shows that the proposed method is able to identify different phases of resource consumption across different workloads. We show that the distinct phases contain specific resource patterns that distinguish them. David Buchaca Prats, Josep Lluís Berral, David Carrera 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2017 | Topology-aware GPU scheduling for learning workloads in cloud environmentsabstractRecent advances in hardware, such as systems with multiple GPUs and their availability in the cloud, are enabling deep learning in various domains including health care, autonomous vehicles, and Internet of Things. Multi-GPU systems exhibit complex connectivity among GPUs and between GPUs and CPUs. Workload schedulers must consider hardware topology and workload communication requirements in order to allocate CPU and GPU resources for optimal execution time and improved utilization in shared cloud environments. Marcelo Amaral, Jorda Polo, David Carrera 0001, Seetharami R. Seelam, Malgorzata Steinder |
SC | 3 |
| 2016 | The state of SQL-on-Hadoop in the cloudabstractManaged Hadoop in the cloud, especially SQL-on-Hadoop, has been gaining attention recently. On Platform-as-a-Service (PaaS), analytical services like Hive and Spark come pre-configured for general-purpose and ready to use. Thus, giving companies a quick entry and on-demand deployment of ready SQL-like solutions for their big data needs. This study evaluates cloud services from an end-user perspective, comparing providers including: Microsoft Azure, Amazon Web Services, Google Cloud, and Rackspace. The study focuses on performance, readiness, scalability, and cost-effectiveness of the different solutions at entry/test level clusters sizes. Results are based on over 15,000 Hive queries derived from the industry standard TPC-H benchmark. The study is framed within the ALOJA research project, which features an open source benchmarking and analysis platform that has been recently extended to support SQL-on-Hadoop engines. The ALOJA Project aims to lower the total cost of ownership (TCO) of big data deployments and study their performance characteristics for optimization. The study benchmarks cloud providers across a diverse range instance types, and uses input data scales from 1GB to 1TB, in order to survey the popular entry-level PaaS SQL-on-Hadoop solutions, thereby establishing a common results-base upon which subsequent research can be carried out by the project. Initial results already show the main performance trends to both hardware and software configuration, pricing, similarities and architectural differences of the evaluated PaaS solutions. Whereas some providers focus on decoupling storage and computing resources while offering network-based elastic storage, others choose to keep the local processing model from Hadoop for high performance, but reducing flexibility. Results also show the importance of application-level tuning and how keeping up-to-date hardware and software stacks can influence performance even more than replicating the on-premises model in the cloud. Nicolás Poggi, Josep Lluís Berral, Thomas Fenech, David Carrera 0001, José A. Blakeley, Umar Farooq Minhas, Nikola Vujic |
IEEE BigData | 4 |
| 2015 | From performance profiling to predictive analytics while evaluating hadoop cost-efficiency in ALOJAabstractDuring the past years the exponential growth of data, its generation speed, and its expected consumption rate presents one of the most important challenges in IT both for industry and research. For these reasons, the ALOJA research project was created by BSC and Microsoft as an open initiative to increase cost-efficiency and the general understanding of Big Data systems via automation and learning. The development of the project over its first year, has resulted in a open source benchmarking platform used to produce the largest public repository of Big Data results1, featuring over 42,000 job execution details. ALOJA also includes web-based analytic tools to evaluate and gather insights about cost-performance of benchmarked systems. The tools offer means to extract knowledge that can lead to optimize configuration and deployment options in the Cloud i.e., selecting the most cost-effective VMs and cluster sizes. This article describes the evolution of the project focus and research lines, for a period of over a year while continuously benchmarking systems for Big Data. As well discusses the motivation - both technical and market-based - of such changes. It also presents the main results from the evaluation of different OS and Hadoop configurations, covering over 100 hardware deployments. During this time, ALOJA's initial target has shifted from a previous low-level profiling of Hadoop runtime with HPC tools, passing through extensive benchmarking and evaluation of a large body of results via aggregation, to currently leveraging Predictive Analytics (PA) techniques. The ongoing efforts in PA show promising results to automatically model the behavior of systems i.e., predicting job execution times with high accuracy or to reduce the number of benchmark runs needed. As well as for Knowledge Discovery (KD) to find relations among software and hardware components. Techniques that jointly support foresighting cost-effectiveness of new defined systems, reducing benchmarking time and costs. Nicolás Poggi, Josep Lluís Berral, David Carrera 0001, Aaron Call, Fabrizio Gagliardi, Rob Reinauer, Nikola Vujic, Daron Green, José A. Blakeley |
IEEE BigData | 3 |
| 2015 | Spark deployment and performance evaluation on the MareNostrum supercomputerabstractIn this paper we present a framework to enable data-intensive Spark workloads on MareNostrum, a petascale supercomputer designed mainly for compute-intensive applications. As far as we know, this is the first attempt to investigate optimized deployment configurations of Spark on a petascale HPC setup. We detail the design of the framework and present some benchmark data to provide insights into the scalabilityof the system. We examine the impact of different configurations including parallelism, storage and networking alternatives, and we discuss several aspects in executing Big Data workloads on a computing system that is based on the compute-centric paradigm. Further, we derive conclusions aiming to pave the way towards systematic and optimized methodologies for fine-tuning data-intensive application on large clusters emphasizing on parallelism configurations. Rubén Tous, Anastasios Gounaris, Carlos Tripiana, Jordi Torres, Sergi Girona, Eduard Ayguadé, Jesús Labarta, Yolanda Becerra 0001, David Carrera 0001, Mateo Valero |
IEEE BigData | 9 |
| 2015 | ALOJA-ML: A Framework for Automating Characterization and Knowledge Discovery in Hadoop DeploymentsabstractThis article presents ALOJA-Machine Learning (ALOJA-ML) an extension to the ALOJA project that uses machine learning techniques to interpret Hadoop benchmark performance data and performance tuning; here we detail the approach, efficacy of the model and initial results. The ALOJA-ML project is the latest phase of a long-term collaboration between BSC and Microsoft, to automate the characterization of cost-effectiveness on Big Data deployments, focusing on Hadoop. Hadoop presents a complex execution environment, where costs and performance depends on a large number of software (SW) configurations and on multiple hardware (HW) deployment choices. Recently the ALOJA project presented an open, vendor-neutral repository, featuring over 16.000 Hadoop executions. These results are accompanied by a test bed and tools to deploy and evaluate the cost-effectiveness of the different hardware configurations, parameter tunings, and Cloud services. Despite early success within ALOJA from expert-guided benchmarking, it became clear that a genuinely comprehensive study requires automation of modeling procedures to allow a systematic analysis of large and resource-constrained search spaces. ALOJA-ML provides such an automated system allowing knowledge discovery by modeling Hadoop executions from observed benchmarks across a broad set of configuration parameters. The resulting empirically-derived performance models can be used to forecast execution behavior of various workloads; they allow a-priori prediction of the execution times for new configurations and HW choices and they offer a route to model-based anomaly detection. In addition, these models can guide the benchmarking exploration efficiently, by automatically prioritizing candidate future benchmark tests. Insights from ALOJA-ML's models can be used to reduce the operational time on clusters, speed-up the data acquisition and knowledge discovery process, and importantly, reduce running costs. In addition to learning from the methodology presented in this work, the community can benefit in general from ALOJA data-sets, framework, and derived insights to improve the design and deployment of Big Data applications. Josep Lluís Berral, Nicolás Poggi, David Carrera 0001, Aaron Call, Rob Reinauer, Daron Green |
KDD | 3 |
| 2015 | Performance Evaluation of Microservices Architectures Using ContainersabstractMicro services architecture has started a new trend for application development for a number of reasons: (1) to reduce complexity by using tiny services, (2) to scale, remove and deploy parts of the system easily, (3) to improve flexibility to use different frameworks and tools, (4) to increase the overall scalability, and (5) to improve the resilience of the system. Containers have empowered the usage of micro services architectures by being lightweight, providing fast start-up times, and having a low overhead. Containers can be used to develop applications based on monolithic architectures where the whole system runs inside a single container or inside a micro services architecture where one or few processes run inside the containers. Two models can be used to implement a micro services architecture using containers: master-slave, or nested-container. The goal of this work is to compare the performance of CPU and network running benchmarks in the two aforementioned models of micro services architecture hence provide a benchmark analysis guidance for system designers. Marcelo Amaral, Jorda Polo, David Carrera 0001, Iqbal Mohomed, Merve Unuvar, Malgorzata Steinder |
NCA | 3 |
| 2014 | ALOJA: A systematic study of Hadoop deployment variables to enable automated characterization of cost-effectivenessabstractThis article presents the ALOJA project, an initiative to produce mechanisms for an automated characterization of cost-effectiveness of Hadoop deployments and reports its initial results. ALOJA is the latest phase of a long-term collaborative engagement between BSC and Microsoft which, over the past 6 years has explored a range of different aspects of computing systems, software technologies and performance profiling. While during the last 5 years, Hadoop has become the de-facto platform for Big Data deployments, still little is understood of how the different layers of the software and hardware deployment options affects its performance. Early ALOJA results show that Hadoop's runtime performance, and therefore its price, are critically affected by relatively simple software and hardware configuration choices e.g., number of mappers, compression, or volume configuration. Project ALOJA presents a vendor-neutral repository featuring over 5000 Hadoop runs, a test bed, and tools to evaluate the cost-effectiveness of different hardware, parameter tuning, and Cloud services for Hadoop. As few organizations have the time or performance profiling expertise, we expect our growing repository will benefit Hadoop customers to meet their Big Data application needs. ALOJA seeks to provide both knowledge and an online service to with which users make better informed configuration choices for their Hadoop compute infrastructure whether this be on-premise or cloud-based. The initial version of ALOJA's Web application and sources are available at http://hadoop.bsc.es Nicolás Poggi, David Carrera 0001, Aaron Call, Sergio Mendoza, Yolanda Becerra 0001, Jordi Torres, Eduard Ayguadé, Fabrizio Gagliardi, Jesús Labarta, Rob Reinauer, Nikola Vujic, Daron Green, José A. Blakeley |
IEEE BigData | 2 |
| 2014 | Adaptive MapReduce Scheduling in Shared EnvironmentsabstractIn this paper we present a MapReduce task scheduler for shared environments in which MapReduce is executed along with other resource-consuming workloads, such as transactional applications. All workloads may potentially share the same data store, some of them consuming data for analytics purposes while others acting as data generators. This kind of scenario is becoming increasingly important in data centers where improved resource utilization can be achieved through workload consolidation, and is specially challenging due to the interaction between workloads of different nature that compete for limited resources. The proposed scheduler aims to improve resource utilization across machines while observing completion time goals. Unlike other MapReduce schedulers, our approach also takes into account the resource demands for non-MapReduce workloads, and assumes that the amount of resources made available to the MapReduce applications is variable over time. As shown in our experiments, our proposal improves the management of MapReduce jobs in the presence of variable resource availability, increasing the accuracy of the estimations made by the scheduler, thus improving completion time goals without an impact on the fairness of the scheduler. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Malgorzata Steinder |
CCGRID | 3 |
| 2014 | Profit-aware cloud resource provisioner for ecommerceabstractIn recent years, the Cloud Computing paradigm has proven effective in scaling dynamically the number of servers according to simple performance metrics and the incoming workload. However while some applications are able to scale-out, as current scaling metrics do not relate system performance to sales, hosting costs and profits are not optimized completely. The following article proposes a novel technique for dynamic resource provisioning based on revenue and cost metrics, to optimize profits for online retailers in the Cloud. The proposal relies on user behavior models that relate Quality-of-Service (QoS) to service capacity, and to the intention of users to buy a product on an Ecommerce site. We show how such metrics can enable profit-aware resource management by setting an optimal number of servers at each time of the day. Experiments are performed on custom, real-life datasets from an Ecommerce retailer contain over two years of access, performance, and sales data from popular travelWeb applications. Nicolás Poggi, David Carrera 0001, Eduard Ayguadé, Jordi Torres |
CLUSTER | 2 |
| 2013 | Business Process Mining from E-Commerce Web Logs
Nicolás Poggi, Vinod Muthusamy, David Carrera 0001, Rania Khalaf |
BPM | 3 |
| 2013 | Enabling Distributed Key-Value Stores with Low Latency-Impact Snapshot SupportabstractCurrent distributed key-value stores generally provide greater scalability at the expense of weaker consistency and isolation. However, additional isolation support is becoming increasingly important in the environments in which these stores are deployed, where different kinds of applications with different needs are executed, from transactional workloads to data analytics. While fully-fledged ACID support may not be feasible, it is still possible to take advantage of the design of these data stores, which often include the notion of multiversion concurrency control, to enable them with additional features at a much lower performance cost and maintaining its scalability and availability. In this paper we explore the effects that additional consistency guarantees and isolation capabilities may have on a state of the art key-value store: Apache Cassandra. We propose and implement a new multiversioned isolation level that provides stronger guarantees without compromising Cassandra's scalability and availability. As shown in our experiments, our version of Cassandra allows Snapshot Isolation-like transactions, preserving the overall performance and scalability of the system. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Mike Spreitzer, Malgorzata Steinder |
NCA | 3 |
| 2013 | Deadline-Based MapReduce Workload ManagementabstractThis paper presents a scheduling technique for multi-job MapReduce workloads that is able to dynamically build performance models of the executing workloads, and then use these models for scheduling purposes. This ability is leveraged to adaptively manage workload performance while observing and taking advantage of the particulars of the execution environment of modern data analytics applications, such as hardware heterogeneity and distributed storage. The technique targets a highly dynamic environment in which new jobs can be submitted at any time, and in which MapReduce workloads share physical resources with other workloads. Thus the actual amount of resources available for applications can vary over time. Beyond the formulation of the problem and the description of the algorithm and technique, a working prototype (called Adaptive Scheduler) has been implemented. Using the prototype and medium-sized clusters (of the order of tens of nodes), the following aspects have been studied separately: the scheduler's ability to meet high-level performance goals guided only by user-defined completion time goals; the scheduler's ability to favor data-locality in the scheduling algorithm; and the scheduler's ability to deal with hardware heterogeneity, which introduces hardware affinity and relative performance characterization for those applications that can benefit from executing on specialized processors. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2012 | Energy accounting for shared virtualized environments under DVFS using PMC-based power models
Ramon Bertran Monfort, Yolanda Becerra 0001, David Carrera 0001, Vicenç Beltran 0001, Marc González 0001, Xavier Martorell, Nacho Navarro, Jordi Torres, Eduard Ayguadé |
Future Gener. Comput. Syst. | 3 |
| 2012 | Autonomic Placement of Mixed Batch and Transactional WorkloadsabstractTo reduce the cost of infrastructure and electrical energy, enterprise datacenters consolidate workloads on the same physical hardware. Often, these workloads comprise both transactional and long-running analytic computations. Such consolidation brings new performance management challenges due to the intrinsically different nature of a heterogeneous set of mixed workloads, ranging from scientific simulations to multitier transactional applications. The fact that such different workloads have different natures imposes the need for new scheduling mechanisms to manage collocated heterogeneous sets of applications, such as running a web application and a batch job on the same physical server, with differentiated performance goals. In this paper, we present a technique that enables existing middleware to fairly manage mixed workloads: long running jobs and transactional applications. Our technique permits collocation of the workload types on the same physical hardware, and leverages virtualization control mechanisms to perform online system reconfiguration. In our experiments, including simulations as well as a prototype system built on top of state-of-the-art commercial middleware, we demonstrate that our technique maximizes mixed workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Resource-Aware Adaptive Scheduling for MapReduce Clusters
Jorda Polo, Claris Castillo, David Carrera 0001, Yolanda Becerra 0001, Ian Whalley, Malgorzata Steinder, Jordi Torres, Eduard Ayguadé |
Middleware | 3 |
| 2011 | Non-intrusive Estimation of QoS Degradation Impact on E-Commerce User SatisfactionabstractWith the massification of high speed Internet access, recent industry consumer reports show that Web site performance is increasingly becoming a key feature in determining user satisfaction, and finally, a decisive factor in whether a user will purchase on a Web site or even return to it. Traditional Web infrastructure capacity planning has focused on maintaining high throughput and availability on Web sites, optimizing the number of servers to serve peak hours to minimize costs. However, as we will show with our study, the conversion rate-the fraction of users that purchase on a site-is higher at peak hours, where systems are more exposed to suffer overload. In this article we propose a methodology to determine the thresholds of user satisfaction as the QoS delivered by an online business degrades, and to estimate its effects on actual sales. The novelty of the presented technique is that it does not involve any intrusive manipulation of production systems, but a learning process over historic sales data that is combined with system performance measurements. The methodology has been applied to Atra palo.com, a top national Travel and Booking site. For our experiments, we were given access to a 3 year long sales history dataset, as well as actual HTTP and resource consumption logs for several weeks. Obtained results enable autonomic resource managers to set best performance goals and optimize the number of server according to the workload, without surpassing the thresholds of user satisfaction and maximizing revenue for the site. Nicolás Poggi, David Carrera 0001, Ricard Gavaldà, Eduard Ayguadé |
NCA | 2 |
| 2011 | Adaptive resource provisioning for read intensive multi-tier applications in the cloud
Waheed Iqbal, Matthew N. Dailey, David Carrera 0001, Paul Janecek |
Future Gener. Comput. Syst. | 3 |
| 2010 | SLA-Driven Dynamic Resource Management for Multi-tier Web Applications in a CloudabstractCurrent service-level agreements (SLAs) offered by cloud providers do not make guarantees about response time of Web applications hosted on the cloud. Satisfying a maximum average response time guarantee for Web applications is difficult due to unpredictable traffic patterns. The complex nature of multi-tier Web applications increases the difficulty of identifying bottlenecks and resolving them automatically. It may be possible to minimize the probability that tiers (hosted on virtual machines) become bottlenecks by optimizing the placement of the virtual machines in a cloud. This research focuses on enabling clouds to offer multi-tier Web application owners maximum response time guarantees while minimizing resource utilization. We present our basic approach, preliminary experiments, and results on a EUCALYPTUS-based testbed cloud. Our preliminary results shows that dynamic bottleneck detection and resolution for multi-tier Web application hosted on the cloud will help to offer SLAs that can offer response time guarantees. Waheed Iqbal, Matthew N. Dailey, David Carrera 0001 |
CCGRID | 3 |
| 2010 | SLA-Driven Automatic Bottleneck Detection and Resolution for Read Intensive Multi-tier Applications Hosted on a Cloud
Waheed Iqbal, Matthew N. Dailey, David Carrera 0001, Paul Janecek |
GPC | 3 |
| 2010 | Performance Management of Accelerated MapReduce Workloads in Heterogeneous ClustersabstractNext generation data centers will be composed of thousands of hybrid systems in an attempt to increase overall cluster performance and to minimize energy consumption. New programming models, such as MapReduce, specifically designed to make the most of very large infrastructures will be leveraged to develop massively distributed services. At the same time, data centers will bring an unprecedented degree of workload consolidation, hosting in the same infrastructure distributed services from many different users. In this paper we present our advancements in leveraging the Adaptive MapReduce Scheduler to meet user defined high level performance goals while transparently and efficiently exploiting the capabilities of hybrid systems. While the Adaptive Scheduler was already able to dynamically allocate resources to co-located MapReduce jobs based on their completion time goals, it was completely unaware of specific hardware capabilities. In our work we describe the changes introduced in the Adaptive Scheduler to enable it with hardware awareness and with the ability to co-schedule accelerable and non-accelerable jobs on the same heterogeneous MapReduce cluster, making the most of the underlying hybrid systems. The developed prototype is tested in a cluster of Cell/BE blades and relies on the use of accelerated and non-accelerated versions of the MapReduce tasks of different deployed applications to dynamically select the best version to run on each node. Decisions are made after workload composition and jobs' completion time goals. Results show that the augmented Adaptive Scheduler provides dynamic resource allocation across jobs, hardware affinity when possible, and is even able to spread jobs' tasks across accelerated and non-accelerated nodes in order to meet performance goals in extreme conditions. To our knowledge this is the first MapReduce scheduler and prototype that is able to manage high-level performance goals even in presence of hybrid systems and accelerable jobs. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 2 |
| 2010 | Performance-driven task co-scheduling for MapReduce environmentsabstractMapReduce is a data-driven programming model proposed by Google in 2004 which is especially well suited for distributed data analytics applications. We consider the management of MapReduce applications in an environment where multiple applications share the same physical resources. Such sharing is in line with recent trends in data center management which aim to consolidate workloads in order to achieve cost and energy savings. In a shared environment, it is necessary to predict and manage the performance of workloads given a set of performance goals defined for them. In this paper, we address this problem by introducing a new task scheduler for a MapReduce framework that allows performance-driven management of MapReduce tasks. The proposed task scheduler dynamically predicts the performance of concurrent MapReduce jobs and adjusts the resource allocation for the jobs. It allows applications to meet their performance objectives without over-provisioning of physical resources. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Malgorzata Steinder, Ian Whalley |
NOMS | 2 |
| 2009 | SLA-Driven Adaptive Resource Management for Web Applications on a Heterogeneous Compute Cloud
Waheed Iqbal, Matthew N. Dailey, David Carrera 0001 |
CloudCom | 3 |
| 2009 | CellMT: A cooperative multithreading library for the Cell/B.EabstractThe Cell BE processor has proved that heterogeneous multi-core systems can provide a huge computational power with high efficiency for a wide range of applications. The simple design of the computational units and the use of small managed local memories is the key to achieve high efficiency and performance at the same time. However, this simple and efficient hardware design comes at the price of higher code complexity. The code written to run in this kind of processors must deal with several issues such as code vectorization, loop unrolling or the explicit management of local memories. Some of these issues such as vectorization or loop unrolling can be partially solved by the compiler, but the overlapping of data transfer and computation times must be manually addressed by the programmer with techniques such as double buffering that increase the code complexity. In this paper we present a user level threading library called CellMT that effectively hide memory latencies. The concurrent execution of several threads inside each SPU naturally overlaps computation and data transfer times without increasing the code complexity. To prove the suitability and feasibility of our multi-threaded library, we perform an exhaustive performance evaluation with a synthetic benchmark and a real application. The experimental results show that the multithreaded approach can outperform a hand-coded double buffering scheme, with speedups from 0.96x to 3.2x, while maintaining the complexity of a naive buffering scheme. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
HiPC | 2 |
| 2009 | Speeding Up Distributed MapReduce Applications Using Hardware AcceleratorsabstractIn an attempt to increase the performance/cost ratio, large compute clusters are becoming heterogeneous at multiple levels: from asymmetric processors, to different system architectures, operating systems and networks. Exploiting the intrinsic multi-level parallelism present in such a complex execution environment has become a challenging task using traditional parallel and distributed programming models. As a result, an increasing need for novel approaches to exploiting parallelism has arisen in these environments. MapReduce is a data-driven programming model originally proposed by Google back in 2004 as a flexible alternative to the existing models, specially devoted to hiding the complexity of both developing and running massively distributed applications in large compute clusters. In some recent works, the MapReduce model has been also used to exploit parallelism in other non-distributed environments, such as multi-cores, heterogeneous processors and GPUs. In this paper we introduce a novel approach for exploiting the heterogeneity of a Cell BE cluster linking an existing MapReduce runtime implementation for distributed clusters and one runtime to exploit the parallelism of the Cell BE nodes. The novel contribution of this work is the design and evaluation of a MapReduce execution environment that effectively exploits the parallelism existing at both the Cell BE cluster level and the heterogeneous processors level. Yolanda Becerra 0001, Vicenç Beltran 0001, David Carrera 0001, Marc González 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 3 |
| 2009 | Batch Job Profiling and Adaptive Profile Enforcement for Virtualized EnvironmentsabstractData center management is driven by high-level performance goals, and it is the responsibility of a management middleware to ensure that those goals are met using dynamic resource allocation. The performance delivered by the heterogeneous setoff applications running in a virtualized enterprise data center must be predicted to make resource allocation decisions. For some of these applications, it is required to produce accurate profiles based on previous executions: that is thecae of batch jobs.In this paper we propose a methodology to produce resource consumption profiles for batch applications running inside of virtual machines and a technique to enforce and adapt the profiles to actual execution conditions and application performance. For this purpose we have developed a testing prototype. The enforcement technique observes the fact that management middleware usually run in control cycles in which the system can be reconfigured, what imposes a tradeoff between the accuracy of the profiles and their applicability in real deployments.The novel contribution of this work is the study of the trade off between accuracy and applicability of workload profiles, what is a necessary step to enable existing management middleware with the performance prediction mechanisms required to perform effective dynamic resource allocation. Yolanda Becerra 0001, David Carrera 0001, Eduard Ayguadé |
PDP | 2 |
| 2008 | Managing SLAs of heterogeneous workloads using dynamic application placementabstractIn this paper we address the problem of managing heterogeneous workloads in a virtualized data center. We consider two different workloads: transactional applications and long-running jobs. We present a technique that permits collocation of these workload types on the same physical hardware. Our technique dynamically modifies workload placement by leveraging control mechanisms such as suspension and migration, and strives to optimally trade off resource allocation among these workloads in spite of their differing characteristics and performance objectives. Our approach builds upon our previous work on dynamically placing transactional workloads. This paper extends our framework with the capability to manage long-running workloads. We achieve this goal by using utility functions, which permit us to compare the performance of various workloads, and which are used to drive allocation decisions. We demonstrate that our technique maximizes heterogeneous workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
HPDC | 1 |
| 2008 | Reducing wasted resources to help achieve green data centersabstractIn this paper we introduce a new approach to the consolidation strategy of a data center that allows an important reduction in the amount of active nodes required to process a heterogeneous workload without degrading the offered service level. This article reflects and demonstrates that consolidation of dynamic workloads does not end with virtualization. If energy-efficiency is pursued, the workloads can be consolidated even more using two techniques, memory compression and request discrimination, which were separately studied and validated in previous work and are now to be combined in a joint effort. We evaluate the approach using a representative workload scenario composed of numerical applications and a real workload obtained from a top national travel website. Our results indicate that an important improvement can be achieved using 20% less servers to do the same work. We believe that this serves as an illustrative example of a new way of management: tailoring the resources to meet high level energy efficiency goals. Jordi Torres, David Carrera 0001, Kevin Hogan, Ricard Gavaldà, Vicenç Beltran 0001, Nicolás Poggi |
IPDPS | 2 |
| 2008 | Enabling Resource Sharing between Transactional and Batch Workloads Using Dynamic Application Placement
David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
Middleware | 1 |
| 2008 | Utility-based placement of dynamic Web applications with fairness goalsabstractWe study the problem of dynamic resource allocation to clustered Web applications. We extend application server middleware with the ability to automatically decide the size of application clusters and their placement on physical machines. Unlike existing solutions, which focus on maximizing resource utilization and may unfairly treat some applications, the approach introduced in this paper considers the satisfaction of each application with a particular resource allocation and attempts to at least equally satisfy all applications. We model satisfaction using utility functions, mapping CPU resource allocation to the performance of an application relative to its objective. The demonstrated online placement technique aims at equalizing the utility value across all applications while also satisfying operational constraints, preventing the over-allocation of memory, and minimizing the number of placement changes. We have implemented our technique in a leading commercial middleware product. Using this real-life testbed and a simulation we demonstrate the benefit of the utility-driven technique as compared to other state-of-the-art techniques. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
NOMS | 1 |
| 2008 | Dynamic CPU provisioning for self-managed secure web applications in SMP hosting platforms
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 2 |
| 2007 | Server virtualization in autonomic management of heterogeneous workloadsabstractServer virtualization opens up a range of new possibilities for autonomic datacenter management, through the availability of new automation mechanisms that can be exploited to control and monitor tasks running within virtual machines. This offers not only new and more flexible control to the operator using a management console, but also more powerful and flexible autonomic control, through management software that maintains the system in a desired state in the face of changing workload and demand. This paper explores in particular the use of server virtualization technology in the autonomic management of data centers running a heterogeneous mix of workloads. We present a system that manages heterogeneous workloads to their performance goals and demonstrate its effectiveness via real-system experiments and simulation. We also present some of the significant challenges to wider usage of virtual servers in autonomic datacenter management. Malgorzata Steinder, Ian Whalley, David Carrera 0001, Ilona Gaweda, David M. Chess |
Integrated Network Management | 3 |
| 2007 | Monitoring and Analysis Framework for Grid MiddlewareabstractAs the use of complex grid middleware becomes widespread and more facilites are offered by these pieces of software, distributed grid applications are becoming more and more popular. But as grid middleware grows in size and offers more advanced features, they become more complex and heavier, as well as harder to tune. Since the performance of a distributed grid application can be strongly influenced by the operation of the underlying grid middleware, it becomes of extreme importance to study and analyse its behaviour and performance. In this paper we present the eDragon monitoring framework (eDMF), a set of tools that can be used for the instrumentation and analysis of grid middleware, and which provides a unique environment to study the performance of grid applications. The eDMF is composed of a set of specialised monitoring tools as well as by a flexible and powerful performance analysis platform. Additionally we also provide a practical application of the eDMF to the Globus toolkit 4 (GT4), one of the most extended and popular grid middleware, showing how it helped us in the detection and resolution of several job management problems observed in the GT4 middleware Ramon Nou, Ferran Julià, David Carrera 0001, Kevin Hogan, Jordi Caubet, Jesús Labarta, Jordi Torres |
PDP | 3 |
| 2007 | Designing an overload control strategy for secure e-commerce applications
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 2 |
| 2005 | A Hybrid Web Server Architecture for Secure e-Business Web Applications
Vicenç Beltran 0001, David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé |
HPCC | 2 |
| 2005 | Session-Based Adaptive Overload Control for Secure Dynamic Web ApplicationsabstractAs dynamic Web content and security capabilities are becoming popular in current Web sites, the performance demand on application servers that host the sites is increasing, leading sometimes these servers to overload. As a result, response times may grow to unacceptable levels and the server may saturate or even crash. In this paper we present a session-based adaptive overload control mechanism based on SSL (secure socket layer) connections differentiation and admission control. The SSL connections differentiation is a key factor because the cost of establishing a new SSL connection is much greater than establishing a resumed SSL connection (it reuses an existing SSL session on server). Considering this big difference, we have implemented an admission control algorithm that prioritizes the resumed SSL connections to maximize performance on session-based environments and limits dynamically the number of new SSL connections accepted depending on the available resources and the current number of connections in the system to avoid server overload. In order to allow the differentiation of resumed SSL connections from new SSL connections we propose a possible extension of the Java Secure Sockets Extension (JSSE) API. Our evaluation on Tomcat server demonstrates the benefit of our proposal for preventing server overload. Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 2 |
| 2004 | Evaluating the Scalability of Java Event-Driven Web ServersabstractThe two major strategies used to construct high-performance Web servers are thread pools and event-driven architectures. The Java platform is commonly used in Web environments but up to the moment it did not provide any standard API to implement event-driven architectures efficiently. The new 1.4 release of the J2SE introduces the NIO (New I/O) API to help in the development of event-driven I/O intensive applications. We evaluate the scalability that this API provides to the Java platform in the field of Web servers, bringing together the majorly used commercial server (Apache) and one experimental server developed using the NIO API. We study the scalability of the NIO-based server as well as of its rival in a number of different scenarios, including uniprocessor, multiprocessor, bandwidth-bounded and CPU-bounded environments. The study concludes that the NIO API can be successfully used to create event-driven Java servers that can scale as well as the best of the commercial native-compiled Web server, at a fraction of its complexity and using only one or two worker threads. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 2 |
| 2003 | Complete instrumentation requirements for performance analysis of Web based technologiesabstractIn this paper we present the eDragon environment, a research platform created to perform complete performance analysis of new Web-based technologies. eDragon enables the understanding of how application servers work in both sequential and parallel platforms offering a new insight in the usage of system resources. The environment is composed of a set of instrumentation modules, a performance analysis and visualization tool and a set of experimental methodologies to perform complete performance analysis of Web-based technologies. This paper describes the design and implementation of this research platform and highlights some of its main functionalities. We will also show how a detailed analytical view can be obtained through the application of a bottom-up strategy, starting with a group of system events and advancing to more complex performance metrics using a continuous derivation process. David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé, Jesús Labarta |
ISPASS | 1 |