VLDB 2026 Research / reviewers in the wild / expert
Jordi Torres
dblp:94/6751
· DBLP profile ↗
79ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9Software engineering, systems software and programming languages · 6Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Computer networks · 5Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FRIDA: Free-rider detection using privacy attacksabstractFederated learning is increasingly popular as it enables multiple parties with limited datasets and resources to train a machine learning model collaboratively. However, similar to other collaborative systems, federated learning is vulnerable to free-riders — participants who benefit from the global model without contributing. Free-riders compromise the integrity of the learning process and slow down the convergence of the global model, resulting in increased costs for honest participants. To address this challenge, we propose FRIDA: f ree- ri der d etection using privacy a ttacks. Instead of focusing on implicit effects of free-riding, FRIDA utilizes membership and property inference attacks to directly infer evidence of genuine client training. Our extensive evaluation demonstrates that FRIDA is effective across a wide range of scenarios. Pol G. Recasens, Ádám Horváth, Alberto Gutierrez-Torre, Jordi Torres, Josep Lluís Berral, Balazs Pejo |
J. Inf. Secur. Appl. | 4 |
| 2025 | Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM InferenceabstractLarge language models have been widely adopted across different tasks, but their auto-regressive generation nature often leads to inefficient resource utilization during inference. While batching is commonly used to increase throughput, per-formance gains plateau beyond a certain batch size, especially with smaller models, a phenomenon that existing literature typically explains as a shift to the compute-bound regime. In this paper, through an in-depth GPU-level analysis, we reveal that large-batch inference remains memory-bound, with most GPU compute capabilities underutilized due to DRAM bandwidth saturation as the primary bottleneck. To address this, we propose a Batching Configuration Advisor (BCA) that optimizes memory allocation, reducing GPU memory requirements with minimal impact on throughput. The freed memory and underutilized GPU compute capabilities can then be leveraged by concurrent workloads. Specifically, we use model replication to improve serving throughput and GPU utilization. Our findings challenge conventional assumptions about LLM inference, offering new in-sights and practical strategies for improving resource utilization, particularly for smaller language models. Pol G. Recasens, Ferran Agullo, Chen Wang 0039, Olivier Tardieu, Jordi Torres, Josep Lluís Berral |
CLOUD | 7 |
| 2025 | Enhancing the output of time series forecasting algorithms for cloud resource provisioningabstractForecasting the resource consumption of workloads is a frequent approach in the cloud provisioning field. Ideally, such predictions allow obtaining a more accurate scheduling and management of resources in a computing cluster. However, the current approaches fail to properly forecast the future consumption in areas where sudden increases of consumption are present, i.e ., spikes. Even, commonly employed metrics lack the ability to properly evaluate sharp behaviours in the traces. This may generate resource starvation problems in the running workloads and decreases the Quality of Service (QoS) provided to external users. To address this issue, we propose two strategies that modify the outputs of forecasting algorithms without changing the algorithms’ internals. The new outputs considerably enhance the prediction of sudden increases, duplicating the F1 score metric in average for all tested algorithms. This improvement in the handling of spikes comes with an increased over-provision of resources. Nevertheless, the proposed strategies give the user an easy way to control this trade-off between predicting spikes and the amount of over-provision. The user can decide which is the right balance that better fits the requirements of its specific scenario. Furthermore, we propose a new evaluation methodology that better assesses the behaviour of forecasting algorithms in cloud traces, especially focused on the performance around increases of consumption, and we give insights on the reasons behind the predictions of the algorithms with the application of explainability techniques. The code repository of this work can be accessed through GitHub at this link https://github.com/FerranAgulloLopez/ResourceForecasting . • Tackling the forecasting of workload resource consumption for cloud provisioning. • The forecasts can improve the sharing of resources between co-allocated workloads. • The current approaches lie far behind when predicting increases of consumption. • The work proposes new strategies and a new evaluation to enhance the forecasts. • Explainability techniques are used to understand the predictions in cloud time series. Ferran Agullo, Alberto Gutierrez-Torre, Jordi Torres, Josep Lluís Berral |
Future Gener. Comput. Syst. | 3 |
| 2023 | Real-World Evidence Inclusion in Guideline-Based Clinical Decision Support Systems: Breast Cancer Use Case
Jordi Torres, Eduardo Alonso 0004, Nekane Larburu |
AIME | 1 |
| 2023 | Wellbeing Recommender System, a User-Centered Framework for Generating a Recommender System for Healthy Aging
Jordi Torres, Meritxell Garcia Perea, Garazi Artola, Teresa García-Navarro, Isabel Amaya, Nekane Larburu |
ICT4AWE | 1 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 87 |
| 2023 | A closer look at referring expressions for video object segmentationabstractAbstract The task of Language-guided Video Object Segmentation (LVOS) aims at generating binary masks for an object referred by a linguistic expression. When this expression unambiguously describes an object in the scene, it is namedreferring expression(RE). Our work argues that existing benchmarks used for LVOS are mainly composed of trivial cases, in which referents can be identified with simple phrases. Our analysis relies on a new categorization of the referring expressions in the DAVIS-2017 and Actor-Action datasets into trivial and non-trivial REs, where the non-trivial REs are further annotated with seven RE semantic categories. We leverage these data to analyze the performance of RefVOS, a novel neural network that obtains competitive results for the task of language-guided image segmentation and state of the art results for LVOS. Our study indicates that the major challenges for the task are related to understanding motion and static actions. Miriam Bellver, Carles Ventura, Carina Silberer, Ioannis Kazakos, Jordi Torres, Xavier Giró-i-Nieto |
Multim. Tools Appl. | 5 |
| 2022 | Personalized Nutritional Guidance System to Prevent Malnutrition in Pluripathological Older PatientsabstractMalnutrition is a frequent problem in the elderly population, who usually is affected by one or more pathologies. The health status of these patients can get worsened if malnutrition is left untreated. Nutritional guidelines have been developed to fulfil the nutritional needs derived from certain pathologies, but still are not easy to use. Digital tools can help implement and use these guidelines in real clinical scenarios. Current solutions are designed around a single pathology or specific scenario, but the pluripathologic scenario presents a challenge when it comes to provide nutritional support. In this paper, we present an adaptative tool that provides personalized nutritional recommendations for pluripathological patients in an efficient way, and can be extended to include other pathologies. Jordi Torres, Garazi Artola, Nekane Larburu, Amaia Agirre, Elixabete Narbaiza, Idoia Berges, Ainhoa Lizaso |
ICT4AWE | 1 |
| 2021 | How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign LanguageabstractOne of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous American Sign Language (ASL) dataset, consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English transcripts, and depth. A three-hour subset was further recorded in the Panoptic studio enabling detailed 3D pose estimation. To evaluate the potential of How2Sign for real-world impact, we conduct a study with ASL signers and show that synthesized videos using our dataset can indeed be understood. The study further gives insights on challenges that computer vision should address in order to make progress in this field. Amanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, Xavier Giró-i-Nieto |
CVPR | 7 |
| 2020 | Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsabstractAcquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research. This problem has been studied through the lens of empowerment, which draws a connection between option discovery and information theory. Information-theoretic skill discovery methods have garnered much interest from the community, but little research has been conducted in understanding their limitations. Through theoretical analysis and empirical evidence, we show that existing algorithms suffer from a common limitation – they discover options that provide a poor coverage of the state space. In light of this, we propose Explore, Discover and Learn (EDL), an alternative approach to information-theoretic skill discovery. Crucially, EDL optimizes the same information-theoretic objective derived from the empowerment literature, but addresses the optimization problem using different machinery. We perform an extensive evaluation of skill discovery methods on controlled environments and show that EDL offers significant advantages, such as overcoming the coverage problem, reducing the dependence of learned skills on the initial state, and allowing the user to define a prior over which behaviors should be learned. Victor Campos 0001, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i-Nieto, Jordi Torres |
ICML | 6 |
| 2020 | Mask-guided sample selection for semi-supervised instance segmentation
Miriam Bellver, Amaia Salvador, Jordi Torres, Xavier Giró-i-Nieto |
Multim. Tools Appl. | 3 |
| 2019 | The OTree: Multidimensional Indexing with efficient data Sampling for HPCabstractSpatial big data is considered an essential trend in future scientific and business applications. Indeed, research instruments, medical devices, and social networks generate hundreds of petabytes of spatial data per year. However, many authors have pointed out that the lack of specialized frameworks for multidimensional Big Data is limiting possible applications and precluding many scientific breakthroughs. Paramount in achieving High-Performance Data Analytics is to optimize and reduce the I/O operations required to analyze large data sets. To do so, we need to organize and index the data according to its multidimensional attributes. At the same time, to enable fast and interactive exploratory analysis, it is vital to generate approximate representations of large datasets efficiently. In this paper, we propose the Outlook Tree (or OTree), a novel Multidimensional Indexing with efficient data Sampling (MIS) algorithm. The OTree enables exploratory analysis of large multidimensional datasets with arbitrary precision, a vital missing feature in current distributed data management solutions. Our algorithm reduces the indexing overhead and achieves high performance even for write-intensive HPC applications. Indeed, we use the OTree to store the scientific results of a study on the efficiency of drug inhalers. Then we compare the OTree implementation on Apache Cassandra, named Qbeast, with PostgreSQL and plain storage. Lastly, we demonstrate that our proposal delivers better performance and scalability. Cesare Cugnasco, Hadrien Calmet, Pol Santamaria, Raül Sirvent, Beatriz Eguzkitza, Guillaume Houzeaux, Yolanda Becerra 0001, Jordi Torres, Jesús Labarta |
IEEE BigData | 8 |
| 2019 | Development of a Gestational Diabetes Computer Interpretable Guideline using Semantic Web TechnologiesabstractThe benefits of following Clinical Practice Guidelines (CPGs) in the daily practice of medicine have been widely studied, being a powerful method for standardization and improvement of medical care quality. However, applying these guidelines to promote evidence-based and up-to-date clinical practice is a known challenge due to the lack of digitalization of clinical guidelines. In order to overcome this issue, the use of Clinical Decision Support Systems (CDSS) has been promoted in clinical centres. Nevertheless, CPGs must be formalized in a computer interpretable way to be implemented within CDSS. Moreover, these systems are usually developed and implemented using local setups, and hence local terminologies, which causes lack of semantic interoperability. In this context, the implementation of Semantic Web Technologies (SWTs) to formalize the concepts used in guidelines promotes the interoperability and standardization of those systems. In this paper, an architecture that allows the formalization of CPGs into Computer Interpretable Guidelines (CIGs) supported by an ontology in the gestational diabetes domain is presented. This CIG has been implemented within a CDSS and a mobile application has been developed for guiding patients based on up-to-date evidence based clinical guidelines. Garazi Artola, Jordi Torres, Nekane Larburu, Roberto Álvarez 0001, Naiara Muro |
KEOD | 2 |
| 2019 | Development and Usability Assessment of a Semantically Validated Guideline-Based Patient-Oriented Gestational Diabetes Mobile App
Garazi Artola, Jordi Torres, Nekane Larburu, Roberto Álvarez 0001, Naiara Muro |
IC3K | 2 |
| 2019 | Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial NetworksabstractSpeech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial Network (GAN) with raw speech input. We propose a deep neural network that is trained from scratch in an end-to-end fashion, generating a face directly from the raw speech waveform without any additional identity information (e.g reference image or one-hot encoding). Our model is trained in a self-supervised approach by exploiting the audio and visual signals naturally aligned in videos. With the purpose of training from video data, we present a novel dataset collected for this work, with high-quality videos of youtubers with notable expressiveness in both the speech and visual signals. Amanda Cardoso Duarte, Francisco Roldan, Miquel Tubau, Janna Escur, Santiago Pascual, Amaia Salvador, Eva Mohedano, Kevin McGuinness, Jordi Torres, Xavier Giró-i-Nieto |
ICASSP | 9 |
| 2018 | Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
Victor Campos 0001, Brendan Jou, Xavier Giró-i-Nieto, Jordi Torres, Shih-Fu Chang |
ICLR (Poster) | 4 |
| 2018 | A Left-to-Right Algorithm for Likelihood Estimation in Gamma-Poisson Factor Analysis
Joan Capdevila, Jesús Cerquides, Jordi Torres, François Petitjean, Wray L. Buntine |
ECML/PKDD (2) | 3 |
| 2018 | Mining urban events from the tweet stream through a probabilistic mixture model
Joan Capdevila, Jesús Cerquides, Jordi Torres |
Data Min. Knowl. Discov. | 3 |
| 2017 | Scaling a Convolutional Neural Network for classification of Adjective Noun Pairs with TensorFlow on GPU ClustersabstractDeep neural networks have gained popularity inrecent years, obtaining outstanding results in a wide range ofapplications such as computer vision in both academia andmultiple industry areas. The progress made in recent years cannotbe understood without taking into account the technologicaladvancements seen in key domains such as High PerformanceComputing, more specifically in the Graphic Processing Unit(GPU) domain. These kind of deep neural networks need massiveamounts of data to effectively train the millions of parametersthey contain, and this training can take up to days or weeksdepending on the computer hardware we are using. In thiswork, we present how the training of a deep neural networkcan be parallelized on a distributed GPU cluster. The effect ofdistributing the training process is addressed from two differentpoints of view. First, the scalability of the task and its performancein the distributed setting are analyzed. Second, the impact ofdistributed training methods on the training times and finalaccuracy of the models is studied. We used TensorFlow on top ofthe GPU cluster of servers with 2 K80 GPU cards, at BarcelonaSupercomputing Center (BSC). The results show an improvementfor both focused areas. On one hand, the experiments showpromising results in order to train a neural network faster. The training time is decreased from 106 hours to 16 hoursin our experiments. On the other hand we can observe howincreasing the numbers of GPUs in one node rises the throughput, images per second, in a near-linear way. Morever an additionaldistributed speedup of 10.3 is achieved with 16 nodes taking asbaseline the speedup of one node. Victor Campos 0001, Francesc Sastre, Maurici Yagües, Jordi Torres, Xavier Giró-i-Nieto |
CCGrid | 4 |
| 2017 | Tweet-SCAN: An event discovery technique for geo-located tweets
Joan Capdevila, Jesús Cerquides, Jordi Nin, Jordi Torres |
Pattern Recognit. Lett. | 4 |
| 2017 | Dynamic Configuration of Partitioning in Spark ApplicationsabstractSpark has become one of the main options for large-scale analytics running on top of shared-nothing clusters. This work aims to make a deep dive into the parallelism configuration and shed light on the behavior of parallel spark jobs. It is motivated by the fact that running a Spark application on all the available processors does not necessarily imply lower running time, while may entail waste of resources. We first propose analytical models for expressing the running time as a function of the number of machines employed. We then take another step, namely to present novel algorithms for configuring dynamic partitioning with a view to minimizing resource consumption without sacrificing running time beyond a user-defined limit. The problem we target is NP-hard. To tackle it, we propose a greedy approach after introducing the notions of dependency graphs and of the benefit from modifying the degree of partitioning at a stage; complementarily, we investigate a randomized approach. Our polynomial solutions are capable of judiciously use the resources that are potentially at user's disposal and strike interesting trade-offs between running time and resource consumption. Their efficiency is thoroughly investigated through experiments based on real execution data. Anastasios Gounaris, Georgia Kougka, Rubén Tous, Carlos Tripiana, Jordi Torres |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | User-generated content curation with deep convolutional neural networksabstractIn this paper, we report a work consisting in using deep convolutional neural networks (CNNs) for curating and filtering photos posted by social media users (Instagram and Twitter). The final goal is to facilitate searching and discovering user-generated content (UGC) with potential value for digital marketing tasks. The images are captured in real time and automatically annotated with multiple CNNs. Some of the CNNs perform generic object recognition tasks while others perform what we call visual brand identity recognition. We report experiments with 5 real brands in which more than 1 million real images were analyzed. In order to speed-up the training of custom CNNs we applied a transfer learning strategy. Rubén Tous, Otto Wüst, Mauro Gomez, Jonatan Poveda, Marc Elena, Jordi Torres, Mouna Makni, Eduard Ayguadé |
IEEE BigData | 6 |
| 2016 | Scaling DBSCAN-like Algorithms for Event Detection Systems in Twitter
Joan Capdevila, Gonzalo Pericacho, Jordi Torres, Jesús Cerquides |
ICA3PP | 3 |
| 2015 | Spark deployment and performance evaluation on the MareNostrum supercomputerabstractIn this paper we present a framework to enable data-intensive Spark workloads on MareNostrum, a petascale supercomputer designed mainly for compute-intensive applications. As far as we know, this is the first attempt to investigate optimized deployment configurations of Spark on a petascale HPC setup. We detail the design of the framework and present some benchmark data to provide insights into the scalabilityof the system. We examine the impact of different configurations including parallelism, storage and networking alternatives, and we discuss several aspects in executing Big Data workloads on a computing system that is based on the compute-centric paradigm. Further, we derive conclusions aiming to pave the way towards systematic and optimized methodologies for fine-tuning data-intensive application on large clusters emphasizing on parallelism configurations. Rubén Tous, Anastasios Gounaris, Carlos Tripiana, Jordi Torres, Sergi Girona, Eduard Ayguadé, Jesús Labarta, Yolanda Becerra 0001, David Carrera 0001, Mateo Valero |
IEEE BigData | 4 |
| 2015 | Experiences of Using Cassandra for Molecular Dynamics SimulationsabstractIn response to the requirements of applications that work with large amounts of data, various NoSQL databases have appeared to deal specifically with these challenges. These systems have become popular in environments such as data analytics and OLTP, however these are not the only data-intensive applications that can benefit from these databases. In the life sciences domain, there are many applications that still use flat files as a medium to store data, and they see themselves very limited in terms of scalability and performance, as well as code complexity. We present an analysis on the viability of using these databases for applications with data demands that differ in some of the characteristics from what these systems were originally designed for. By using these databases, we can also observe that the design of the data model, queries and other configuration parameters can have a considerable impact on performance, thus we present examples of different data and system configurations to analyse their effects on performance. With the executions that are presented in this paper we can see performance gaps of a factor of up to almost 5 between using different models, queries and configuration parameters. Roger Hernandez, Cesare Cugnasco, Yolanda Becerra 0001, Jordi Torres, Eduard Ayguadé |
PDP | 4 |
| 2015 | Matching renewable energy supply and demand in green datacenters
Íñigo Goiri, Md. Enamul Haque, Kien Le, Ryan Beauchea, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
Ad Hoc Networks | 7 |
| 2014 | ALOJA: A systematic study of Hadoop deployment variables to enable automated characterization of cost-effectivenessabstractThis article presents the ALOJA project, an initiative to produce mechanisms for an automated characterization of cost-effectiveness of Hadoop deployments and reports its initial results. ALOJA is the latest phase of a long-term collaborative engagement between BSC and Microsoft which, over the past 6 years has explored a range of different aspects of computing systems, software technologies and performance profiling. While during the last 5 years, Hadoop has become the de-facto platform for Big Data deployments, still little is understood of how the different layers of the software and hardware deployment options affects its performance. Early ALOJA results show that Hadoop's runtime performance, and therefore its price, are critically affected by relatively simple software and hardware configuration choices e.g., number of mappers, compression, or volume configuration. Project ALOJA presents a vendor-neutral repository featuring over 5000 Hadoop runs, a test bed, and tools to evaluate the cost-effectiveness of different hardware, parameter tuning, and Cloud services for Hadoop. As few organizations have the time or performance profiling expertise, we expect our growing repository will benefit Hadoop customers to meet their Big Data application needs. ALOJA seeks to provide both knowledge and an online service to with which users make better informed configuration choices for their Hadoop compute infrastructure whether this be on-premise or cloud-based. The initial version of ALOJA's Web application and sources are available at http://hadoop.bsc.es Nicolás Poggi, David Carrera 0001, Aaron Call, Sergio Mendoza, Yolanda Becerra 0001, Jordi Torres, Eduard Ayguadé, Fabrizio Gagliardi, Jesús Labarta, Rob Reinauer, Nikola Vujic, Daron Green, José A. Blakeley |
IEEE BigData | 6 |
| 2014 | Adaptive MapReduce Scheduling in Shared EnvironmentsabstractIn this paper we present a MapReduce task scheduler for shared environments in which MapReduce is executed along with other resource-consuming workloads, such as transactional applications. All workloads may potentially share the same data store, some of them consuming data for analytics purposes while others acting as data generators. This kind of scenario is becoming increasingly important in data centers where improved resource utilization can be achieved through workload consolidation, and is specially challenging due to the interaction between workloads of different nature that compete for limited resources. The proposed scheduler aims to improve resource utilization across machines while observing completion time goals. Unlike other MapReduce schedulers, our approach also takes into account the resource demands for non-MapReduce workloads, and assumes that the amount of resources made available to the MapReduce applications is variable over time. As shown in our experiments, our proposal improves the management of MapReduce jobs in the presence of variable resource availability, increasing the accuracy of the estimations made by the scheduler, thus improving completion time goals without an impact on the fairness of the scheduler. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Malgorzata Steinder |
CCGRID | 4 |
| 2014 | Profit-aware cloud resource provisioner for ecommerceabstractIn recent years, the Cloud Computing paradigm has proven effective in scaling dynamically the number of servers according to simple performance metrics and the incoming workload. However while some applications are able to scale-out, as current scaling metrics do not relate system performance to sales, hosting costs and profits are not optimized completely. The following article proposes a novel technique for dynamic resource provisioning based on revenue and cost metrics, to optimize profits for online retailers in the Cloud. The proposal relies on user behavior models that relate Quality-of-Service (QoS) to service capacity, and to the intention of users to buy a product on an Ecommerce site. We show how such metrics can enable profit-aware resource management by setting an optimal number of servers at each time of the day. Experiments are performed on custom, real-life datasets from an Ecommerce retailer contain over two years of access, performance, and sales data from popular travelWeb applications. Nicolás Poggi, David Carrera 0001, Eduard Ayguadé, Jordi Torres |
CLUSTER | 4 |
| 2014 | Building Green Cloud Services at Low CostabstractInterest in powering data enters at least partially using on-site renewable sources, e.g. solar or wind, has been growing. In fact, researchers have studied distributed services comprising networks of such "green" data centers, and load distribution approaches that "follow the renewables" to maximize their use. However, prior works have not considered where to site such a network for efficient production of renewable energy, while minimizing both data center and renewable plant building costs. Moreover, researchers have not built real load management systems for follow-the-renewables services. Thus, in this paper, we propose a framework, optimization problem, and solution approach for sitting and provisioning green data centers for a follow-the-renewables HPC cloud service. We illustrate the location selection tradeoffs by quantifying the minimum cost of achieving different amounts of renewable energy. Finally, we design and implement a system capable of migrating virtual machines across the green data centers to follow the renewables. Among other interesting results, we demonstrate that one can build green HPC cloud services at a relatively low additional cost compared to existing services. Josep Lluís Berral, Íñigo Goiri, Thu D. Nguyen, Ricard Gavaldà, Jordi Torres, Ricardo Bianchini |
ICDCS | 5 |
| 2014 | Towards the Cloudification of the Social Networks Analytics
Daniel Cea, Jordi Nin, Rubén Tous, Jordi Torres, Eduard Ayguadé |
MDAI | 4 |
| 2013 | Power-Aware Multi-data Center Management Using Machine LearningabstractThe cloud relies upon multi-data center (multi-DC) infrastructures distributed along the world, where people and enterprises pay for resources to offer their web-services to worldwide clients. Intelligent management is required to automate and manage these infrastructures, as the amount of resources and data to manage exceeds the capacities of human operators. Also, it must take into account the cost of running the resources (energy) and the quality of service towards web-services and clients. (De-)consolidation and priming proximity to clients become two main strategies to allocate resources and properly place these web-services in the multi-DC network. Here we present a mathematical model to describe the scheduling problem given web-services and hosts across a multi-DC system, enhancing the decision makers with models for the system behavior obtained using machine learning. After running the system on real DC infrastructures we see that the model drives web-services to the best locations given quality of service, energy consumption, and client proximity, also (de-)consolidating according to the resources required for each web-service given its load. Josep Lluís Berral, Ricard Gavaldà, Jordi Torres |
ICPP | 3 |
| 2013 | Enabling Distributed Key-Value Stores with Low Latency-Impact Snapshot SupportabstractCurrent distributed key-value stores generally provide greater scalability at the expense of weaker consistency and isolation. However, additional isolation support is becoming increasingly important in the environments in which these stores are deployed, where different kinds of applications with different needs are executed, from transactional workloads to data analytics. While fully-fledged ACID support may not be feasible, it is still possible to take advantage of the design of these data stores, which often include the notion of multiversion concurrency control, to enable them with additional features at a much lower performance cost and maintaining its scalability and availability. In this paper we explore the effects that additional consistency guarantees and isolation capabilities may have on a state of the art key-value store: Apache Cassandra. We propose and implement a new multiversioned isolation level that provides stronger guarantees without compromising Cassandra's scalability and availability. As shown in our experiments, our version of Cassandra allows Snapshot Isolation-like transactions, preserving the overall performance and scalability of the system. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé, Mike Spreitzer, Malgorzata Steinder |
NCA | 4 |
| 2013 | Deadline-Based MapReduce Workload ManagementabstractThis paper presents a scheduling technique for multi-job MapReduce workloads that is able to dynamically build performance models of the executing workloads, and then use these models for scheduling purposes. This ability is leveraged to adaptively manage workload performance while observing and taking advantage of the particulars of the execution environment of modern data analytics applications, such as hardware heterogeneity and distributed storage. The technique targets a highly dynamic environment in which new jobs can be submitted at any time, and in which MapReduce workloads share physical resources with other workloads. Thus the actual amount of resources available for applications can vary over time. Beyond the formulation of the problem and the description of the algorithm and technique, a working prototype (called Adaptive Scheduler) has been implemented. Using the prototype and medium-sized clusters (of the order of tens of nodes), the following aspects have been studied separately: the scheduler's ability to meet high-level performance goals guided only by user-defined completion time goals; the scheduler's ability to favor data-locality in the scheduling algorithm; and the scheduler's ability to deal with hardware heterogeneity, which introduces hardware affinity and relative performance characterization for those applications that can benefit from executing on specialized processors. Jorda Polo, Yolanda Becerra 0001, David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2012 | GreenHadoop: leveraging green energy in data-processing frameworksabstractInterest has been growing in powering datacenters (at least partially) with renewable or "green" sources of energy, such as solar or wind. However, it is challenging to use these sources because, unlike the "brown" (carbon-intensive) energy drawn from the electrical grid, they are not always available. This means that energy demand and supply must be matched, if we are to take full advantage of the green energy to minimize brown energy consumption. In this paper, we investigate how to manage a datacenter's computational workload to match the green energy supply. In particular, we consider data-processing frameworks, in which many background computations can be delayed by a bounded amount of time. We propose GreenHadoop, a MapReduce framework for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenHadoop predicts the amount of solar energy that will be available in the near future, and schedules the MapReduce jobs to maximize the green energy consumption within the jobs' time bounds. If brown energy must be used to avoid time bound violations, GreenHadoop selects times when brown energy is cheap, while also managing the cost of peak brown power consumption. Our experimental results demonstrate that GreenHadoop can significantly increase green energy consumption and decrease electricity cost, compared to Hadoop. Íñigo Goiri, Kien Le, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
EuroSys | 5 |
| 2012 | Energy accounting for shared virtualized environments under DVFS using PMC-based power models
Ramon Bertran Monfort, Yolanda Becerra 0001, David Carrera 0001, Vicenç Beltran 0001, Marc González 0001, Xavier Martorell, Nacho Navarro, Jordi Torres, Eduard Ayguadé |
Future Gener. Comput. Syst. | 8 |
| 2012 | Energy-efficient and multifaceted resource management for profit-driven virtualized data centers
Íñigo Goiri, Josep Lluís Berral, Josep Oriol Fitó, Ferran Julià, Ramon Nou, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
Future Gener. Comput. Syst. | 8 |
| 2012 | Autonomic Placement of Mixed Batch and Transactional WorkloadsabstractTo reduce the cost of infrastructure and electrical energy, enterprise datacenters consolidate workloads on the same physical hardware. Often, these workloads comprise both transactional and long-running analytic computations. Such consolidation brings new performance management challenges due to the intrinsically different nature of a heterogeneous set of mixed workloads, ranging from scientific simulations to multitier transactional applications. The fact that such different workloads have different natures imposes the need for new scheduling mechanisms to manage collocated heterogeneous sets of applications, such as running a web application and a batch job on the same physical server, with differentiated performance goals. In this paper, we present a technique that enables existing middleware to fairly manage mixed workloads: long running jobs and transactional applications. Our technique permits collocation of the workload types on the same physical hardware, and leverages virtualization control mechanisms to perform online system reconfiguration. In our experiments, including simulations as well as a prototype system built on top of state-of-the-art commercial middleware, we demonstrate that our technique maximizes mixed workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2011 | Intelligent Placement of Datacenters for Internet ServicesabstractPopular Internet services are hosted by multiple geographically distributed data centers. The location of the data centers has a direct impact on the services' response times, capital and operational costs, and (indirect) carbon dioxide emissions. Selecting a location involves many important considerations, including its proximity to population centers, power plants, and network backbones, the source of the electricity in the region, the electricity, land, and water prices at the location, and the average temperatures at the location. As there can be many potential locations and many issues to consider for each of them, the selection process can be extremely involved and time-consuming. In this paper, we focus on the selection process and its automation. Specifically, we propose a framework that formalizes the process as a non-linear cost optimization problem, and approaches for solving the problem. Based on the framework, we characterize areas across the United States as potential locations for data centers, and delve deeper into seven interesting locations. Using the framework and our solution approaches, we illustrate the selection trade offs by quantifying the minimum cost of (1) achieving different response times, availability levels, and consistency times, and (2) restricting services to green energy and chiller-less data centers. Among other interesting results, we demonstrate that the intelligent placement of data centers can save millions of dollars under a variety of conditions. We also demonstrate that the selection process is most efficient and accurate when it uses a novel combination of linear programming and simulated annealing. Íñigo Goiri, Kien Le, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
ICDCS | 4 |
| 2011 | Optimal Resource Allocation in a Virtualized Software Aging Platform with Software RejuvenationabstractNowadays, virtualized platforms have become the most popular option to deploy complex enough services. The reason is that virtualization allows resource providers to increase resource utilization. Deployed services are expected to be always available, but these long-running services are especially sensitive to suffer from software aging phenomenon. This term refers to an accumulation of errors, which usually causes resource exhaustion, and eventually makes the service hang/crash. To counteract this phenomenon, a preventive approach to fault management, called software rejuvenation has been proposed. In this paper, we propose a framework which provides transparent and predictive software rejuvenation to web services that suffer software aging on virtualized platforms, achieving high levels of availability. To exploit the provider resources, the framework also seeks to maximize the number of services running simultaneously on the platform, while guaranteeing the resources needed by each service. Javier Alonso 0001, Íñigo Goiri, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
ISSRE | 5 |
| 2011 | Resource-Aware Adaptive Scheduling for MapReduce Clusters
Jorda Polo, Claris Castillo, David Carrera 0001, Yolanda Becerra 0001, Ian Whalley, Malgorzata Steinder, Jordi Torres, Eduard Ayguadé |
Middleware | 7 |
| 2011 | GreenSlot: scheduling energy consumption in green datacentersabstractIn this paper, we propose GreenSlot, a parallel batch job scheduler for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenSlot predicts the amount of solar energy that will be available in the near future, and schedules the workload to maximize the green energy consumption while meeting the jobs' deadlines. If grid energy must be used to avoid deadline violations, the scheduler selects times when it is cheap. Our results for production scientific workloads demonstrate that Green-Slot can increase green energy consumption by up to 117% and decrease energy cost by up to 39%, compared to a conventional scheduler. Based on these positive results, we conclude that green datacenters and green-energy-aware scheduling can have a significant role in building a more sustainable IT ecosystem. Íñigo Goiri, Ryan Beauchea, Kien Le, Thu D. Nguyen, Md. Enamul Haque, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
SC | 7 |
| 2011 | A path to achieving a self-managed Grid middleware
Ramon Nou, Ferran Julià, Kevin Hogan, Jordi Torres |
Future Gener. Comput. Syst. | 4 |
| 2010 | Characterizing Cloud Federation for Enhancing Providers' ProfitabstractCloud federation has been proposed as a new paradigm that allows providers to avoid the limitation of owning only a restricted amount of resources, which forces them to reject new customers when they have not enough local resources to fulfill their customers' requirements. Federation allows a provider to dynamically outsource resources to other providers in response to demand variations. It also allows a provider that has underused resources to rent part of them to other providers. Both things could make the provider to get more profit when used adequately. This requires that the provider has a clear understanding of the potential of each federation decision, in order to choose the most convenient depending on the environment conditions. In this paper, we present a complete characterization of providers' federation in the Cloud, including decision equations to outsource resources to other providers, rent free resources to other providers (i.e. insourcing), or shutdown unused nodes to save power, and we characterize these decisions as a function of several parameters. Then, we demonstrate in the evaluation section how a provider can enhance its profit by using these equations to exploit federation, and how the different parameters influence which is the best decision on each situation. Íñigo Goiri, Jordi Guitart, Jordi Torres |
IEEE CLOUD | 3 |
| 2010 | Energy-Aware Scheduling in Virtualized DatacentersabstractThe reduction of energy consumption in large-scale datacenters is being accomplished through an extensive use of virtualization, which enables the consolidation of multiple workloads in a smaller number of machines. Nevertheless, virtualization also incurs some additional overheads (e.g. virtual machine creation and migration) that can influence what is the best consolidated configuration, and thus, they must be taken into account. In this paper, we present a dynamic job scheduling policy for power-aware resource allocation in a virtualized datacenter. Our policy tries to consolidate workloads from separate machines into a smaller number of nodes, while fulfilling the amount of hardware resources needed to preserve the quality of service of each job. This allows turning off the spare servers, thus reducing the overall datacenter power consumption. As a novelty, this policy incorporates all the virtualization overheads in the decision process. In addition, our policy is prepared to consider other important parameters for a datacenter, such as reliability or dynamic SLA enforcement, in a synergistic way with power consumption. The introduced policy is evaluated comparing it against common policies in a simulated environment that accurately models HPC jobs execution in a virtualized datacenter including power consumption modeling and obtains a power consumption reduction of 15% with respect to typical policies. Íñigo Goiri, Ferran Julià, Ramon Nou, Josep Lluís Berral, Jordi Guitart, Jordi Torres |
CLUSTER | 6 |
| 2010 | Adaptive on-line software aging prediction based on machine learningabstractThe growing complexity of software systems is resulting in an increasing number of software faults. According to the literature, software faults are becoming one of the main sources of unplanned system outages, and have an important impact on company benefits and image. For this reason, a lot of techniques (such as clustering, fail-over techniques, or server redundancy) have been proposed to avoid software failures, and yet they still happen. Many software failures are those due to the software aging phenomena. In this work, we present a detailed evaluation of our chosen machine learning prediction algorithm (M5P) in front of dynamic and non-deterministic software aging. We have tested our prediction model on a three-tier web J2EE application achieving acceptable prediction accuracy against complex scenarios with small training data sets. Furthermore, we have found an interesting approach to help to determine the root cause failure: The model generated by machine learning algorithms. Javier Alonso 0001, Jordi Torres, Josep Lluís Berral, Ricard Gavaldà |
DSN | 2 |
| 2010 | Performance Management of Accelerated MapReduce Workloads in Heterogeneous ClustersabstractNext generation data centers will be composed of thousands of hybrid systems in an attempt to increase overall cluster performance and to minimize energy consumption. New programming models, such as MapReduce, specifically designed to make the most of very large infrastructures will be leveraged to develop massively distributed services. At the same time, data centers will bring an unprecedented degree of workload consolidation, hosting in the same infrastructure distributed services from many different users. In this paper we present our advancements in leveraging the Adaptive MapReduce Scheduler to meet user defined high level performance goals while transparently and efficiently exploiting the capabilities of hybrid systems. While the Adaptive Scheduler was already able to dynamically allocate resources to co-located MapReduce jobs based on their completion time goals, it was completely unaware of specific hardware capabilities. In our work we describe the changes introduced in the Adaptive Scheduler to enable it with hardware awareness and with the ability to co-schedule accelerable and non-accelerable jobs on the same heterogeneous MapReduce cluster, making the most of the underlying hybrid systems. The developed prototype is tested in a cluster of Cell/BE blades and relies on the use of accelerated and non-accelerated versions of the MapReduce tasks of different deployed applications to dynamically select the best version to run on each node. Decisions are made after workload composition and jobs' completion time goals. Results show that the augmented Adaptive Scheduler provides dynamic resource allocation across jobs, hardware affinity when possible, and is even able to spread jobs' tasks across accelerated and non-accelerated nodes in order to meet performance goals in extreme conditions. To our knowledge this is the first MapReduce scheduler and prototype that is able to manage high-level performance goals even in presence of hybrid systems and accelerable jobs. Jorda Polo, David Carrera 0001, Yolanda Becerra 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 5 |
| 2010 | Checkpoint-based fault-tolerant infrastructure for virtualized service providersabstractCrash and omission failures are common in service providers: a disk can break down or a link can fail anytime. In addition, the probability of a node failure increases with the number of nodes. Apart from reducing the provider's computation power and jeopardizing the fulfillment of his contracts, this can also lead to computation time wasting when the crash occurs before finishing the task execution. In order to avoid this problem, efficient checkpoint infrastructures are required, especially in virtualized environments where these infrastructures must deal with huge virtual machine images. This paper proposes a smart checkpoint infrastructure for virtualized service providers. It uses Another Union File System to differentiate read-only from read-write parts in the virtual machine image. In this way, read-only parts can be checkpointed only once, while the rest of checkpoints must only save the modifications in read-write parts, thus reducing the time needed to make a checkpoint. The checkpoints are stored in a Hadoop Distributed File System. This allows resuming a task execution faster after a node crash and increasing the fault tolerance of the system, since checkpoints are distributed and replicated in all the nodes of the provider. This paper presents a running implementation of this infrastructure and its evaluation, demonstrating that it is an effective way to make faster checkpoints with low interference on task execution and efficient task recovery after a node failure. Íñigo Goiri, Ferran Julià, Jordi Guitart, Jordi Torres |
NOMS | 4 |
| 2010 | Exploiting semantics and virtualization for SLA-driven resource allocation in service providersabstractAbstract Resource management is a key challenge that service providers must adequately face in order to accomplish their business goals. This paper introduces a framework, the semantically enhanced resource allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by the service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide an application specific and isolated virtual environment for each task. In addition, the system supports fine‐grain dynamic resource distribution among these virtual environments based on Service‐Level Agreements. The required adaptation is implemented using agents, guarantying enough resources to each task in order to meet the agreed performance goals. Copyright © 2009 John Wiley & Sons, Ltd. Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres |
Concurr. Comput. Pract. Exp. | 7 |
| 2010 | A survey on performance management for internet applicationsabstractAbstract Internet applications have become indispensable for many business and personal processes, turning the performance of these applications into a key issue. For this reason, recent research has comprehensively explored mechanisms for managing the performance of these applications, with special focus on dealing with overload situations and providing QoS guarantees to clients. This paper makes a survey on the different proposals in the literature for managing Internet applications' performance. We present a complete taxonomy that characterizes and classifies these proposals into several categories including request scheduling, admission control, service differentiation, dynamic resource management, service degradation, control theoretic approaches, works using queuing models, observation‐based approaches that use runtime measurements, and overall approaches combining several mechanisms. For each work, we provide a brief description in order to provide the reader with a global understanding of the research progress in this area. Copyright © 2009 John Wiley & Sons, Ltd. Jordi Guitart, Jordi Torres, Eduard Ayguadé |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Maximizing revenue in Grid markets using an economically enhanced resource managerabstractAbstract Traditional resource management has had as its main objective the optimization of throughput, based on parameters such as CPU, memory, and network bandwidth. With the appearance of Grid markets, new variables that determine economic expenditure, benefit and opportunity must be taken into account. The Self‐organizing ICT Resource Management (SORMA) project aims at allowing resource owners and consumers to exploit market mechanisms to sell and buy resources across the Grid. SORMA's motivation is to achieve efficient resource utilization by maximizing revenue for resource providers and minimizing the cost of resource consumption within a market environment. An overriding factor in Grid markets is the need to ensure that the desired quality of service levels meet the expectations of market participants. This paper explains the proposed use of an economically enhanced resource manager (EERM) for resource provisioning based on economic models. In particular, this paper describes techniques used by the EERM to support revenue maximization across multiple service level agreements and provides an application scenario to demonstrate its usefulness and effectiveness. Copyright © 2008 John Wiley & Sons, Ltd. Mario Macías, Omer F. Rana, Garry Smith, Jordi Guitart, Jordi Torres |
Concurr. Comput. Pract. Exp. | 5 |
| 2009 | CellMT: A cooperative multithreading library for the Cell/B.EabstractThe Cell BE processor has proved that heterogeneous multi-core systems can provide a huge computational power with high efficiency for a wide range of applications. The simple design of the computational units and the use of small managed local memories is the key to achieve high efficiency and performance at the same time. However, this simple and efficient hardware design comes at the price of higher code complexity. The code written to run in this kind of processors must deal with several issues such as code vectorization, loop unrolling or the explicit management of local memories. Some of these issues such as vectorization or loop unrolling can be partially solved by the compiler, but the overlapping of data transfer and computation times must be manually addressed by the programmer with techniques such as double buffering that increase the code complexity. In this paper we present a user level threading library called CellMT that effectively hide memory latencies. The concurrent execution of several threads inside each SPU naturally overlaps computation and data transfer times without increasing the code complexity. To prove the suitability and feasibility of our multi-threaded library, we perform an exhaustive performance evaluation with a synthetic benchmark and a real application. The experimental results show that the multithreaded approach can outperform a hand-coded double buffering scheme, with speedups from 0.96x to 3.2x, while maintaining the complexity of a naive buffering scheme. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
HiPC | 3 |
| 2009 | Speeding Up Distributed MapReduce Applications Using Hardware AcceleratorsabstractIn an attempt to increase the performance/cost ratio, large compute clusters are becoming heterogeneous at multiple levels: from asymmetric processors, to different system architectures, operating systems and networks. Exploiting the intrinsic multi-level parallelism present in such a complex execution environment has become a challenging task using traditional parallel and distributed programming models. As a result, an increasing need for novel approaches to exploiting parallelism has arisen in these environments. MapReduce is a data-driven programming model originally proposed by Google back in 2004 as a flexible alternative to the existing models, specially devoted to hiding the complexity of both developing and running massively distributed applications in large compute clusters. In some recent works, the MapReduce model has been also used to exploit parallelism in other non-distributed environments, such as multi-cores, heterogeneous processors and GPUs. In this paper we introduce a novel approach for exploiting the heterogeneity of a Cell BE cluster linking an existing MapReduce runtime implementation for distributed clusters and one runtime to exploit the parallelism of the Cell BE nodes. The novel contribution of this work is the design and evaluation of a MapReduce execution environment that effectively exploits the parallelism existing at both the Cell BE cluster level and the heterogeneous processors level. Yolanda Becerra 0001, Vicenç Beltran 0001, David Carrera 0001, Marc González 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 5 |
| 2009 | Introducing Virtual Execution Environments for Application Lifecycle Management and SLA-Driven Resource Distribution within Service ProvidersabstractResource management is a key challenge that service providers must adequately face in order to ensure their profitability. This paper describes a proof-of-concept framework for facilitating resource management in service providers, which allows reducing costs and at the same time fulfilling the quality of service agreed with the customers. This is accomplished by means of virtualization. Our approach provides application-specific virtual environments and consolidates them in order to achieve a better utilization of the providers resources. In addition, it implements self-adaptive capabilities for dynamically distributing the providers resources among these virtual environments based on Service Level Agreements. The proposed solution has been implemented as a part of the Semantically-Enhanced Resource Allocator prototype developed within the BREIN European project. The evaluation shows that our prototype is able to react in very short time under changing conditions and avoid SLA violations by rescheduling efficiently the resources. Íñigo Goiri, Ferran Julià, Jorge Ejarque, Marc de Palol, Rosa M. Badia, Jordi Guitart, Jordi Torres |
NCA | 7 |
| 2009 | Self-adaptive utility-based web session management
Nicolás Poggi, Toni Moreno, Josep Lluís Berral, Ricard Gavaldà, Jordi Torres |
Comput. Networks | 5 |
| 2009 | Autonomic QoS control in enterprise Grid environments using online simulation
Ramon Nou, Samuel Kounev, Ferran Julià, Jordi Torres |
J. Syst. Softw. | 4 |
| 2009 | Using Virtualization to Improve Software RejuvenationabstractIn this paper, we present an approach for software rejuvenation based on automated self-healing techniques that can be easily applied to off-the-shelf application servers. Software aging and transient failures are detected through continuous monitoring of system data and performability metrics of the application server. If some anomalous behavior is identified, the system triggers an automatic rejuvenation action. This self-healing scheme is meant to disrupt the running service for a minimal amount of time, achieving zero downtime in most cases. In our scheme, we exploit the usage of virtualization to optimize the self-recovery actions. The techniques described in this paper have been tested with a set of open-source Linux tools and the XEN virtualization middleware. We conducted an experimental study with two application benchmarks (Tomcat/Axis and TPC-W). Our results demonstrate that virtualization can be extremely helpful for failover and software rejuvenation in the occurrence of transient failures and software aging. Luís Moura Silva, Javier Alonso 0001, Jordi Torres |
IEEE Trans. Computers | 3 |
| 2008 | SLA-Driven Semantically-Enhanced Dynamic Resource Allocator for Virtualized Service ProvidersabstractIn order to be profitable, service providers must be able to undertake complex management tasks such as provisioning, deployment, execution and adaptation in an autonomic way. This paper introduces a framework, the Semantically-Enhanced Resource Allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide a full-customized and isolated virtual environment for each task. In addition, the system supports fine-grain dynamic resource distribution among these virtual environments based on SLAs. The required adaptation is implemented using agents, guarantying to each task enough resources to meet the agreed performance goals. Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres |
eScience | 7 |
| 2008 | Managing SLAs of heterogeneous workloads using dynamic application placementabstractIn this paper we address the problem of managing heterogeneous workloads in a virtualized data center. We consider two different workloads: transactional applications and long-running jobs. We present a technique that permits collocation of these workload types on the same physical hardware. Our technique dynamically modifies workload placement by leveraging control mechanisms such as suspension and migration, and strives to optimally trade off resource allocation among these workloads in spite of their differing characteristics and performance objectives. Our approach builds upon our previous work on dynamically placing transactional workloads. This paper extends our framework with the capability to manage long-running workloads. We achieve this goal by using utility functions, which permit us to compare the performance of various workloads, and which are used to drive allocation decisions. We demonstrate that our technique maximizes heterogeneous workload performance while providing service differentiation based on high-level performance goals. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
HPDC | 4 |
| 2008 | Improving Web Server Performance Through Main Memory CompressionabstractCurrent web servers are highly multithreaded applications whose scalability benefits from the current multi-core/multiprocessor trend. However, some workloads cannot capitalize on this because their performance is limited by the available memory and/or the disk bandwidth, which prevents the server from taking advantage of the computing resources provided by the system. To solve this situation we propose the use of main memory compression techniques to increment the available memory and mitigate the disk band-width problem, allowing the web server to improve its use of CPU system resources. In this paper we implement to the Linux OS a full SMP capable main memory compression subsystem to increase the performance of a web server running the SPEC web 2005 benchmark. Although main memory compression is not a new technique perse, its use in a multicore environment running heavily multithreaded applications like a webserver introduces new challenges in the technique, such as scalability issues and the trade-off between the compressed memory size and the computational power required to achieve it. Finally, the evaluation of our implementaiton shows promising results such as a 30% web server throughput improvement and a 70% reduction in the disk bandwidth usage. Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPADS | 2 |
| 2008 | Understanding tuning complexity in multithreaded and hybrid web serversabstractAdequately setting up a multi-threaded Web server is a challenging task because its performance is determined by a combination of configurable Web server parameters and unsteady external factors like the workload type, workload intensity and machine resources available. Usually administrators set up Web server parameters like the keep-alive timeout and number of worker threads based on their experience and judgment, expecting that this configuration will perform well for the guessed uncontrollable factors. The nontrivial interaction between the configuration parameters of a multi-threaded Web server makes it a hard task to properly tune it for a given workload, but the burst nature of the Internet quickly change the uncontrollable factors and make it impossible to obtain an optimal configuration that will always perform well. In this paper we show the complexity of optimally configuring a multi-threaded Web server for different workloads with an exhaustive study of the interactions between the keep-alive timeout and the number of worker threads for a wide range of workloads. We also analyze the Hybrid Web server architecture (multi-threaded and event-driven) as a feasible solution to simplify Web server tuning and obtain the best performance for a wide range of workloads that can dynamically change in intensity and type. Finally, we compare the performance of the optimally tuned multithreaded Web server and the hybrid Web server with different workloads to validate our assertions. We conclude from our study that the hybrid architecture clearly outperforms the multi-threaded one, not only in terms of performance, but also in terms of its tuning complexity and its adaptability over different workload types. In fact, from the obtained results, we expect that the hybrid architecture is well suited to simplify the self configuration of complex application servers. Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
IPDPS | 2 |
| 2008 | Reducing wasted resources to help achieve green data centersabstractIn this paper we introduce a new approach to the consolidation strategy of a data center that allows an important reduction in the amount of active nodes required to process a heterogeneous workload without degrading the offered service level. This article reflects and demonstrates that consolidation of dynamic workloads does not end with virtualization. If energy-efficiency is pursued, the workloads can be consolidated even more using two techniques, memory compression and request discrimination, which were separately studied and validated in previous work and are now to be combined in a joint effort. We evaluate the approach using a representative workload scenario composed of numerical applications and a real workload obtained from a top national travel website. Our results indicate that an important improvement can be achieved using 20% less servers to do the same work. We believe that this serves as an illustrative example of a new way of management: tailoring the resources to meet high level energy efficiency goals. Jordi Torres, David Carrera 0001, Kevin Hogan, Ricard Gavaldà, Vicenç Beltran 0001, Nicolás Poggi |
IPDPS | 1 |
| 2008 | Enabling Resource Sharing between Transactional and Batch Workloads Using Dynamic Application Placement
David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
Middleware | 4 |
| 2008 | Utility-based placement of dynamic Web applications with fairness goalsabstractWe study the problem of dynamic resource allocation to clustered Web applications. We extend application server middleware with the ability to automatically decide the size of application clusters and their placement on physical machines. Unlike existing solutions, which focus on maximizing resource utilization and may unfairly treat some applications, the approach introduced in this paper considers the satisfaction of each application with a particular resource allocation and attempts to at least equally satisfy all applications. We model satisfaction using utility functions, mapping CPU resource allocation to the performance of an application relative to its objective. The demonstrated online placement technique aims at equalizing the utility value across all applications while also satisfying operational constraints, preventing the over-allocation of memory, and minimizing the number of placement changes. We have implemented our technique in a leading commercial middleware product. Using this real-life testbed and a simulation we demonstrate the benefit of the utility-driven technique as compared to other state-of-the-art techniques. David Carrera 0001, Malgorzata Steinder, Ian Whalley, Jordi Torres, Eduard Ayguadé |
NOMS | 4 |
| 2008 | Dynamic CPU provisioning for self-managed secure web applications in SMP hosting platforms
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 4 |
| 2007 | Should the grid middleware look to self-managing capabilities?abstractGrid technologies have enabled the clustering of a wide variety of geographically distributed resources and services. While providing this, the grid middleware layers that make up the supporting platform for grid applications have taken an increased level of importance in the overall performance of such distributed applications. In this paper we discuss the necessity of introducing self-managing capabilities to the core functionalities of the grid middleware. Our opinions are based on a simple example that shows how Globus Toolkit 4 (GT4) can be driven to a state of unavailability under certain overloading conditions, and how a simple but effective self-managing policy applied to its resource management mechanisms could overcome such an unfavourable scenario. We present an approach showing the benefits that a resource management introduced in the middleware can provide to the user. This approach is the base of a self-managing layer that is being developed for use under more generic conditions Ramon Nou, Ferran Julià, Jordi Torres |
ISADS | 3 |
| 2007 | Using Virtualization to Improve Software RejuvenationabstractIn this paper, we present an approach for software rejuvenation based on automated self-healing techniques that can be easily applied to off-the-shelf Application Servers and Internet sites. Software aging and transient failures are detected through continuous monitoring of system data and performability metrics of the application server. If some anomalous behavior is identified the system triggers an automatic rejuvenation action. This self-healing scheme is meant to be the less disruptive as possible for the running service and to get a zero downtime for most of the cases. In our scheme, we exploit the usage of virtualization to optimize the self-recovery actions. The techniques described in this paper have been tested with a set of open-source Linux tools and the XEN virtualization middleware. We conducted an experimental study with two applications benchmarks (Tomcat/Axis and TPC-W). Our results demonstrate that virtualization can be extremely helpful for software rejuvenation and fail-over in the occurrence of transient application failures and software aging. Luís Moura Silva, Javier Alonso 0001, Jordi Torres, Artur Andrzejak 0001 |
NCA | 4 |
| 2007 | Differentiated Quality of Service for e-Commerce Applications through Connection Scheduling based on System-Level Thread PrioritiesabstractThe e-commerce Web sites receive a great and varied number of visitors every day. These visitors share the application server's limited resources and when there are too many clients connecting to the Web site, it is possible that they hinder between them, even to overload the application server. These visitors can be divided in different categories, depending on their importance from site viewpoint. Considering the importance that in these Web sites some client connections (e.g. buyers' connections) finish successfully before other connections, in this paper we propose a mechanism to provide different quality of service to the different client categories by assigning different priorities to the threads attending the connections. After observing that Java thread priorities are only applied within the JVM, and moreover, these priorities do not reach the O.S. threads, we propose to schedule threads using the Linux Real Time priorities. Our results demonstrate that different quality of service classes can be supported using this mechanism Javier Alonso 0001, Jordi Guitart, Jordi Torres |
PDP | 3 |
| 2007 | Monitoring and Analysis Framework for Grid MiddlewareabstractAs the use of complex grid middleware becomes widespread and more facilites are offered by these pieces of software, distributed grid applications are becoming more and more popular. But as grid middleware grows in size and offers more advanced features, they become more complex and heavier, as well as harder to tune. Since the performance of a distributed grid application can be strongly influenced by the operation of the underlying grid middleware, it becomes of extreme importance to study and analyse its behaviour and performance. In this paper we present the eDragon monitoring framework (eDMF), a set of tools that can be used for the instrumentation and analysis of grid middleware, and which provides a unique environment to study the performance of grid applications. The eDMF is composed of a set of specialised monitoring tools as well as by a flexible and powerful performance analysis platform. Additionally we also provide a practical application of the eDMF to the Globus toolkit 4 (GT4), one of the most extended and popular grid middleware, showing how it helped us in the detection and resolution of several job management problems observed in the GT4 middleware Ramon Nou, Ferran Julià, David Carrera 0001, Kevin Hogan, Jordi Caubet, Jesús Labarta, Jordi Torres |
PDP | 7 |
| 2007 | Economically Enhanced Resource Management for Internet Service Utilities
Tim Püschel, Nikolay Borissov, Mario Macías, Dirk Neumann 0001, Jordi Guitart, Jordi Torres |
WISE | 6 |
| 2007 | Designing an overload control strategy for secure e-commerce applications
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 4 |
| 2005 | A Hybrid Web Server Architecture for Secure e-Business Web Applications
Vicenç Beltran 0001, David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé |
HPCC | 4 |
| 2005 | Session-Based Adaptive Overload Control for Secure Dynamic Web ApplicationsabstractAs dynamic Web content and security capabilities are becoming popular in current Web sites, the performance demand on application servers that host the sites is increasing, leading sometimes these servers to overload. As a result, response times may grow to unacceptable levels and the server may saturate or even crash. In this paper we present a session-based adaptive overload control mechanism based on SSL (secure socket layer) connections differentiation and admission control. The SSL connections differentiation is a key factor because the cost of establishing a new SSL connection is much greater than establishing a resumed SSL connection (it reuses an existing SSL session on server). Considering this big difference, we have implemented an admission control algorithm that prioritizes the resumed SSL connections to maximize performance on session-based environments and limits dynamically the number of new SSL connections accepted depending on the available resources and the current number of connections in the system to avoid server overload. In order to allow the differentiation of resumed SSL connections from new SSL connections we propose a possible extension of the Java Secure Sockets Extension (JSSE) API. Our evaluation on Tomcat server demonstrates the benefit of our proposal for preventing server overload. Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 4 |
| 2004 | Evaluating the Scalability of Java Event-Driven Web ServersabstractThe two major strategies used to construct high-performance Web servers are thread pools and event-driven architectures. The Java platform is commonly used in Web environments but up to the moment it did not provide any standard API to implement event-driven architectures efficiently. The new 1.4 release of the J2SE introduces the NIO (New I/O) API to help in the development of event-driven I/O intensive applications. We evaluate the scalability that this API provides to the Java platform in the field of Web servers, bringing together the majorly used commercial server (Apache) and one experimental server developed using the NIO API. We study the scalability of the NIO-based server as well as of its rival in a number of different scenarios, including uniprocessor, multiprocessor, bandwidth-bounded and CPU-bounded environments. The study concludes that the NIO API can be successfully used to create event-driven Java servers that can scale as well as the best of the commercial native-compiled Web server, at a fraction of its complexity and using only one or two worker threads. Vicenç Beltran 0001, David Carrera 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 3 |
| 2003 | Complete instrumentation requirements for performance analysis of Web based technologiesabstractIn this paper we present the eDragon environment, a research platform created to perform complete performance analysis of new Web-based technologies. eDragon enables the understanding of how application servers work in both sequential and parallel platforms offering a new insight in the usage of system resources. The environment is composed of a set of instrumentation modules, a performance analysis and visualization tool and a set of experimental methodologies to perform complete performance analysis of Web-based technologies. This paper describes the design and implementation of this research platform and highlights some of its main functionalities. We will also show how a detailed analytical view can be obtained through the application of a bottom-up strategy, starting with a group of system events and advancing to more complex performance metrics using a continuous derivation process. David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé, Jesús Labarta |
ISPASS | 3 |
| 2001 | Performance Analysis Tools for Parallel Java Applications on Shared-memory SystemsabstractIn this paper we describe an instrumentation environment for the performance analysis and visualization of parallel applications written in JOMP, an OpenMP-like interface for Java. The environment includes two complementary approaches. The first one has been designed to provide a detailed analysis of the parallel behavior at the JOMP programming model level. At this level, the user is faced with parallel, work-sharing and synchronization constructs, which are the core of JOMP. The second mechanism has been designed to support an in-depth analysis of the threaded execution inside the Java virtual machine (JVM). At this level of analysis, the user is faced with the supporting threads layer monitors and conditional variables. The paper discusses the implementation of both mechanisms and evaluates the overhead incurred by them. Jordi Guitart, Jordi Torres, Eduard Ayguadé, J. Mark Bull |
ICPP | 2 |
| 2001 | Strategies for the efficient exploitation of loop-level parallelism in JavaabstractAbstract This paper analyzes the overheads incurred in the exploitation of loop‐level parallelism using Java Threads and proposes some code transformations that minimize them. The transformations avoid the intensive use of Java Threads and reduce the number of classes used to specify the parallelism in the application (which reduces the time for class loading). The use of such transformations results in promising performance gains that may encourage the use of Java for exploiting loop‐level parallelism in the framework of OpenMP. On average, the execution time for our synthetic benchmarks is reduced by 50% from the simplest transformation when eight threads are used. The paper explores some possible enhancements to the Java threading API oriented towards improving the application–runtime interaction. Copyright © 2001 John Wiley & Sons, Ltd. José Oliver 0002, Jordi Guitart, Eduard Ayguadé, Nacho Navarro, Jordi Torres |
Concurr. Comput. Pract. Exp. | 5 |
| 1993 | Partitioning the Statement per Iteration Space Using Non-Singular MatricesabstractIn this paper we generalize the framework of linear loop transformations: we consider loop alignment as a new component in the transformation process. The aim is to exploit the additional inherent statement-level parallelism and reduce the amount of interprocessor synchronization and communication when a coarse-grain MIMD execution model is considered. The transformation process is modelled with non-singular matrices and we use the ideas recently proposed in this field to generate an efficient transformed code. However, additional aspects have to be studied when statements are considered in the process. We try to reduce the overhead due to conditionals that appear in the loop body of the transformed loops. Eduard Ayguadé, Jordi Torres |
International Conference on Supercomputing | 2 |
| 1989 | GTS: parallelization and vectorization of tight recurrencesabstractIn this paper we present a new method for extracting the maximum parallelism or vector operations out of DO loops with tight recurrences using sequential programming languages. We have named the method Graph Traverse Scheduling (GTS). It is devised to produce code for shared memory multiprocessors or vector machines. When parallelizing, hardware support for fast synchronization is assumed. Eduard Ayguadé, Jesús Labarta, Jordi Torres, Patricia Borensztejn |
SC | 3 |