Rosa M. Badia

dblp:88/929 · also Rosa Maria Badia · DBLP profile ↗
← Back
118ranked-venue papers
12as first author
21since 2021 · last 2026
0000-0003-2941-5499ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 95 · 11 first-author · 18 since 2021Software engineering, systems software and programming languages · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 HP2C-DT: High-Precision High-Performance Computer-enabled Digital Twin
E. Iraola, Mauro Garcia Lorenzo, Francesc Lordan, F. Rossi, Eduardo Prieto-Araujo, Rosa M. Badia
Future Gener. Comput. Syst.6
2026 A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva
Future Gener. Comput. Syst.4
2024 Performance Analysis of Distributed GPU-Accelerated Task-Based Workflows
Marcos N. L. Carvalho, Anna Queralt, Oscar Romero 0001, Alkis Simitsis, Cristian Tatu, Rosa M. Badia
EDBT6
2024 GPU Cache System for COMPSs: A Task-Based Distributed Computing Framework
Cristian Tatu, Javier Conejero, Fernando Vázquez-Novoa, Rosa M. Badia
Euro-Par (3)4
2024 Quantum optimization algorithms: Energetic implications
abstract
Summary Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. In this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least to more energy efficient than their classical counterparts; why feedback‐based quantum optimization is energy‐inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of 20. Finally, we use the NchooseK high‐level programming model to run optimization problems on both gate‐based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate‐based counterparts.
Rolando P. Hong Enriquez, Rosa M. Badia, Barbara M. Chapman, Kirk Bresniker, Scott Pakin, Alok Mishra 0002, Pedro Bruel, Aditya Dhakal, Gourav Rattihalli, Ninad Hogade, Eitan Frachtenberg, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.2
2024 Portability and scalability evaluation of large-scale statistical modeling and prediction software through HPC-ready containers
abstract
HPC-based applications often have complex workflows with many software dependencies that hinder their portability on contemporary HPC architectures. In addition, these applications often require extraordinary efforts to deploy and execute at performance potential on new HPC systems, while the users expert in these applications generally have less expertise in HPC and related technologies. This paper provides a dynamic solution that facilitates containerization for transferring HPC software onto diverse parallel systems . The study relies on the HPC Workflow as a Service (HPCWaaS) paradigm proposed by the EuroHPC eFlows4HPC project. It offers to deploy workflows through containers tailored for any of a number of specific HPC systems. Traditional container image creation tools rely on OS system packages compiled for generic architecture families (x86_64, amd64, ppc64, …) and specific MPI or GPU runtime library versions. The containerization solution proposed in this paper leverages HPC Builders such as Spack or Easybuild and multi-platform builders such as buildx to create a service for automating the creation of container images for the software specific to each hardware architecture, aiming to sustain the overall performance of the software. We assess the efficiency of our proposed solution for porting the geostatistics ExaGeoStat software on various parallel systems while preserving the computational performance. The results show that the performance of the generated images is comparable with the native execution of the software on the same architectures. On the distributed-memory system, the containerized version can scale up to 256 nodes without impacting performance.
Sameh Abdulah, Jorge Ejarque, Omar Marzouk, Hatem Ltaief, Ying Sun 0002, Marc G. Genton, Rosa M. Badia, David E. Keyes
Future Gener. Comput. Syst.7
2024 Extreme-scale workflows: A perspective from the JLESC international community
Orcun Yildiz, Amal Gueroudji, Julien Bigot, Bruno Raffin, Rosa M. Badia, Tom Peterka
Future Gener. Comput. Syst.5
2024 Boosting HPC data analysis performance with the ParSoDA-Py library
abstract
Abstract Developing and executing large-scale data analysis applications in parallel and distributed environments can be a complex and time-consuming task. Developers often find themselves diverted from their application logic to handle technical details about the underlying runtime and related issues. To simplify this process, ParSoDA, a Java library, has been proposed to facilitate the development of parallel data mining applications executed on HPC systems. It simplifies the process by providing built-in scalability mechanisms relying on the Hadoop and Spark frameworks. This paper presents ParSoDA-Py, the Python version of the ParSoDA library, which allows for further support of commonly used runtimes and libraries for big data analysis. After a complete library redesign, ParSoDA can be now easily integrated with other Python-based distributed runtimes for HPC systems, such as COMPSs and Apache Spark, and with the large ecosystem of Python-based data processing libraries. The paper discusses the adaptation process, which takes into consideration the new technical requirements, and evaluates both usability and scalability through some case study applications.
Loris Belcastro, Salvatore Giampà, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Rosa M. Badia, Jorge Ejarque, Nihad Mammadli
J. Supercomput.6
2023 Hierarchical Management of Extreme-Scale Task-Based Applications
Francesc Lordan, Gabriel Puigdemunt, Pere Vergés, Javier Conejero, Jorge Ejarque, Rosa M. Badia
Euro-Par6
2023 Scalable Random Forest with Data-Parallel Computing
Fernando Vázquez-Novoa, Javier Conejero, Cristian Tatu, Rosa M. Badia
Euro-Par4
2023 Special Issue 19th international workshop on algorithms, models and tools for parallel computing on heterogeneous platforms (HeteroPar'21)
abstract
Heterogeneity has emerged as one of the most profound and challenging characteristics of today's parallel environments. From the macro level, where networks of distributed computers, composed of diverse node architectures, are interconnected with potentially heterogeneous networks, to the micro level, where deeper memory hierarchies and various accelerator architectures are increasingly common, the impact of heterogeneity on all computing tasks is increasing rapidly. Traditional parallel algorithms, programming environments, and tools, designed for legacy homogeneous multiprocessors, achieve a small fraction of the efficiency and the potential performance that can be obtained in current and future heterogeneous computing platforms. New ideas, innovative algorithms, and specialized programming environments and tools are needed to efficiently use these modern parallel and heterogeneous architectures. The International workshop on algorithms, models and tools for parallel computing on heterogeneous platforms (HeteroPar) is a forum for researchers working on algorithms, programming languages, tools, and theoretical models for efficiently solving complex problems on heterogeneous parallel platforms. HeteroPar 2021 took place (virtually) in Lisbon, Portugal, organized for the 13th time in conjunction with the Euro-Par annual international conference. The format of the workshop included one keynote and 11 technical presentations. The workshop papers represented an interesting mix of topics, addressing the implementation of algorithms and kernels for heterogeneous computing, programming models, data management, runtime and resource management, energy efficiency, cloud computing, and artificial intelligence-based methods oriented towards heterogeneous platforms, as the basis for the next generation exascale computers. The selected papers cover a good spectrum of research topics in the area of heterogeneous computing showing the challenges present on these modern platforms.
Rosa M. Badia
Concurr. Comput. Pract. Exp.1
2023 The EU Center of Excellence for Exascale in Solid Earth (ChEESE): Implementation, results, and roadmap for the second phase
abstract
The EU Center of Excellence for Exascale in Solid Earth (ChEESE) develops exascale transition capabilities in the domain of Solid Earth, an area of geophysics rich in computational challenges embracing different approaches to exascale (capability, capacity, and urgent computing). The first implementation phase of the project (ChEESE-1P; 2018–2022) addressed scientific and technical computational challenges in seismology, tsunami science, volcanology, and magnetohydrodynamics, in order to understand the phenomena, anticipate the impact of natural disasters, and contribute to risk management. The project initiated the optimisation of 10 community flagship codes for the upcoming exascale systems and implemented 12 Pilot Demonstrators that combine the flagship codes with dedicated workflows in order to address the underlying capability and capacity computational challenges. Pilot Demonstrators reaching more mature Technology Readiness Levels (TRLs) were further enabled in operational service environments on critical aspects of geohazards such as long-term and short-term probabilistic hazard assessment, urgent computing, and early warning and probabilistic forecasting. Partnership and service co-design with members of the project Industry and User Board (IUB) leveraged the uptake of results across multiple research institutions, academia, industry, and public governance bodies (e.g. civil protection agencies). This article summarises the implementation strategy and the results from ChEESE-1P, outlining also the underpinning concepts and the roadmap for the on-going second project implementation phase (ChEESE-2P; 2023–2026).
Arnau Folch, Claudia Abril, Michael Afanasiev, Giorgio Amati, Michael Bader, Rosa M. Badia, Hafize B. Bayraktar, Sara Barsotti, Roberto Basili 0002, Fabrizio Bernardi, Christian Boehm, Beatriz Brizuela, Federico Brogi, Eduardo Cabrera, Emanuele Casarotti, Manuel Jesús Castro Díaz, Matteo Cerminara, Antonella Cirella, Alexey Cheptsov, Javier Conejero, Antonio Costa 0002, Marc de la Asunción, Josep de la Puente, Marco Djuric, Ravil Dorozhinskii, Gabriela Espinosa, Tomaso Esposti Ongaro, Joan Farnós, Nathalie Favretto-Cristini, Andreas Fichtner, Alexandre Fournier, Alice-Agnes Gabriel, Jean-Matthieu Gallard, Steven J. Gibbons, Sylfest Glimsdal, José Manuel González-Vida, José Gracia, Rose Gregorio, Natalia Gutiérrez, Benedikt Halldorsson, Okba Hamitou, Guillaume Houzeaux, Stephan Jaure, Mouloud Kessar, Lukas Krenz, Lion Krischer, Soline Laforet, Piero Lanucara, Bo Li 0147, Maria Concetta Lorenzino, Stefano Lorito, Finn Løvholt, Giovanni Macedonio, Jorge Macías Sánchez, Guillermo Marin, Beatriz Martínez Montesinos, Leonardo Mingari, Geneviève Moguilny, Vadim Montellier, Marisol Monterrubio Velasco, Georges-Emmanuel Moulard, Masaru Nagaso, Massimo Nazaria, Christoph Niethammer, Federica Pardini, Marta Pienkowska, Luca Pizzimenti, Natalia Poiata, Leonhard Rannabauer, Otilio Rojas, Juan Esteban Rodriguez, Fabrizio Romano, Oleksandr Rudyy, Vittorio Ruggiero, Philipp Samfass, Carlos Sánchez-Linares, Sabrina Sanchez, Laura Sandri, Antonio Scala, Nathanaël Schaeffer, Joseph Schuchart, Jacopo Selva, Amadine Sergeant, Angela Stallone, Matteo Taroni, Solvi Thrastarson, Manuel Titos, Nadia Tonelllo, Roberto Tonini, Thomas Ulrich, Jean-Pierre Vilotte, Malte Vöge, Manuela Volpe, Sara Aniko Wirp, Uwe Wössner
Future Gener. Comput. Syst.6
2022 The BioExcel methodology for developing dynamic, scalable, reliable and portable computational biomolecular workflows
abstract
Developing complex biomolecular workflows is not always straightforward. It requires tedious developments to enable the interoperability between the different biomolecular simulation and analysis tools. Moreover, the need to execute the pipelines on distributed systems increases the complexity of these developments. To address these issues, we propose a methodology to simplify the implementation of these workflows on HPC infrastructures. It combines a library, the BioExcel Building Blocks (BioBBs), that allows scientists to implement biomolecular pipelines as Python scripts, and the PyCOMPSs programming framework which allows to easily convert Python scripts into task-based parallel workflows executed in distributed computing systems such as HPC clusters, clouds, containerized platforms, etc. Using this methodology, we have implemented a set of computational molecular workflows and we have performed several experiments to validate its portability, scalability, reliability and malleability.
Jorge Ejarque, Pau Andrio, Adam Hospital, Javier Conejero, Daniele Lezzi, Josep Lluís Gelpí, Rosa M. Badia
e-Science7
2022 Enabling dynamic and intelligent workflows for HPC, data analytics, and AI convergence
Jorge Ejarque, Rosa M. Badia, Loïc Albertin, Giovanni Aloisio, Enrico Baglione, Yolanda Becerra 0001, Stefan Boschert, Julian R. Berlin, Alessandro D'Anca, Donatello Elia, François Exertier, Sandro Fiore, José Flich, Arnau Folch, Steven J. Gibbons, Nikolay Koldunov, Francesc Lordan, Stefano Lorito, Finn Løvholt, Jorge Macías Sánchez, Fabrizio Marozzo, Alberto Michelini, Marisol Monterrubio Velasco, Marta Pienkowska, Josep de la Puente, Anna Queralt, Enrique S. Quintana-Ortí, Juan Esteban Rodriguez, Fabrizio Romano, Jedrzej Rybicki, Miroslaw Kupczyk, Jacopo Selva, Domenico Talia, Roberto Tonini, Paolo Trunfio, Manuela Volpe
Future Gener. Comput. Syst.2
2022 Storage-Heterogeneity Aware Task-based Programming Models to Optimize I/O Intensive Applications
abstract
Task-based programming models have enabled the optimized execution of the computation workloads of applications. These programming models can take advantage of large-scale distributed infrastructures by allowing the parallel and distributed execution of applications in high-level work components calledtasks. Nevertheless, in the era of Big Data and Exascale, the amount of data produced by modern scientific applications has already surpassed terabytes and is rapidly increasing. Hence, I/O performance became the bottleneck to overcome in order to achieve more total performance improvement. New storage technologies offer higher bandwidth and faster solutions than traditional Parallel File Systems (PFS). Such storage devices are deployed in modern day infrastructures to boost I/O performance by offering a fast layer that absorbs the generated data. Therefore, it is necessary for any programming model targeting more performance to manage this heterogeneity and take advantage of it to improve the I/O performance of applications. Towards this goal, we propose in this article a set of programming model capabilities that we refer to asStorage-Heterogeneity Awareness. Such capabilities include: (i) abstracting the heterogeneity of storage systems, and (ii) optimizing I/O performance by supporting dedicated I/O schedulers and an automatic data flushing technique. The evaluation section of this article presents the performance results of different applications on the MareNostrum CTE-Power heterogeneous storage cluster. Our experiments demonstrate that a storage-heterogeneity aware programming model can achieve up to almost 5x I/O performance speedup and 48% total time improvement compared to the reference PFS-based usage of the execution infrastructure.
Hatem Elshazly, Jorge Ejarque, Rosa M. Badia
IEEE Trans. Parallel Distributed Syst.3
2021 Advancing Design and Runtime Management of AI Applications with AI-SPRINT (Position Paper)
abstract
The adoption of Artificial intelligence (AI) technologies is steadily increasing. However, to become fully pervasive, AI needs resources at the edge of the network. The cloud can provide the processing power needed for big data, but edge computing is close to where data are produced and therefore crucial to their timely, flexible, and secure management. In this paper, we introduce the AI-SPRINT project, which will provide solutions to seamlessly design, partition, and run AI applications in computing continuum environments. AI-SPRINT will offer novel tools for AI applications development, secure execution, easy deployment, as well as runtime management and optimization: AI-SPRINT design tools will allow trading-off application performance (in terms of end-to-end latency or throughput), energy efficiency, and AI models accuracy while providing security and privacy guarantees. The runtime environment will support live data protection, architecture enhancement, agile delivery, runtime optimization, and continuous adaptation.
Hamta Sedghani, Danilo Ardagna, Matteo Matteucci, Giulio Fontana, Giacomo Verticale, Fabrizio Amarilli, Rosa M. Badia, Daniele Lezzi, Ignacio Blanquer, André Martin, Konrad Wawruch
COMPSAC7
2021 Colony: Parallel Functions as a Service on the Cloud-Edge Continuum
Francesc Lordan, Daniele Lezzi, Rosa M. Badia
Euro-Par3
2021 Superscalar Programming Models: A Perspective from Barcelona
abstract
The importance of the programming model in the development of applications has been increasingly more important with the evolution of computing architectures and infrastructures. Aspects such as the number of cores and heterogeneity in the computing nodes, the increase in scale, and new highly distributed environments (the so-called computing continuum) make it even more critical.
Rosa M. Badia
HPDC1
2021 An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks
abstract
Deep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using high dimension inputs. With the steady increase in datasets and model sizes, model/hybrid parallelism is deemed to have an important role in the future of distributed training of DNNs. We analyze the compute, communication, and memory requirements of Convolutional Neural Networks (CNNs) to understand the trade-offs between different parallelism approaches on performance and scalability. We leverage our model-driven analysis to be the basis for an oracle utility which can help in detecting the limitations and bottlenecks of different parallelism approaches at scale. We evaluate the oracle on six parallelization strategies, with four CNN models and multiple datasets (2D and 3D), on up to 1024 GPUs. The results demonstrate that the oracle has an average accuracy of about 86.74% when compared to empirical results, and as high as 97.57% for data parallelism.
Albert Kahira, Truong Thao Nguyen, Leonardo Arturo Bautista-Gomez, Ryousei Takano, Rosa M. Badia, Mohamed Wahib
HPDC5
2021 Towards enabling I/O awareness in task-based programming models
Hatem Elshazly, Jorge Ejarque, Francesc Lordan, Rosa M. Badia
Future Gener. Comput. Syst.4
2021 DDF Library: Enabling functional programming in a task-based model
Lucas M. Ponce, Daniele Lezzi, Rosa M. Badia, Dorgival O. Guedes
J. Parallel Distributed Comput.3
2020 Managing Failures in Task-Based Parallel Workflows in Distributed Computing Environments
Jorge Ejarque, Marta Bertran, Javier Álvarez Cid-Fuentes, Javier Conejero, Rosa M. Badia
Euro-Par5
2020 Performance Meets Programmabilty: Enabling Native Python MPI Tasks In PyCOMPSs
abstract
The increasing complexity of modern and future computing systems makes it challenging to develop applications that aim for maximum performance. Hybrid parallel programming models offer new ways to exploit the capabilities of the underlying infrastructure. However, the performance gain is sometimes accompanied by increased programming complexity. We introduce an extension to PyCOMPSs, a high-level task-based parallel programming model for Python applications, to support tasks that use MPI natively as part of the task model. Without compromising application's programmability, using Native MPI tasks in PyCOMPSs offers up to 3x improvement in total performance for compute intensive applications and up to 1.9x improvement in total performance for I/O intensive applications over sequential implementation of the tasks.
Hatem Elshazly, Francesc Lordan, Jorge Ejarque, Rosa M. Badia
PDP4
2020 Efficient development of high performance data analytics in Python
abstract
Our society is generating an increasing amount of data at an unprecedented scale, variety, and speed. This also applies to numerous research areas, such as genomics, high energy physics, and astronomy, for which large-scale data processing has become crucial. However, there is still a gap between the traditional scientific computing ecosystem and big data analytics tools and frameworks. On the one hand, high performance computing (HPC) programming models lack productivity, and do not provide means for processing large amounts of data in a simple manner. On the other hand, existing big data processing tools have performance issues in HPC environments, and are not general-purpose. In this paper, we propose and evaluate PyCOMPSs, a task-based programming model for Python, as an excellent solution for distributed big data processing in HPC infrastructures. Among other useful features, PyCOMPSs offers a highly productive general-purpose programming model, is infrastructure-agnostic, and provides transparent data management with support for distributed storage systems. We show how two machine learning algorithms (Cascade SVM and K-means) can be developed with PyCOMPSs, and evaluate PyCOMPSs’ productivity based on these algorithms. Additionally, we evaluate PyCOMPSs performance on an HPC cluster using up to 1,536 cores and 320 million input vectors. Our results show that PyCOMPSs achieves similar performance and scalability to MPI in HPC infrastructures, while providing a much more productive interface that allows the easy development of data analytics algorithms.
Javier Álvarez Cid-Fuentes, Pol Álvarez, Ramon Amela, Kuninori Ishii, Rafael K. Morizawa, Rosa M. Badia
Future Gener. Comput. Syst.6
2020 A programming model for Hybrid Workflows: Combining task-based workflows and dataflows all-in-one
Cristian Ramon-Cortes, Francesc Lordan, Jorge Ejarque, Rosa M. Badia
Future Gener. Comput. Syst.4
2020 Energy-Aware Self-Adaptation for Application Execution on Heterogeneous Parallel Architectures
abstract
Hardware in High Performance Computing environments in recent years have increasingly become more heterogeneous in order to improve computational performance. An additional aspect of such systems is the management of power and energy consumption. The increase in heterogeneity requires middleware and programming model abstractions to eliminate additional complexities that it brings, while also offering opportunities such as improved power management. In this paper, we explore application level self-adaptation including aspects such as automated configuration and deployment of applications to different heterogeneous infrastructure and for their redeployment. This therefore not only mitigates complexities associated with heterogeneous devices but aims to take advantage of the heterogeneity. The overall result of this paper is a self-adaptive framework that manages application Quality of Service (QoS) at runtime, which includes the automatic migration of applications between different accelerated infrastructures. Discussion covers when this migration is appropriate and quantifies the likely benefits.
Richard E. Kavanagh, Karim Djemame, Jorge Ejarque, Rosa M. Badia, David García-Pérez
IEEE Trans. Sustain. Comput.4
2019 dislib: Large Scale High Performance Machine Learning in Python
abstract
In recent years, machine learning has proven to be an extremely useful tool for extracting knowledge from data. This can be leveraged in numerous research areas, such as genomics, earth sciences, and astrophysics, to gain valuable insight. At the same time, Python has become one of the most popular programming languages among researchers due to its high productivity and rich ecosystem. Unfortunately, existing machine learning libraries for Python do not scale to large data sets, are hard to use by non-experts, and are difficult to set up in high performance computing clusters. These limitations have prevented scientists to exploit the full potential of machine learning in their research. In this paper, we present and evaluate dislib, a distributed machine learning library on top of PyCOMPSs programming model that addresses the issues of other existing libraries. In our evaluation, we show that dislib can be up to 9 times faster, and can process data sets up to 16 times larger than other popular distributed machine learning libraries, such as MLlib. In addition to this, we also show how dislib can be used to reduce the computation time of a real scientific application from 18 hours to 17 minutes.
Javier Álvarez Cid-Fuentes, Salvi Solà, Pol Álvarez, Alfred Castro-Ginard, Rosa M. Badia
eScience5
2019 Workflow Environments for Advanced Cyberinfrastructure Platforms
abstract
Progress in science is deeply bound to the effective use of high-performance computing infrastructures and to the efficient extraction of knowledge from vast amounts of data. Such data comes from different sources that follow a cycle composed of pre-processing steps for data curation and preparation for subsequent computing steps, and later analysis and analytics steps applied to the results. However, scientific workflows are currently fragmented in multiple components, with different processes for computing and data management, and with gaps in the viewpoints of the user profiles involved. Our vision is that future workflow environments and tools for the development of scientific workflows should follow a holistic approach, where both data and computing are integrated in a single flow built on simple, high-level interfaces. The topics of research that we propose involve novel ways to express the workflows that integrate the different data and compute processes, dynamic runtimes to support the execution of the workflows in complex and heterogeneous computing infrastructures in an efficient way, both in terms of performance and energy. These infrastructures include highly distributed resources, from sensors and instruments, and devices in the edge, to High-Performance Computing and Cloud computing resources. This paper presents our vision to develop these workflow environments and also the steps we are currently following to achieve it.
Rosa M. Badia, Jorge Ejarque, Francesc Lordan, Daniele Lezzi, Javier Conejero, Javier Álvarez Cid-Fuentes, Yolanda Becerra 0001, Anna Queralt
ICDCS1
2019 Extension of a Task-Based Model to Functional Programming
abstract
Recently, efforts have been made to bring together the areas of high-performance computing (HPC) and massive data processing (Big Data). Traditional HPC frameworks, like COMPSs, are mostly task-based, while popular big-data environments, like Spark, are based on functional programming principles. The earlier are know for their good performance for regular, matrix-based computations; on the other hand, for fine-grained, data-parallel workloads, the later has often been considered more successful. In this paper we present our experience with the integration of some dataflow techniques into COMPSs, a task-based framework, in an effort to bring together the best aspects of both worlds. We present our API, called DDF, which provides a new data abstraction that addresses the challenges of integrating Big Data application scenarios into COMPSs. DDF has a functional-based interface, similar to many Data Science tools, that allows us to use dynamic evaluation to adapt the task execution in runtime. Besides the performance optimization it provides, the API facilitates the development of applications by experts in the application domain. In this paper we evaluate DDF's effectiveness by comparing the resulting programs to their original versions in COMPSs and Spark. The results show that DDF can improve COMPSs execution time and even outperform Spark in many use cases.
Lucas M. Ponce, Daniele Lezzi, Rosa M. Badia, Dorgival O. Guedes
SBAC-PAD3
2019 BIGSEA: A Big Data analytics platform for public transportation information
Andy S. Alic, Jussara M. Almeida, Giovanni Aloisio, Nazareno Andrade, Nuno Antunes, Danilo Ardagna, Rosa M. Badia, Tânia Basso, Ignacio Blanquer, Tarciso Braz, Andrey Brito, Donatello Elia, Sandro Fiore, Dorgival O. Guedes, Marco Lattuada 0001, Daniele Lezzi, Matheus Maciel, Wagner Meira Jr., Demetrio Gomes Mestre, Regina Lúcia de Oliveira Moraes, Fábio Morais 0001, Carlos Eduardo S. Pires, Nádia P. Kozievitch, Walter Santos, Paulo Silva 0002, Marco Vieira
Future Gener. Comput. Syst.7
2019 On the maturity of parallel applications for asymmetric multi-core processors
Kallia Chronaki, Miquel Moretó, Marc Casas, Alejandro Rico, Rosa M. Badia, Eduard Ayguadé, Mateo Valero
J. Parallel Distributed Comput.5
2018 Programmability versus performance tradeoff: overcoming the hardware challenges from a task-based approach
abstract
Programming languages that offer simple, elegant interfaces with strong semantics are valued by the applications developers. Python is one example of such a programming language, adopted both by the High Performance Computing and Data Analytics communitites, with a design philosophy that emphasizes code readibilty and a syntax that allows programmers to express concepts in fewer lines of code, while still offering object-orientation and advanced programming features such as generators and list comprehensions. However, Python is an interpreted language and concurreny is ill-supported. This talk will be based on PyCOMPSs, a task-based programming model that aims to parallelize Python sequential codes and to execute them in distributed computing platforms. The talk will overview the system, and present how different hardware challenges are overcome: multicore architectures, accelerators such as GPUs with specific APIs, memory hierarchy. distributed computing, or distributed file systems.
Rosa M. Badia
CF1
2018 Boosting Atmospheric Dust Forecast with PyCOMPSs
abstract
Task-based programming is becoming a tool of large interest for boosting High-Performance Computing (HPC) and Big Data applications. In particular, COMP Superscalar (COMPSs), is showing to be an effective task-based programming model for distributed computing of Big Data applications within HPC environments. Applications like NMMB-MONARCH, which is a dust forecast application composed by a set of steps (being some of them binaries with or without MPI), are perfect candidates for PyCOMPSs, the Python binding of COMPSs. This paper describes the success story of the adaptation of the NMMB-MONARCH online multi-scale atmospheric dust model to PyCOMPSs in order to exploit its inherent parallelism with the minimal developer effort. The paper also includes an evaluation of this implementation in the Nord3 supercomputer, a scalability analysis and an in-depth behaviour study. The main results presented in this paper are: (1) PyCOMPSs is able to extract the parallelism from the NMMB-MONARCH application; (2) it is able to improve the dust forecasting in terms of performance when compared with previous versions, and (3) PyCOMPSs is able to interact and share the resources with MPI applications when included in the workflow as tasks. Finally, we present the keys for exporting the knowledge of this experience to other applications in order to benefit from using PyCOMPSs.
Javier Conejero, Cristian Ramon-Cortes, Kim Serradell, Rosa M. Badia
eScience4
2018 Dynamic energy-aware scheduling for parallel task-based application in cloud computing
Fredy Juarez, Jorge Ejarque, Rosa M. Badia
Future Gener. Comput. Syst.3
2018 Towards Mobile Cloud Computing with Single Sign-on Access
Francesc Lordan, Jens Jensen, Rosa M. Badia
J. Grid Comput.3
2018 Transparent Orchestration of Task-based Parallel Applications in Containers Platforms
Cristian Ramon-Cortes, Albert Serven, Jorge Ejarque, Daniele Lezzi, Rosa M. Badia
J. Grid Comput.5
2017 Transparent Execution of Task-Based Parallel Applications in Docker with COMP Superscalar
abstract
This paper presents a framework to easily build and execute parallel applications in container-based distributed computing platforms in a user transparent way. The proposed framework is a combination of the COMP Superscalar and Docker. We have built a prototype in order to evaluate how it performs by evaluating the overhead in the building, deployment and execution phases. We have observed an important gain compared with cloud environments during the building and deployment phases. In contrast, we have detected an extra overhead during the execution, which is mainly due to the multi-host Docker networking.
Victor Anton, Cristian Ramon-Cortes, Jorge Ejarque, Rosa M. Badia
PDP4
2017 COMPSs-Mobile: Parallel Programming for Mobile Cloud Computing
Francesc Lordan, Rosa M. Badia
J. Grid Comput.2
2017 Task Scheduling Techniques for Asymmetric Multi-Core Systems
abstract
As performance and energy efficiency have become the main challenges for next-generation high-performance computing, asymmetric multi-core architectures can provide solutions to tackle these issues. Parallel programming models need to be able to suit the needs of such systems and keep on increasing the application’s portability and efficiency. This paper proposes two task scheduling approaches that target asymmetric systems. These dynamic scheduling policies reduce total execution time either by detecting the longest or the critical path of the dynamic task dependency graph of the application, or by finding the earliest executor of a task. They use dynamic scheduling and information discoverable during execution, fact that makes them implementable and functional without the need of off-line profiling. In our evaluation we compare these scheduling approaches with two existing state-of the art heterogeneous schedulers and we track their improvement over a FIFO baseline scheduler. We show that the heterogeneous schedulers improve the baseline by up to 1.45$\times$in a real 8-core asymmetric system and up to 2.1$\times$in a simulated 32-core asymmetric chip.
Kallia Chronaki, Alejandro Rico, Marc Casas, Miquel Moretó, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero
IEEE Trans. Parallel Distributed Syst.5
2016 POSTER: Exploiting Asymmetric Multi-Core Processors with Flexible System Sofware
abstract
Energy efficiency has become the main challenge for high performance computing (HPC). The use of mobile asymmetric multi-core architectures to build future multi-core systems is an approach towards energy savings while keeping high performance. However, it is not known yet whether such systems are ready to handle parallel applications.
Kallia Chronaki, Miquel Moretó, Marc Casas, Alejandro Rico, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero
PACT5
2016 COMPSs-Mobile: Parallel Programming for Mobile-Cloud Computing
abstract
The advent of Cloud and the popularization of mobile devices have led us to a shift in computing access. Computing users will have an interaction display while the real computation will be performed remotely, in the Cloud. COMPSs-Mobile is a framework that aims to ease the development of energy-efficient and high-performing applications for this environment. The framework provides an infrastructure-unaware programming model that allows developers to code regular Android applications that, transparently, are parallelized, and partially offloaded to remote resources. This paper gives an overview of the programming model and describes the internal components of the toolkit which supports it focusing on the offloading and checkpointing mechanisms. It also presents the results of some tests conducted to evaluate the behavior of the solution and to measure the potential benefits in Android applications.
Francesc Lordan, Rosa M. Badia
CCGrid2
2016 CATA: Criticality Aware Task Acceleration for Multicore Processors
abstract
Managing criticality in task-based programming models opens a wide range of performance and power optimization opportunities in future manycore systems. Criticality aware task schedulers can benefit from these opportunities by scheduling tasks to the most appropriate cores. However, these schedulers may suffer from priority inversion and static binding problems that limit their expected improvements. Based on the observation that task criticality information can be exploited to drive hardware reconfigurations, we propose a Criticality Aware Task Acceleration (CATA) mechanism that dynamically adapts the computational power of a task depending on its criticality. As a result, CATA achieves significant improvements over a baseline static scheduler, reaching average improvements up to 18.4% in execution time and 30.1% in Energy-Delay Product (EDP) on a simulated 32-core system. The cost of reconfiguring hardware by means of a software-only solution rises with the number of cores due to lock contention and reconfiguration overhead. Therefore, novel architectural support is proposed to eliminate these overheads on future manycore systems. This architectural support minimally extends hardware structures already present in current processors, which allows further improvements in performance with negligible overhead. As a consequence, average improvements of up to 20.4% in execution time and 34.0% in EDP are obtained, outperforming state-of-the-art acceleration proposals not aware of task criticality.
Emilio Castillo, Miquel Moretó, Marc Casas, Lluc Alvarez, Enrique Vallejo 0001, Kallia Chronaki, Rosa M. Badia, José Luis Bosque, Ramón Beivide, Eduard Ayguadé, Jesús Labarta, Mateo Valero
IPDPS7
2016 Energy-Aware Programming Model for Distributed Infrastructures
abstract
Day after day, cloud technologies are more and more adopted by very diverse types of stakeholders, and this success creates a side-effect problem: the energy spent by this kind of infrastructures is growing bigger every day. With the objective of reducing energy consumption when programming applications for cloud infrastructures, we have implemented energy-aware mechanisms in the COMPSs Programming Model, inside the context of the ASCETiC Project. In this paper, we demonstrate that application-level scheduling can have a big impact on the energy consumed by an application when executed in a heterogeneous cloud. We have implemented an energy-aware scheduling mechanism in COMPSs, together with a versioning technique, and we have run experiments with a use case coming from the real estate sector that proves our hypotheses.
Francesc Lordan, Jorge Ejarque, Raül Sirvent, Rosa M. Badia
PDP4
2016 Web Services as Building Blocks for Science Gateways in Astrophysics
Susana Sánchez-Expósito, Pablo Martín, José Enrique Ruiz, Lourdes Verdes-Montenegro, Julián Garrido, Raül Sirvent, Antonio Ruiz Falcó, Rosa M. Badia, Daniele Lezzi
J. Grid Comput.8
2016 Exploiting task and data parallelism in ILUPACK's preconditioned CG solver on NUMA architectures and many-core accelerators
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
Parallel Comput.2
2015 Towards Automatic Application Migration to Clouds
abstract
Porting applications to Clouds is one of the key challenges in software industry. The available approaches to perform this task are basically either services derived from alliances of major software vendors and Cloud providers focusing on their own products, or small platform providers focusing on the most popular software stacks. For migrating other types of software, the options are limited to Infrastructure-as-a-Service (IaaS) solutions which require a lot of programming effort for adapting the software to a Cloud provider's API. Moreover, if it must be deployed in different providers, new integration procedures must be designed and implemented which could be a nightmare. This paper presents a solution for facilitating the migration of any application to the cloud, inferring the most suitable deployment model for the application and automatically deploying it in the available Cloud providers.
Jorge Ejarque, András Micsik, Rosa M. Badia
CLOUD3
2015 Criticality-Aware Dynamic Task Scheduling for Heterogeneous Architectures
abstract
Current and future parallel programming models need to be portable and efficient when moving to heterogeneous multi-core systems. OmpSs is a task-based programming model with dependency tracking and dynamic scheduling. This paper describes the OmpSs approach on scheduling dependent tasks onto the asymmetric cores of a heterogeneous system. The proposed scheduling policy improves performance by prioritizing the newly-created tasks at runtime, detecting the longest path of the dynamic task dependency graph, and assigning critical tasks to fast cores. While previous works use profiling information and are static, this dynamic scheduling approach uses information that is discoverable at runtime which makes it implementable and functional without the need of an oracle or profiling. The evaluation results show that our proposal outperforms a dynamic implementation of Heterogeneous Earliest Finish Time by up to 1.15x, and the default breadth-first OmpSs scheduler by up to 1.3x in an 8-core heterogeneous platform and up to 2.7x in a simulated 128-core chip.
Kallia Chronaki, Alejandro Rico, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero
ICS3
2015 Supporting biodiversity studies with the EUBrazilOpenBio Hybrid Data Infrastructure
abstract
Summary EUBrazilOpenBio is a collaborative initiative addressing strategic barriers in biodiversity research by integrating open access data and user‐friendly tools widely available in Brazil and Europe. The project deploys the EU‐Brazil Hybrid Data Infrastructure that allows the sharing of hardware, software and data on‐demand. This infrastructure provides access to several integrated services and resources to seamlessly aggregate taxonomic, biodiversity and climate data, used by processing services implementing checklist cross‐mapping and ecological niche modelling. A Virtual Research Environment was created to provide users with a single entry point to processing and data resources. This article describes the architecture, demonstration use cases and some experimental results and validation. Copyright © 2014 John Wiley & Sons, Ltd.
Rafael Amaral, Rosa M. Badia, Ignacio Blanquer, Ricardo Braga-Neto, Leonardo Candela, Donatella Castelli, Christina Flann, Renato De Giovanni, W. Alex Gray, Andrew C. Jones, Daniele Lezzi, Pasquale Pagano, Vanderlei Perez Canhos, Francisco Quevedo, Roger Rafanell, Vinod E. F. Rebello, Mariane S. Sousa-Baena, Erik Torres
Concurr. Comput. Pract. Exp.2
2015 Special Section on Terascale Computing
Stefan Wesner, Lutz Schubert, Rosa M. Badia, Antonio Rubio 0001, Pier Stanislao Paolucci, Roberto Giorgi
Future Gener. Comput. Syst.3
2015 Picos: A hardware runtime architecture support for OmpSs
Fahimeh Yazdanpanah, Carlos Álvarez 0001, Daniel Jiménez-González, Rosa M. Badia, Mateo Valero
Future Gener. Comput. Syst.4
2015 Resource-Aware Task Scheduling
abstract
Dependency-aware task-based parallel programming models have proven to be successful for developing efficient application software for multicore-based computer architectures. The programming model is amenable to programmers, thereby supporting productivity, whereas hardware performance is achieved through a runtime system that dynamically schedules tasks onto cores in such a way that all dependencies are respected. However, even if the scheduling is completely successful with respect to load balancing, the scaling with the number of cores may be suboptimal due to resource contention. Here we consider the problem of scheduling tasks not only with respect to their interdependencies but also with respect to their usage of resources, such as memory and bandwidth. At the software level, this is achieved by user annotations of the task resource consumption. In the runtime system, the annotations are translated into scheduling constraints. Experimental results for different hardware, demonstrating performance gains both for model examples and real applications, are presented. Furthermore, we provide a set of tools to detect resource sensitivity and predict the performance improvements that can be achieved by resource-aware scheduling. These tools are solely based on parallel execution traces and require no instrumentation or modification of the application code.
Martin Tillenius, Elisabeth Larsson, Rosa M. Badia, Xavier Martorell
ACM Trans. Embed. Comput. Syst.3
2014 Leveraging Task-Parallelism with OmpSs in ILUPACK's Preconditioned CG Method
abstract
In this paper we describe how to efficiently exploit task parallelism for the solution of sparse linear systems on multithreaded processors via ILUPACK's multi-level preconditioned CG method. Using a pair of data structures, we capture the task dependencies that appear in the two most challenging operations in the method (calculation of the preconditioned and its application), passing this information to the OmpSs runtime which can then implement a correct and efficient schedule of the entire solver. Our results with high-end multicore platforms equipped with Intel and AMD processors report significant performance gains, demonstrating that OmpSs provides an efficient and close-to seamless means to leverage the concurrency in a complex scientific code like ILUPACK.
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Enrique S. Quintana-Ortí
SBAC-PAD2
2014 ServiceSs: An Interoperable Programming Framework for the Cloud
Francesc Lordan, Enric Tejedor, Jorge Ejarque, Roger Rafanell, Javier Álvarez Cid-Fuentes, Fabrizio Marozzo, Daniele Lezzi, Raül Sirvent, Domenico Talia, Rosa M. Badia
J. Grid Comput.10
2014 Leveraging task-parallelism in message-passing dense matrix factorizations using SMPSs
Alberto F. Martín, Ruymán Reyes, Rosa M. Badia, Enrique S. Quintana-Ortí
Parallel Comput.3
2013 The TERAFLUX Project: Exploiting the DataFlow Paradigm in Next Generation Teradevices
abstract
Thanks to the improvements in semiconductor technologies, extreme-scale systems such as teradevices (i.e., composed by 1000 billion of transistors) will enable systems with 1000+ general purpose cores per chip, probably by 2020. Three major challenges have been identified: programmability, manageable architecture design, and reliability. TERAFLUX is a Future and Emerging Technology (FET) large-scale project funded by the European Union, which addresses such challenges at once by leveraging the dataflow principles. This paper describes the project and provides an overview of the research carried out by the TERAFLUX consortium.
Marco Solinas, Rosa M. Badia, François Bodin, Albert Cohen 0001, Paraskevas Evripidou, Paolo Faraboschi, Bernhard Fechner, Guang R. Gao, Arne Garbade, Sylvain Girbal, Daniel Goodman 0001, Behram Khan, Souad Koliai, Feng Li 0016, Mikel Luján, Laurent Morin, Avi Mendelson, Nacho Navarro, Antoniu Pop, Pedro Trancoso, Theo Ungerer, Mateo Valero, Sebastian Weis, Ian Watson, Stéphane Zuckerman, Roberto Giorgi
DSD2
2013 Topic 6: Grid, Cluster and Cloud Computing - (Introduction)
Erwin Laure, Odej Kao, Rosa M. Badia, Laurent Lefèvre, Beniamino Di Martino, Radu Prodan, Matteo Turilli, Daniel Warneke
Euro-Par3
2013 Loop level speculation in a task based programming model
abstract
Uncountable loops (such as while loops in C) and if-conditions are some of the most common constructs in programming. While-loops are widely used to determine the convergence in linear algebra algorithms or goal finding problems from graph algorithms, to name a few. In general while-loops are used whenever the loop iteration space, the number of iterations a loop executes is unknown. Usually in while-loops, the execution of the next iteration is decided inside the current loop iteration (i.e. the execution of iteration i depends on the values computed in iteration i-1). This precludes their parallel execution in today's ubiquitous multi-core architectures. In this paper a technique to speculatively create parallel tasks from the next iterations before the current one completes is proposed. If consecutive loop-iterations are only control dependent, then multiple iterations can be executed simultaneously; later in the execution path, the runtime system will decide to either commit the results of such speculatively executed iterations or undo the changes made by them. Data dependences within or between non-speculative and speculative work are honored to guarantee correctness. The proposed technique is implemented in SMPSs, a task-based dataflow programming model for shared-memory multiprocessor architectures. The approach is evaluated on a set of applications from graph algorithms and linear algebra. Results are promising with an average increase in the speedup of 1.2x with 16 threads when compared to non speculative execution of the applications. The increase in the speedup is significant, since the performance gain is achieved over an already parallelized version of the benchmarks.
Rahulkumar Gayatri, Rosa M. Badia, Eduard Ayguadé
HiPC2
2013 Implementing OmpSs support for regions of data in architectures with multiple address spaces
abstract
The need for features for managing complex data accesses in modern programming models has increased due to the emerging hardware architectures. HPC hardware has moved towards clusters of accelerators and/or multicores, architectures with a complex memory hierarchy exposed to the programmer.
Javier Bueno, Xavier Martorell, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta
ICS3
2013 Programmable and Scalable Reductions on Clusters
abstract
Reductions matter and they are here to stay. Wide adoption of parallel processing hardware in a broad range of computer applications has encouraged recent research efforts on their efficient parallelization. Furthermore, trends towards high productivity languages in mainstream computing increases the demand for efficient programming support. In this paper we present a new approach on parallel reductions for distributed memory systems that provides both scalability and programmability. Using OmpSs, a task-based parallel programming model, the developer has the ability to express scalable reductions through a single pragma annotation. This pragma annotation is applicable for tasks as well as for work-sharing constructs (with implicit tasking) and instructs the compiler to generate the required runtime calls. The supporting runtime handles data and task distribution, parallel execution and data reduction. Scalability is achieved through a software cache that maximizes local and temporal data reuse and allows overlapped computation and communication. Results confirm scalability for up to 32 12-core cluster nodes.
Jan Ciesko, Javier Bueno, Nikola Puzovic, Alex Ramírez, Rosa M. Badia, Jesús Labarta
IPDPS5
2013 Self-Adaptive OmpSs Tasks in Heterogeneous Environments
abstract
As new heterogeneous systems and hardware accelerators appear, high performance computers can reach a higher level of computational power. Nevertheless, this does not come for free: the more heterogeneity the system presents, the more complex becomes the programming task in terms of resource management. OmpSs is a task-based programming model and framework focused on the runtime exploitation of parallelism from annotated sequential applications. This paper presents a set of extensions to this framework: we show how the application programmer can expose different specialized versions of tasks (i.e. pieces of specific code targeted and optimized for a particular architecture) and how the system can choose between these versions at runtime to obtain the best performance achievable for the given application. From the results obtained in a multi-GPU system, we prove that our proposal gives flexibility to application's source code and can potentially increase application's performance.
Judit Planas, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta
IPDPS2
2012 Transactional Access to Shared Memory in StarSs, a Task Based Programming Model
Rahulkumar Gayatri, Rosa M. Badia, Eduard Ayguadé, Mikel Luján, Ian Watson
Euro-Par2
2012 Enabling Cloud Interoperability with COMPSs
Fabrizio Marozzo, Francesc Lordan, Roger Rafanell, Daniele Lezzi, Domenico Talia, Rosa M. Badia
Euro-Par6
2012 Tools for Power-Energy Modelling and Analysis of Parallel Scientific Applications
abstract
Understanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core.
Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes
ICPP2
2012 Productive Programming of GPU Clusters with OmpSs
abstract
Clusters of GPUs are emerging as a new computational scenario. Programming them requires the use of hybrid models that increase the complexity of the applications, reducing the productivity of programmers. We present the implementation of OmpSs for clusters of GPUs, which supports asynchrony and heterogeneity for task parallelism. It is based on annotating a serial application with directives that are translated by the compiler. With it, the same program that runs sequentially in a node with a single GPU can run in parallel in multiple GPUs either local (single node) or remote (cluster of GPUs). Besides performing a task-based parallelization, the runtime system moves the data as needed between the different nodes and GPUs minimizing the impact of communication by using affinity scheduling, caching, and by overlapping communication with the computational task. We show several applications programmed with OmpSs and their performance with multiple GPUs in a local node and in remote nodes. The results show good tradeoff between performance and effort from the programmer.
Javier Bueno, Judit Planas, Alejandro Duran, Rosa M. Badia, Xavier Martorell, Eduard Ayguadé, Jesús Labarta
IPDPS4
2012 Cloud Application Resource Mapping and Scaling Based on Monitoring of QoS Constraints
Xabriel J. Collazo-Mojica, Seyed Masoud Sadjadi, Jorge Ejarque, Rosa M. Badia
SEKE4
2012 A high-productivity task-based programming model for clusters
abstract
SUMMARY Programming for large‐scale, multicore‐based architectures requires adequate tools that offer ease of programming and do not hinder application performance. StarSs is a family of parallel programming models based on automatic function‐level parallelism that targets productivity. StarSs deploys a data‐flow model: it analyzes dependencies between tasks and manages their execution, exploiting their concurrency as much as possible. This paper introduces Cluster Superscalar (ClusterSs), a new StarSs member designed to execute on clusters of SMPs (Symmetric Multiprocessors). ClusterSs tasks are asynchronously created and assigned to the available resources with the support of the IBM APGAS runtime, which provides an efficient and portable communication layer based on one‐sided communication. We present the design of ClusterSs on top of APGAS, as well as the programming model and execution runtime for Java applications. Finally, we evaluate the productivity of ClusterSs, both in terms of programmability and performance and compare it to that of the IBM X10 language. Copyright © 2012 John Wiley & Sons, Ltd.
Enric Tejedor, Montse Farreras, David Grove, Rosa M. Badia, Gheorghe Almási 0001, Jesús Labarta
Concurr. Comput. Pract. Exp.4
2012 OPTIMIS: A holistic approach to cloud service provisioning
Ana Juan Ferrer, Francisco Hernández-Rodriguez, Johan Tordsson, Erik Elmroth, Ahmed Ali-Eldin, Csilla Zsigri, Raül Sirvent, Jordi Guitart, Rosa M. Badia, Karim Djemame, Wolfgang Ziegler, Theodosis Dimitrakos, Srijith Krishnan Nair, George Kousiouris, Kleopatra Konstanteli, Theodora A. Varvarigou, Benoit Hudzia, Alexander Kipp, Stefan Wesner, Marcelo Corrales, Nikolaus Forgó, Tabassum Sharif, Craig Sheridan
Future Gener. Comput. Syst.9
2011 Parallel Implementation of the Integral Histogram
Pieter Bellens, Kannappan Palaniappan, Rosa M. Badia, Guna Seetharaman, Jesús Labarta
ACIVS3
2011 A Rule-based Approach for Infrastructure Providers' Interoperability
abstract
Cloud Computing is a new computing paradigm where a large amount of computing capacity is offered on demand and only paying for what you use. Several Infrastructure Providers have adopted this approach offering resources which are easily managed by means of web-based APIs. However, if a user wants to use different providers, the resource management becomes tedious because providers define different API requiring a special implementation for interacting with each of them. In this paper, we present a methodology for making the provider interoperability easier. In this methodology, each provider's API is modeled by an ontology. Equivalences between these ontologies are modeled by rules, and messages used in a provider's API are converted in calls to another provider's API applying these rules. With our approach, users interact with Infrastructure Providers using their most familiar API and the translation to the other APIs is automatically done by the system.
Jorge Ejarque, Javier Álvarez Cid-Fuentes, Raül Sirvent, Rosa M. Badia
CloudCom4
2011 A Cloud-unaware Programming Model for Easy Development of Composite Services
abstract
Cloud computing is inherently service-oriented: cloud applications are delivered to consumers as services via the Internet. Therefore, these applications can potentially benefit from the Service-Oriented Architecture (SOA) principles: they can be programmed as added-value services composed by pre-existing ones, thus favouring code reuse. However, new programming models are required to simplify their development, along with systems that are capable of orchestrating the execution of the resulting SaaS in the Cloud. In that regard, this paper presents Service Super scalar (Servicess), an alternative to existing PaaS which provides a programming model and execution runtime to ease the development and execution of service-based applications in clouds. Servicess is a task-based model: the user is only required to select the tasks, which can be services or regular methods, to be spawned asynchronously. The application, a composite service, is programmed in a totally sequential way and no API call must be included in the code. The runtime is in charge of automatically orchestrating the execution of the tasks in the Cloud, as well as of elastically deploying new virtual resources depending on the load. After describing the main characteristics of the programming model and the runtime, we evaluate the productivity of Servicess and show how it offers a good trade-off between programmability and runtime performance.
Enric Tejedor, Jorge Ejarque, Francesc Lordan, Roger Rafanell, Javier Álvarez Cid-Fuentes, Daniele Lezzi, Raül Sirvent, Rosa M. Badia
CloudCom8
2011 Introduction
Rosa M. Badia, Fabrice Huet, Rob van Nieuwpoort, Rainer Keller
Euro-Par (1)1
2011 Productive Cluster Programming with OmpSs
Javier Bueno, Luis Martinell, Alejandro Duran, Montse Farreras, Xavier Martorell, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta
Euro-Par (1)6
2011 ClusterSs: a task-based programming model for clusters
abstract
Programming for large-scale, multicore-based architectures requires adequate tools that offer ease of programming while not hindering application performance. StarSs is a family of parallel programming models based on automatic function level parallelism that targets productivity. StarSs deploys a data-flow model: it analyses dependencies between tasks and manages their execution, exploiting their concurrency as much as possible. We introduce Cluster Superscalar (ClusterSs), a new StarSs member designed to execute on clusters of SMPs. ClusterSs tasks are asynchronously created and assigned to the available resources with the support of the IBM APGAS runtime, which provides an efficient and portable communication layer based on one-sided communication.This short paper gives an overview of the ClusterSs design on top of APGAS, as well as the conclusions of a productivity study; in this study, ClusterSs was compared to the IBM X10 language, both in terms of programmability and performance. A technical report is available with the details.
Enric Tejedor, Montse Farreras, David Grove, Rosa M. Badia, Gheorghe Almási 0001, Jesús Labarta
HPDC4
2011 Poster: programming clusters of GPUs with OMPSs
abstract
OmpSs is a programming model that provides an environment to develop parallel applications for cluster environments with heterogeneous architectures. Based on OpenMP and StarSs, it offers a set of compiler directives that can be used to annotate a sequential code. Additional features have been added to support the use of accelerators like GPUs. This schema offers a high productivity environment due to its simplicity compared to other models like MPI. Our current implementation has shown a good performance when running different benchmarks.
Javier Bueno, Alejandro Duran, Xavier Martorell, Eduard Ayguadé, Rosa M. Badia, Jesús Labarta
ICS5
2011 A Study of Speculative Distributed Scheduling on the Cell/B.E
abstract
Star Superscalar's (StarSs) programming model converts a sequential application in C or Fortran into an efficient parallel program. The resulting parallel code is highly dynamic in the sense that data analysis and task scheduling occur at run-time, while the application executes. In this paper we compare this approach to the strategy adopted by other multi-core programming environments. The prize to pay for dynamic scheduling and dependence tracking is higher runtime overhead. We propose a distributed scheduler for Task Dependence Graphs (TDGs) to attenuate the scheduling cost in heterogeneous multi-core architectures. This scheduler allows the cores to speculatively select tasks from a conservative estimate of the TDG. In case of conflicts or lack of tasks a lightweight centralized scheduler services the faulting core after which the latter resumes its participation in the distributed scheme. Experiments with Cell Super scalar (CellSs) on a representative set of benchmarks demonstrate the reduction in runtime overhead achieved by the distributed scheduler. This reduction in runtime overhead carries over directly to a performance improvement for a large fraction of the benchmarks.
Pieter Bellens, Josep M. Pérez, Rosa M. Badia, Jesús Labarta
IPDPS3
2011 Job Scheduling with License Reservation: A Semantic Approach
abstract
The license management is one of the main concerns when Independent Software Vendors (ISV) try to distribute their software in computing platforms such as Clouds. They want to be sure that customers use their software according to their license terms. The work presented in this paper tries to solve part of this problem extending a semantic resource allocation approach for supporting the scheduling of job taking into account software licenses. This approach defines the licenses as another type of computational resource which is available in the system and must be allocated to the different jobs requested by the users. License terms are modeled as resource properties, which describe the license constraints. A resource ontology has been extended in order to model the relations between customers, providers, jobs, resources and licenses in detail and make them machine processable. The license scheduling has been introduced in a semantic resource allocation process by providing a set of rules, which evaluate the semantic license terms during the job scheduling.
Jorge Ejarque, András Micsik, Raül Sirvent, Peter Pallinger, Rosa M. Badia
PDP6
2011 Extracting the optimal sampling frequency of applications using spectral analysis
abstract
SUMMARY The research community have agreed on several applications as benchmarks to evaluate the adequateness of architectures and high performance computing infrastructures. The performance of these benchmarks is used to determine the weaknesses and strengths of novel designs. Therefore, the performance evaluation of benchmarks is a key factor in the process of designing new architectures. In this paper, we propose a new method based on spectral analysis that allows to perform an automatic analysis of benchmarks' executions. The output of the new method is a representative segment of the benchmarks' executions. Given the nature of the method, the optimal sampling interval length of applications is obtained. This method complements and improves existing techniques focused on the reduction of the application's instruction execution stream of sequential benchmarks and enables the extraction of significant performance information of parallel benchmarks without executing the whole application. The results obtained with the SPEC CPU2000 and the NAS Parallel Benchmarks demonstrate the efficiency and benefits of the approach. Copyright © 2011 John Wiley & Sons, Ltd.
Marc Casas, Harald Servat, Rosa M. Badia, Jesús Labarta
Concurr. Comput. Pract. Exp.3
2010 A Multi-agent Approach for Semantic Resource Allocation
abstract
This paper presents a new approach of the Semantically Enhanced Resource Allocation (SERA) distributed as a multi-agent system. It presents a distributed resource allocation process which combines the benefits of semantic web for making easier the integration between multiple resource providers in the Cloud and agent technologies for coordinating and adapting the execution accross the different providers. The allocation process is based on the negotiation of different agents which allows the combination of customer and providers policies getting scheduling results which satisfies both parts. The SERA agents can be deployed in multiple locations improving the system scalability. The new approach makes the SERA suitable for working as a scheduler inside a Service Provider as well as a metascheduler integrating resources from different providers and platforms (clusters, grids, clouds,...).
Jorge Ejarque, Raül Sirvent, Rosa M. Badia
CloudCom3
2010 Handling task dependencies under strided and aliased references
abstract
The emergence of multicore processors has increased the need for simple parallel programming models usable by nonexperts. The ability to specify subparts of a bigger data structure is an important trait of High Productivity Programming Languages. Such a concept can also be applied to dependency-aware task-parallel programming models. In those paradigms, tasks may have data dependencies, and those are used for scheduling them in parallel.
Josep M. Pérez, Rosa M. Badia, Jesús Labarta
ICS2
2010 Task Superscalar: An Out-of-Order Task Pipeline
abstract
We present \emph{Task Super scalar}, an abstraction of instruction-level out-of-order pipeline that operates at the task-level. Like ILP pipelines, which uncover parallelism in a sequential instruction stream, task super scalar uncovers task-level parallelism among tasks generated by a sequential thread. Utilizing intuitive programmer annotations of task inputs and outputs, the task super scalar pipeline dynamically detects inter-task data dependencies, identifies task-level parallelism, and executes tasks out-of-order. Furthermore, we propose a design for a distributed task super scalar pipeline front end, that can be embedded into any many core fabric, and manages cores as functional units. We show that our proposed mechanism is capable of driving hundreds of cores simultaneously with non-speculative tasks, which allows our pipeline to sustain work windows consisting of tens of thousands of tasks. We further show that our pipeline can maintain a decode rate faster than 60ns per task and dynamically uncover data dependencies among as many as ~50,000 in-flight tasks, using 7MB of on-chip eDRAM storage. This configuration achieves speedups of 95-255x (average 183x) over sequential execution for nine scientific benchmarks, running on a simulated CMP with 256 cores. Task super scalar thus enables programmers to exploit many core systems effectively, while simultaneously simplifying their programming model.
Yoav Etsion, Felipe Cabarcas, Alejandro Rico, Alex Ramírez, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero
MICRO5
2010 Cell BE and Bluetooth applied to Digital TV
abstract
This paper presents the results of a research in progress about a system for multiple users' behavior analysis and program recommendation for digital TV called SmarTV. The research explores a multi-user's behavior analysis and program recommendation using Bluetooth cell phones and a Playstation 3 working as a set-top box connected to a digital TV network with return channel or an IPTV network. Supposing that every viewer has a Bluetooth enabled cell phone, each mobile will identify one spectator to the set-top box wirelessly. Since the set-top box used is a Playstation 3, it is a full computer with IBM Cell BE processor capable of executing computer-intensive algorithms. This device was used to store the history of programs watched by every person near the TV connected to the set-top box, and to identify their profile and preferences using data mining algorithms. The program recommendation system is a hybrid of a content-based approach with a collaborative approach using a rate system, and uses the k-Means and k-NN data mining algorithms modified to explore the Cell BE features.
Aislan Gomide Foina, Francisco Javier Ramirez Fernandez, Rosa M. Badia
NOMS3
2010 Exploiting semantics and virtualization for SLA-driven resource allocation in service providers
abstract
Abstract Resource management is a key challenge that service providers must adequately face in order to accomplish their business goals. This paper introduces a framework, the semantically enhanced resource allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by the service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide an application specific and isolated virtual environment for each task. In addition, the system supports fine‐grain dynamic resource distribution among these virtual environments based on Service‐Level Agreements. The required adaptation is implemented using agents, guarantying enough resources to each task in order to meet the agreed performance goals. Copyright © 2009 John Wiley & Sons, Ltd.
Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres
Concurr. Comput. Pract. Exp.6
2010 Scheduling dense linear algebra operations on multicore processors
abstract
Abstract State‐of‐the‐art dense linear algebra software, such as the LAPACK and ScaLAPACK libraries, suffers performance losses on multicore processors due to their inability to fully exploit thread‐level parallelism. At the same time, the coarse–grain dataflow model gains popularity as a paradigm for programming multicore architectures. This work looks at implementing classic dense linear algebra workloads, the Cholesky factorization, the QR factorization and the LU factorization, using dynamic data‐driven execution. Two emerging approaches to implementing coarse–grain dataflow are examined, the model of nested parallelism, represented by the Cilk framework, and the model of parallelism expressed through an arbitrary Direct Acyclic Graph, represented by the SMP Superscalar framework. Performance and coding effort are analyzed and compared against code manually parallelized at the thread level. Copyright © 2009 John Wiley & Sons, Ltd.
Jakub Kurzak, Hatem Ltaief, Jack J. Dongarra, Rosa M. Badia
Concurr. Comput. Pract. Exp.4
2010 Monitoring and steering Grid applications with GRID superscalar
Sebastián Reyes, Camelia Muñoz-Caro, Alfonso Niño, Raül Sirvent, Rosa M. Badia
Future Gener. Comput. Syst.5
2010 Perspectives on grid computing
Uwe Schwiegelshohn, Rosa M. Badia, Marian Bubak, Marco Danelutto, Schahram Dustdar, Fabrizio Gagliardi, Alfred Geiger, Ladislav Hluchý, Dieter Kranzlmüller, Erwin Laure, Thierry Priol, Alexander Reinefeld, Michael M. Resch, Andreas Reuter 0001, Otto Rienhoff, Thomas Rüter, Peter M. A. Sloot, Domenico Talia, Klaus Ullmann, Ramin Yahyapour
Future Gener. Comput. Syst.2
2009 An Extension of the StarSs Programming Model for Platforms with Multiple GPUs
Eduard Ayguadé, Rosa M. Badia, Francisco D. Igual, Jesús Labarta, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Euro-Par2
2009 Graph-Based Task Replication for Workflow Applications
abstract
The Grid is an heterogeneous and dynamic environment which enables distributed computation. This makes it a technology prone to failures. Some related work uses replication to overcome failures in a set of independent tasks, and in workflow applications, but they do not consider possible resource limitations when scheduling the replicas. In this paper, we focus on the use of task replication techniques for workflow applications, trying to achieve not only tolerance to the possible failures in an execution, but also to speed up the computation without demanding the user to implement an application-level checkpoint, which may be a difficult task depending on the application. Moreover, we also study what to do when there are not enough resources for replicating all running tasks. We establish different priorities of replication depending on the graph of the workflow application, giving more priority to tasks with a higher output degree. We have implemented our proposed policy in the GRID superscalar system, and we have run the fastDNAml as an experiment to prove our objectives are reached. Finally, we have identified and studied a problem which may arise due to the use of replication in workflow applications: the replication wait time.
Raül Sirvent, Rosa M. Badia, Jesús Labarta
HPCC2
2009 Introducing Virtual Execution Environments for Application Lifecycle Management and SLA-Driven Resource Distribution within Service Providers
abstract
Resource management is a key challenge that service providers must adequately face in order to ensure their profitability. This paper describes a proof-of-concept framework for facilitating resource management in service providers, which allows reducing costs and at the same time fulfilling the quality of service agreed with the customers. This is accomplished by means of virtualization. Our approach provides application-specific virtual environments and consolidates them in order to achieve a better utilization of the providers resources. In addition, it implements self-adaptive capabilities for dynamically distributing the providers resources among these virtual environments based on Service Level Agreements. The proposed solution has been implemented as a part of the Semantically-Enhanced Resource Allocator prototype developed within the BREIN European project. The evaluation shows that our prototype is able to react in very short time under changing conditions and avoid SLA violations by rescheduling efficiently the resources.
Íñigo Goiri, Ferran Julià, Jorge Ejarque, Marc de Palol, Rosa M. Badia, Jordi Guitart, Jordi Torres
NCA5
2009 Impact of the Memory Hierarchy on Shared Memory Architectures in Multicore Programming Models
abstract
Many and multicore architectures put a big pressure in parallel programming but gives a unique opportunity to propose new programming models that automatically exploit the parallelism of these architectures. Open MP is a very well known standard that exploits parallelism in shared memory architectures. SMPSs has recently been proposed as a task based programming model that exploits the parallelism at the task level and takes into account data dependencies between tasks. However, besides parallelism in the programming, the memory hierarchy impact in many/multi core architectures is a feature of large importance. This paper presents an evaluation of these two programming models with regard to the impact of different levels of the memory hierarchy in the duration of the application. The evaluation is based on trace-files with hardware counters on the execution of a memory intensive benchmark in both programming models.
Rosa M. Badia, Josep M. Pérez, Eduard Ayguadé, Jesús Labarta
PDP1
2009 Parallelizing dense and banded linear algebra libraries using SMPSs
abstract
Abstract The promise of future many‐core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP‐like pragmas and a run‐time system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS‐level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column‐major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run‐time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms‐by‐blocks in the FLAME argot) for most operations and storage‐by‐blocks (or block data layout) is already in place. Copyright © 2009 John Wiley & Sons, Ltd.
Rosa M. Badia, José R. Herrero 0001, Jesús Labarta, Josep M. Pérez, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2008 COMP Superscalar: Bringing GRID Superscalar and GCM Together
abstract
This paper presents the design, implementation and evaluation of COMP Superscalar, a new and componentised version of the GRID superscalar framework that enables the easy development of Grid-unaware applications. By means of a simple programming model, COMP Superscalar keeps the Grid as transparent as possible to the programmer. Moreover, the performance of the applications is optimized by exploiting their inherent concurrency when executing them on the Grid. The runtime of COMP Superscalar has been designed to follow the Grid Component Model (GCM) and is therefore formed by several components, each one encapsulating a given functionality identified in GRID superscalar.
Enric Tejedor, Rosa M. Badia
CCGRID2
2008 Prediction of behavior of MPI applications
abstract
Scalability and performance of applications is a very important issue today. As more complex have become high performance architectures, it is more complex to predict the behavior of a given application running on them. In this paper, we propose a methodology which automatically and quickly predicts, from a very limited number of runs using very few processors, the scalability and performance of a given application in a wide range of supercomputers taking into account details of the architecture and the network of the machines.
Marc Casas, Rosa M. Badia, Jesús Labarta
CLUSTER2
2008 A dependency-aware task-based programming environment for multi-core architectures
abstract
Parallel programming on SMP and multi-core architectures is hard. In this paper we present a programming model for those environments based on automatic function level parallelism that strives to be easy, flexible, portable, and performant. Its main trait is its ability to exploit task level parallelism by analyzing task dependencies at run time. We present the programming environment in the context of algorithms from several domains and pinpoint its benefits compared to other approaches. We discuss its execution model and its scheduler. Finally we analyze its performance and demonstrate that it offers reasonable performance without tuning, and that it can rival highly tuned libraries with minimal tuning effort.
Josep M. Pérez, Rosa M. Badia, Jesús Labarta
CLUSTER2
2008 SLA-Driven Semantically-Enhanced Dynamic Resource Allocator for Virtualized Service Providers
abstract
In order to be profitable, service providers must be able to undertake complex management tasks such as provisioning, deployment, execution and adaptation in an autonomic way. This paper introduces a framework, the Semantically-Enhanced Resource Allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide a full-customized and isolated virtual environment for each task. In addition, the system supports fine-grain dynamic resource distribution among these virtual environments based on SLAs. The required adaptation is implemented using agents, guarantying to each task enough resources to meet the agreed performance goals.
Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres
eScience6
2008 Integration of GRID Superscalar and GridWay Metascheduler with the DRMAA OGF Standard
Rosa M. Badia, D. Du, Eduardo Huedo, Antonis C. Kokossis, Ignacio Martín Llorente, Rubén S. Montero, Marc de Palol, Raül Sirvent, Constantino Vázquez
Euro-Par1
2008 Automatic analysis of speedup of MPI applications
abstract
The intricacy of high performance computing applications has been growing veryfast in the last years. Only skilled analysts are able to determine the factors that are undermining the performance of up-to-date applications. Analyst time is a very expensive resource and, for that reason, a strong effort to develop automatic performance analysis methodologies has been made by the scientific community. In this paper, we propose a methodology that is able to automatically detect the main performance problems of applications. This methodology is based on, first, a size reduction of the performance data obtained from the executions and, second, an analytical model obtained from this performance data which fits the speedup of the applications in terms of several parameters related to several performance issues. The paper also shows results obtained from real up-to-date applications and validates the conclusions automatically derived from the methodology.
Marc Casas, Rosa M. Badia, Jesús Labarta
ICS2
2008 Special section: Selected papers from the 7th IEEE/ACM international conference on grid computing (Grid2006)
Rosa M. Badia, Dennis Gannon, Craig A. Lee
Future Gener. Comput. Syst.1
2007 Topic 6 Grid and Cluster Computing
Rosa M. Badia, Christian Pérez, Artur Andrzejak 0001, Álvaro Enrique Arenas
Euro-Par1
2007 Automatic Structure Extraction from MPI Applications Tracefiles
Marc Casas, Rosa M. Badia, Jesús Labarta
Euro-Par2
2007 Improving Separation of Concerns in the Development of Scientific Applications
Seyed Masoud Sadjadi, T. Soldo, L. Atencio, Rosa M. Badia, Jorge Ejarque
SEKE5
2007 Performance of computationally intensive parameter sweep applications on Internet-based Grids of computers: the mapping of molecular potential energy hypersurfaces
abstract
Abstract This work focuses on the use of computational Grids for processing the large set of jobs arising in parameter sweep applications. In particular, we tackle the mapping of molecular potential energy hypersurfaces. For computationally intensive parameter sweep problems, performance models are developed to compare the parallel computation in a multiprocessor system with the computation on an Internet‐based Grid of computers. We find that the relative performance of the Grid approach increases with the number of processors, being independent of the number of jobs. The experimental data, obtained using electronic structure calculations, fit the proposed performance expressions accurately. To automate the mapping of potential energy hypersurfaces, an application based on GRID superscalar is developed. It is tested on the prototypical case of the internal dynamics of acetone. Copyright © 2006 John Wiley & Sons, Ltd.
Sebastián Reyes, Camelia Muñoz-Caro, Alfonso Niño, Rosa M. Badia, José María Cela
Concurr. Comput. Pract. Exp.4
2006 Including SMP in Grids as Execution Platform and Other Extensions in GRID Superscalar
abstract
GRID superscalar provides a very easy to use programming environment for enabling applications on the grid. Although the system already has many features, there are some areas that we wanted to enhance. In this paper we present a new version of GRID superscalar based on code annotations that includes full renaming support for scalar, array and structure parameters. We also present a tracing mechanism that allows fine tuning GRID superscalar applications, and improved support for running on SMP hosts.
Josep M. Pérez, Rosa M. Badia, Jesús Labarta
e-Science2
2006 Memory - CellSs: a programming model for the cell BE architecture
abstract
In this work we present Cell superscalar (CellSs) which addresses the automatic exploitation of the functional parallelism of a sequential program through the different processing elements of the Cell BE architecture. The focus in on the simplicity and flexibility of the programming model. Based on a simple annotation of the source code, a source to source compiler generates the necessary code and a runtime library exploits the existing parallelism by building at runtime a task dependency graph. The runtime takes care of the task scheduling and data handling between the different processors of this heterogeneous architecture. Besides, a locality-aware task scheduling has been implemented to reduce the overhead of data transfers. The approach has been implemented and tested with a set of examples and the results obtained since now are promising.
Pieter Bellens, Josep M. Pérez, Rosa M. Badia, Jesús Labarta
SC3
2006 Automatic Grid workflow based on imperative programming languages
abstract
Abstract GRID superscalar is a Grid programming environment that enables one to parallelize the execution of sequential applications in computational Grids. The run‐time library automatically builds a task data‐dependence graph of the application and it can be seen as an implicit workflow system. The current interface supports C/C++ and Perl applications. The run‐time library is based on Globus Toolkit 2.x using GRAM and GSIFTP services. In this document we describe the GRID superscalar basics emphasizing those aspects related to Grid workflow, in particular the flexibility of using an imperative language to describe the application. Copyright © 2005 John Wiley & Sons, Ltd.
Raül Sirvent, Josep M. Pérez, Rosa M. Badia, Jesús Labarta
Concurr. Comput. Pract. Exp.3
2006 System-level power-performance tradeoffs for reconfigurable computing
abstract
In this paper, we propose a configuration-aware data-partitioning approach for reconfigurable computing. We show how the reconfiguration overhead impacts the data-partitioning process. Moreover, we explore the system-level power-performance tradeoffs available when implementing streaming embedded applications on fine-grained reconfigurable architectures. For a certain group of streaming applications, we show that an efficient hardware/software partitioning algorithm is required when targeting low power. However, if the application objective is performance, then we propose the use of dynamically reconfigurable architectures. We propose a design methodology that adapts the architecture and algorithms to the application requirements. The methodology has been proven to work on a real research platform based on Xilinx devices. Finally, we have applied our methodology and algorithms to the case study of image sharpening, which is required nowadays in digital cameras and mobile phones
Juanjo Noguera, Rosa M. Badia
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Implementing phylogenetic inference with GRID superscalar
abstract
The grid has appeared recently as a new computing paradigm. However, to make the use of the grid available to the scientific community, frameworks that enable to easily write applications and to run them efficiently in the grid should be provided. GRID superscalar has been specially designed to satisfy the two requirements mentioned above. This paper presents an implementation of a biological application, fastDNAml, using GRID superscalar. The objective is not only to demonstrate the performance that can be achieved, but the programmability of the framework. The description contains details of the fastDNAml implementation, new features of GRID superscalar and summary of results obtained.
Vasilis Dialinos, Rosa M. Badia, Raül Sirvent, Josep M. Pérez, Jesús Labarta
CCGRID2
2005 Performance and Energy Analysis of Task-Level Graph Transformation Techniques for Dynamically Reconfigurable Architectures
abstract
In this paper, we present an analysis of the impact in both performance and energy of several task-level graph transformation techniques to exploit the parallel processing capabilities of run-time partially reconfigurable architectures. The proposed techniques have been applied to an image processing application (i.e., image sharpening), which has been implemented in a real research platform.
Juanjo Noguera, Rosa M. Badia
FPL2
2005 Data Distribution Strategies for Domain Decomposition Applications in Grid Environments
Beatriz Otero, José María Cela, Rosa M. Badia, Jesús Labarta
ICA3PP3
2004 Generation of Simple Analytical Models for Message Passing Applications
Rosa M. Badia, Jesús Labarta
Euro-Par2
2004 Multitasking on reconfigurable architectures: microarchitecture support and dynamic scheduling
abstract
Dynamic scheduling for system-on-chip (SoC) platforms has become an important field of research due to the emerging range of applications with dynamic behavior (e.g., MPEG-4). Dynamically reconfigurable architectures are an interesting solution for this type of applications. Scheduling for dynamically reconfigurable architectures might be classified in two major broad categories: (1) static scheduling techniques or (2) use of an operating system (OS) for reconfigurable computing. However, research efforts demonstrate a trend to move tasks traditionally assigned to the OS into hardware (thus increasing performance and reducing power).In this paper, we introduce a methodology for dynamically reconfigurable architectures. The dynamic scheduling of tasks to several reconfigurable units is performed by a hardware-based multitasking support unit. Two different versions of the microarchitecture are possible (with or without a hardware configuration prefetch unit). The dynamic scheduling algorithms are also explained. Both algorithms try to minimize the reconfiguration overhead by overlapping the execution of tasks with device reconfigurations.An exhaustive study (using the developed simulation and performance analysis framework) of this novel proposal is presented, and the effect of the microarchitecture parameters has been studied. Results demonstrate the benefits of our approach (achieving similar performance to a static configuration solution but using half of the resources). The hardware configuration prefetch unit is useful (i.e., minimize the execution time) in applications with low level of parallelism.
Juanjo Noguera, Rosa M. Badia
ACM Trans. Embed. Comput. Syst.2
2003 System-level power-performance trade-offs in task scheduling for dynamically reconfigurable architectures
abstract
Dynamic scheduling for System-on-Chip (SoC) platforms has become an important field of research due to the emerging range of applications with dynamic behavior (e.g. MPEG-4). Dynamically reconfigurable architectures are an interesting solution for this type of applications.However, dynamic scheduling for run-time reconfigurable architectures with power-performance trade-offs has not been addressed in previous research efforts. In this paper, we address this open issue using a system-level approach. Within our approach, we have used clock-gating and frequency-scaling strategies for power consumption minimization, jointly with our proposed architecture and scheduling algorithms.Device reconfiguration is a high-power consumption process. Thus reducing the number of device reconfigurations not only helps to reduce the reconfiguration overhead penalty (minimizing the application execution time), but also helps to reduce the system-level power consumption. Thus, dynamic task scheduling and reconfiguration context scheduling become a critical issue for power-performance trade-offs in embedded systems design.
Juanjo Noguera, Rosa M. Badia
CASES2
2003 Programming Grid Applications with GRID Superscalar
Rosa M. Badia, Jesús Labarta, Raül Sirvent, Josep M. Pérez, José María Cela, Rogeli Grima
J. Grid Comput.1
2002 A framework for performance modeling and prediction
abstract
Cycle-accurate simulation is far too slow for modeling the expected performance of full parallel applications on large HPC systems. And just running an application on a system and observing wallclock time tells you nothing about why the application performs as it does (and is anyway impossible on yet-to-be-built systems). Here we present a framework for performance modeling and prediction that is faster than cycle-accurate simulation, more informative than simple benchmarking, and is shown useful for performance investigations in several dimensions.
Allan Snavely, Laura Carrington, Nicole Wolter, Jesús Labarta, Rosa M. Badia, Avi Purkayastha
SC5
2002 HW/SW codesign techniques for dynamically reconfigurable architectures
abstract
Hardware/software (HW/SW) codesign and reconfigurable computing are commonly used methodologies for digital-systems design. However, no previous work has been carried out in order to define a HW/SW codesign methodology with dynamic scheduling for run-time reconfigurable architectures. In addition, all previous approaches to reconfigurable computing multicontext scheduling are based on static-scheduling techniques. In this paper, we present three main contributions: 1) a novel HW/SW codesign methodology with dynamic scheduling for discrete event systems using dynamically reconfigurable architectures; 2) a new dynamic approach to reconfigurable computing multicontext scheduling; and 3) a HW/SW partitioning algorithm for dynamically reconfigurable architectures. We have developed a whole codesign framework, where we have applied our methodology and algorithms to the case study of software acceleration. An exhaustive study has been carried out, and the obtained results demonstrate the benefits of our approach.
Juanjo Noguera, Rosa M. Badia
IEEE Trans. Very Large Scale Integr. Syst.2
2001 A HW/SW partitioning algorithm for dynamically reconfigurable architectures
abstract
"System-On-Chip" has become a reality, and recently new reconfigurable devices have appeared. However, few efforts have been carried out in order to define HW/SW codesign methodologies and algorithms which address the challenges presented by new reconfigurable devices. In this paper we address this open problem and present a novel HW/SW partitioning algorithm for dynamically reconfigurable architectures. The algorithm is a constructive algorithm, which obtains an initial solution and afterwards tries to optimize it. The HW/SW partitioning is done taking into account the features of the dynamically reconfigurable devices, and its final goal is to minimize the reconfiguration latency. The partitioning algorithm has been implemented and integrated into our developed codesign environment, where several experiments have been carried out. The results obtained demonstrate the benefits of the algorithm.
Juanjo Noguera, Rosa M. Badia
DATE2
1999 Optimal exploration of the unrolling degree for software pipelining
Fermín Sánchez, Jordi Cortadella, Rosa M. Badia
J. Syst. Archit.3
1993 Glass: a graph-theoretical approach for global binding
Rosa M. Badia, Jordi Cortadella
Microprocess. Microprogramming1
1991 Scheduling in a continuous area-time design space
Jordi Cortadella, Rosa M. Badia, Eduard Ayguadé
Microprocessing and Microprogramming2