EDBT 2026 Demo / reviewers in the wild / expert
Jesús Carretero 0001
dblp:c/JCarretero · also Jesús Carretero Pérez
· DBLP profile ↗
160ranked-venue papers
17as first author
41since 2021 · last 2026
0000-0002-1413-4793ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 103 · 14 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Computer networks · 10Software engineering, systems software and programming languages · 10 · 3 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OrbiFaaS: An orbital method to build Continuum Earth Observation SystemsabstractEarth observation (EO) problems require processing large volumes of data, such as those involved in environmental studies using diverse spatiotemporal variables. In this context, the computing continuum offers a promising model to support EO tasks. It enables low-latency data processing near ground stations and broad data sharing through the cloud. Nevertheless, designing systems in the computing continuum introduces challenges. One is the coordination of distributed entities across different computational environments (e.g., the edge, the fog, or the cloud). Another is managing data across heterogeneous infrastructures. This paper presents OrbiFaaS, a serverless-based method for building Continuum Earth Observation Systems (EOS) using an orbital model. OrbiFaaS organizes a continuum EOS as a system of multiple orbits, each representing a different computational environment arranged around a core of data sources. Orbits near the core correspond to low-latency environments (e.g., the edge), while those farther away exhibit higher latencies (e.g., the cloud). Each orbit contains satellites, which represent service provider infrastructures. Organizations can use these satellites to deploy multiple data management services, such as containerized microservices or functions. Satellites are interconnected using data containers that establish communication channels based on available filesystem, memory, and network resources. We evaluate OrbiFaaS in a case study involving the processing of satellite images across multiple environments. Our experimental results demonstrate the feasibility and efficiency of OrbiFaaS for constructing Earth observation systems. Catherine Alessandra Torres Charles, Dante D. Sánchez-Gallegos, Diana Carrizales-Espinoza, José Luis González 0002, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 5 |
| 2026 | HERCULES: A scalable and elastic ad-hoc file system for large-scale computing systemsabstractThe increasing demand for data processing by new, data-intensive applications is placing significant strain on the performance and capacity of HPC storage systems. Advancements in storage technologies, such as NVMe and persistent memory, have been introduced to address these demands. However, relying exclusively on ultra-fast storage devices is not cost-effective, necessitating multi-tier storage hierarchies to manage data based on its usage. In response, ad-hoc file systems have been proposed as a solution. These systems use the storage resources available in compute nodes, including memory and persistent storage, to create temporary file systems that adapt to application behavior in the HPC environment. This work presents the design, implementation, and evaluation of HERCULES, a distributed ad-hoc in-memory storage system, with a focus on its new metadata and elasticity model. HERCULES takes advantage of the Unified Communication X (UCX) framework, leveraging RDMA protocols such as Infiniband, Omnipath, shared-memory, and zero-copy transfers for data transfer. It includes elasticity features at runtime and fault-tolerant facilities. The elasticity features, together with flexible policies for data allocation, allow HERCULES to migrate data so that the available resources can be efficiently used. Our exhaustive evaluation results demonstrate a better performance than Lustre and BeeGFS, two parallel file systems heavily used in High-Performance Computing systems, and GekkoFS, an ad-hoc state-of-the-art solution. Genaro Sanchez-Gallegos, Cosmin Petre, Francisco Javier García Blas, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 4 |
| 2026 | Nez: A design-driven skeleton model for building continuum AI-based and analytic systemsabstractContext: Organizations increasingly rely on artificial intelligence (AI) and machine learning (ML) to process data, automate tasks, and enhance decision-making. At the same time, the computing continuum enables AI and ML to be deployed closer to data sources, thereby reducing system latency and response time. Objective: Managing applications across this distributed environment is challenging due to the need for manual deployment, integration, and compliance with non-functional requirements (NFRs) such as security and fault tolerance. Therefore, there is a need for frameworks that automate the deployment and execution of computing continuum systems while integrating both functional and non-functional requirements. Method: This paper presents Nez , a design-driven skeleton model for building continuum AI and analytics systems. Nez construction model automatically and transparently integrates AI/ML applications with non-functional components to create continuum systems that are deployed dynamically across multiple distributed infrastructures. Results: We conducted case studies on the processing of medical imagery and satellite imagery to provide automatic and continuous support for decision-makers. Nez has already been deployed at the Mexican hospital, Instituto Nacional de Rehabilitación Luis Gerardo Ibarra Ibarra , to create an AI-based data flow supporting bone cancer diagnosis. The evaluation shows that Nez outperforms state-of-the-art tools such as Nextflow, Makeflow, and Parsl, achieving improvements in response time of 28.46%, 17.46%, and 23.54%, respectively. Conclusion: Nez efficiently transforms organizational data flow designs into continuum computing services. This enables organizations to construct continuum AI-based and analytical systems that account for both functional and non-functional requirements. Dante D. Sánchez-Gallegos, Diana Carrizales-Espinoza, José Luis González 0002, Marco Antonio Núñez-Gaona, Heriberto Aguirre-Meneses, Jesús Carretero 0001 |
Inf. Softw. Technol. | 6 |
| 2025 | Detection of Active and Latent Tuberculosis with Explainable Deep Learning EnsemblesabstractTuberculosis is one of the deadliest diseases in the world, despite being treatable, the number of new cases increases yearly. The situation is even more concerning due to antimicrobial resistance (AMR). To avoid undetected cases and breaking transmission, the research community is using artificial intelligence (AI) to develop new and fast diagnosis tools. Nevertheless, the creation of good quality and well-balanced datasets for training AI tools is challenging. It is especially difficult to collect data on asymptomatic patients who do not go to the professional as they do not feel the symptoms, such as the one who suffered latent tuberculosis. In this work, we propose an explainable deep learning ensemble of convolutional neural networks (CNNs) to classify tuberculosis chest X-ray (CXRs) images. The ensemble was designed to alleviate being influenced by unbalanced data due to the small number of latent patients included in the dataset. The system was trained to detect tuberculosis against healthy patients and other diseases with CXRs as the only input, so as to inform tuberculosis patients of the stage of the disease (active or latent). The model was designed with two parallel CNNs in an ensemble that used a random forest (RF) to overcome the imbalance. The system reported a performance of 96 % accuracy, showing that the ensemble can improve the performance of the underrepresented class. Further, the RF completed the interpretability of the system provided by grad-CAM heatmaps, supporting the evolution and integration of the system in clinical environments. Lara Visuña, Francisco Javier García Blas, Jesús Carretero 0001 |
CBMS | 3 |
| 2025 | Dynostore: A Wide-Area Distribution System for the Management of Data Over Heterogeneous StorageabstractData distribution across different facilities offers benefits such as enhanced resource utilization, increased resilience through replication, and improved performance by processing data near its source. However, managing such data is challenging due to heterogeneous access protocols, disparate authentication models, and the lack of a unified coordination framework. This paper presents DynoStore, a system that manages data across heterogeneous storage systems. At the core of DynoStore are data containers, an abstraction that provides standardized interfaces for seamless data management, irrespective of the underlying storage systems. Multiple data container connections create a cohesive wide-area storage network, ensuring resilience using erasure coding policies. Furthermore, a load-balancing algorithm ensures equitable and efficient utilization of storage resources. We evaluate DynoStore using benchmarks and realworld case studies, including the management of medical and satellite data across geographically distributed environments. Our results demonstrate a 10 % performance improvement compared to centralized cloud-hosted systems while maintaining competitive performance with state-of-the-art solutions such as Redis and IPFS. DynoStore also exhibits superior fault tolerance, withstanding more failures than traditional systems. Dante D. Sánchez-Gallegos, José Luis González 0002, Maxime Gonthier, Valérie Hayot-Sasson, J. Gregory Pauloski, Haochen Pan, Kyle Chard, Jesús Carretero 0001, Ian T. Foster |
CCGrid | 8 |
| 2025 | D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed StorageabstractThe exponential growth of data necessitates distributed storage models, such as peer-to-peer systems and data federations.While distributed storage can reduce costs and increase reliability, the heterogeneity in storage capacity, I/O performance, and failure rates of storage resources makes their efficient use a challenge.Further, node failures are common and can lead to data unavailability and even data loss. Maxime Gonthier, Dante D. Sánchez-Gallegos, Haochen Pan, Bogdan Nicolae, Hai Nguyen 0005, Valérie Hayot-Sasson, J. Gregory Pauloski, Jesús Carretero 0001, Kyle Chard, Ian T. Foster |
ICS | 9 |
| 2025 | A-Flow: managing dataflows on the computing continuum using abstract communication channelsabstractComputing continuum systems are emerging as a solution for organizations to process data across diverse infrastructures, reducing latency compared to traditional cloud computing. However, managing I/O operations in such distributed and heterogeneous environments remains an open research challenge. In this paper, we present A-Flow, a model for constructing I/O systems to manage data exchange in computing continuum environments. These systems are built around data distribution patterns defined by structures called abstract communication channels (ACCs). ACCs are established between processing stages using memory, file system, and network resource connections. To prevent resource overload during execution, A-Flow automatically selects the appropriate communication channel based on user-defined criteria such as throughput or resource utilization. We implemented this model in a prototype and evaluated it through a case study focused on managing medical data in HDF5 format across different environments. The evaluation revealed that A-Flow’s ACCs can be integrated into existing stage-based systems found in the state of the art. The results highlight the efficiency and effectiveness of A-Flow in enabling dataflows across heterogeneous infrastructures, addressing key challenges in computing continuum. Catherine Alessandra Torres Charles, Dante D. Sánchez-Gallegos, Diana Carrizales-Espinoza, José Luis González 0002, Jesús Carretero 0001 |
SBAC-PAD | 5 |
| 2025 | Automated monitoring of Mycobacterium population growth in time-lapse microscopy with Deep Learning pseudo-supervised algorithmabstractThe understanding of bacteria evolution through time and the analysis of their interaction with the medium is a main step in biomedical research. Bacteria cause in humans different infectious diseases, such as tuberculosis. Tuberculosis is one of the leading causes of death worldwide, caused by Mycobacterium tuberculosis ( Mtb ). Observation of Mycobacteria under a microscope provides essential information for disease-fighting. Nevertheless, Mtb presents a very particular cell division pattern characterized by a high entanglement in which the frontier between dividing cells are almost indistinguishable to the human eye, which makes the manual analysis of cultures a very challenging task, resulting in extremely long processing times. In this paper, we propose a novel approach to automatically detect bacterial cells and infer the growth rate and growth speed over a time-lapse microscopy (TLM) sequence. This method leverages a semi-supervised Deep Learning approach with a lightweight U-Net for efficient and fast processing. To reduce efforts of database construction, we deploy and exploit a mix of manually and synthetically annotated databases of two different classes of Mycobacterium , i.e. Mtb and Mycobacterium smegmatis ( Msm ). The experiments demonstrated the efficiency of the detection of both bacterial reproduction rate and bacterial death in Mycobacterium species, with the Phase-contrast (PhC) microscopy channels as a unique input. Reaching a segmentation accuracy of 96.3%, recall of 78.2%, and precision of 85.5% on Mtb detection for the model trained with a mix of manually and synthetically annotated databases. The experiments validate as well the synthetic label generation technique, which provides more conscience and accurate results than the manually segmented techniques. The mask generation workflow opens new opportunities for the microbial AI development field. The system implemented may change the paradigm in preclinical anti-mycobacteria drug development for since enables the adoption of TLM, a technology that provides extremely valuable information on drug efficacy but is seldom leveraged due to the complex image processing it involves. • A semi-supervised Deep Learning U-Net based model for TLM bacterial culture analysis. • The method reports bacterial growth and death with PhC as a single input. • Evaluation in two different Mycobacterium ( M. tuberculosis and M. smegmatis ). • Creating a methodology for synthetic annotation of the bacterial population in TLM. • Comparative analysis of manual and synthetic databases. Lara Visuña, Francisco Javier García Blas, Santiago Ferrer-Bazaga, Patricio Lopez-Exposito, Jesús Carretero 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | An avatar cloud service based method for supervising and interacting with containerized applications
Juan Armando Barron-Lugo, Ivan López-Arévalo, José Luis González 0002, Jose Carlos Morin Garcia, Melesio Crespo-Sanchez, Jesús Carretero 0001 |
Expert Syst. Appl. | 6 |
| 2025 | A comparative study of ad-hoc file systems for extreme scale computing
Njoud O. Almaaitah, Francisco Javier García Blas, Genaro Sanchez-Gallegos, Jesús Carretero 0001, Marc-Andre Vef, André Brinkmann |
Future Gener. Comput. Syst. | 4 |
| 2025 | Edge-cloud solutions for big data analysis and distributed machine learning - 2abstractIn recent years, edge-cloud solutions have gained widespread adoption for efficiently collecting and analyzing IoT-generated data across various domains like urban mobility, healthcare, and smart cities. These solutions integrate resources from edge to cloud to support real-time processing and analysis tasks, reducing latency and network congestion. Big data analysis within this paradigm involves sophisticated techniques for distributed data processing, enabling applications such as predictive maintenance and smart grid management. Nevertheless, carrying out big data analysis within the edge-cloud presents several challenges, including data privacy and security, interoperability, scalability, and energy efficiency. Addressing these challenges is imperative for providing efficient and scalable solutions for data-intensive applications like federated learning, social data analysis, smart city services, and text mining. The special issue concludes with 27 scientific papers, divided into two parts for a streamlined editorial process. This editorial, as part two, presents 12 rigorously peer-reviewed papers, complementing the 15 papers covered in the previous editorial. Loris Belcastro, Jesús Carretero 0001, Domenico Talia |
Future Gener. Comput. Syst. | 2 |
| 2025 | Improving I/O performance in HPC environments using the Expand Ad-Hoc file systemabstractThis work introduces an exhaustive evaluation, including both benchmarks and real-world data-intensive applications, performed in MareNostrum 4 and HPC4AI Laboratory supercomputers using the Expand Ad-Hoc file system. Expand Ad-Hoc is an ad-hoc file system that dynamically virtualizes the local storage available on compute nodes (i.e., SSD, SHM, etc.) into a fast storage volume to reduce congestion on parallel file systems used as backends in High-Performance Computing (HPC) environments. The main contributions of this work include the design of a new ad-hoc parallel file system, called Expand Ad-Hoc , compatible with POSIX and MPI-IO, and an exhaustive and comprehensive evaluation of Expand. The evaluation compares the performance obtained using IOR and DLIO benchmarks and Nek5000 and Remote Sensing real-world applications on Expand Ad-Hoc , GekkoFS, GPFS, and BeeGFS, showing satisfactory results proving that Expand Ad-Hoc can be used in HPC environments to improve the I/O performance of data-intensive applications transparently. Diego Camarmas-Alonso, Félix García Carballeira, Alejandro Calderón 0001, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2025 | Guest Editorial: New Tools and Techniques for the Distributed Computing Continuum
Jesús Carretero 0001, Francisco Javier García Blas, Sameer Shende |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | I/O Behind the Scenes: Bandwidth Requirements of HPC Applications with Asynchronous I/OabstractI/O bandwidth is a critical resource in an HPC cluster. As with all shared resources, its availability is impacted significantly by the users and the applications they execute. Without proper restrictions, jobs consuming more prominent portions of the I/O bandwidth can severely affect other jobs by notably prolonging their runtime. In such a context, applications that perform asynchronous I/O bring unique properties that allow for the reduction of such effects. That is, by limiting the bandwidth to the required one to perform the I/O in the background of the compute phases, I/O bursts can be flattened without significantly prolonging the application time, if at all. Hence, the bandwidth consumption of such applications is limited to what they need, sparing much of the system bandwidth to other applications. At the same time, these applications achieve higher parallel efficiency due to the overlapping of different resources (e.g., compute and I/O). This paper shows these aspects and demonstrates our approach to finding the required bandwidth for applications that use asynchronous I/O. Moreover, we apply it automatically using an MPI implementation of a bandwidth limitation approach at the application level. We validate our approach with several experiments on a large production cluster and show the impact of our approach on the application behavior and its importance for the system throughput. Ahmad Tarraf, Javier Fernández 0001, David E. Singh, Taylan Özden, Jesús Carretero 0001, Felix Wolf 0001 |
CLUSTER | 5 |
| 2024 | Fault Tolerant in the Expand Ad-Hoc Parallel File SystemabstractAbstract In the last years, applications related to Artificial Intelligence and big data, among others, have been involved. There is a need to improve I/O operations to avoid bottlenecks in accessing a larger amount of data. For this purpose, the Expand Ad-Hoc parallel file system is being designed and developed. Since these applications have very long execution times, fault tolerance mechanisms in the file system are necessary to allow them to continue running in the presence of failures. This work introduces a fault-tolerant design based on data replication for the Expand Ad-Hoc parallel file system and an initial evaluation conducted on the HPC4AI Laboratory supercomputer in Torino. The evaluation of Expand Ad-Hoc with fault-tolerant found that, despite data replication, its performance and scalability are generally better than those of other parallel file systems without fault-tolerant. Dario Muñoz-Muñoz, Félix García Carballeira, Diego Camarmas-Alonso, Alejandro Calderón 0001, Jesús Carretero 0001 |
Euro-Par (2) | 5 |
| 2024 | Edge-Cloud Solutions for Big Data Analysis and Distributed Machine Learning - 1
Loris Belcastro, Jesús Carretero 0001, Domenico Talia |
Future Gener. Comput. Syst. | 2 |
| 2024 | Cluster and cloud computing for life sciences
Jesús Carretero 0001, Dagmar Krefting |
Future Gener. Comput. Syst. | 1 |
| 2024 | StructMesh: A storage framework for serverless computing continuumabstractComputing continuum is becoming a solution for organizations to process and analyze data for supporting decision-making processes. In this context, serverless paradigm is arising as a solution to manage continuum computing. However, the management of data storage still represents an obstacle for integrating continuum computing and serverless paradigms into a single solution, as this has to be performed transparently to users through multiple infrastructures. This paper presents StructMesh, a storage framework for serverless continuum systems. This framework is based on a processing plane where functions are managed as patterns, and a data plane based on storage meshes that represent maps of storage resources available in a given infrastructure. The logical interconnection of storage meshes enables organizations to integrate storage resources into a single unified storage service, which creates data exchange channels for continuum processing throughout multiple infrastructures. These meshes include load-balancing and data allocation/location algorithms for transparently and automatically managing the inputs/outputs of functions throughout these channels, as well as non-functional requirement schemes for organizations to manage sensitive data. We developed a framework prototype that harmonizes processing serverless functions with storage functions for building serverless pipeline services. A case study was conducted by using these services for processing meteorological and earth observation data throughout multiple infrastructures. The evaluation revealed the efficiency of StructMesh when managing data through fog and cloud infrastructures. It also showed the feasibility of StructMesh to create and enable continuum data exchange channels for serverless pipelines. Diana Carrizales-Espinoza, Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 4 |
| 2024 | Performance-driven scheduling for malleable workloadsabstractAbstract The development of adaptive scheduling algorithms that take advantage of malleability has become a crucial area of research in many large-scale projects. Malleable workloads can improve the system’s performance but, at the same time, provide an extra dimension to the scheduling problem. This paper proposes an adaptive, performance-based job scheduling method that emphasizes the backfilling concept with malleability. The proposed method performs the malleability operations only when the estimated execution time of the involved applications is better than or equal to the execution time on the allocated resources without reconfiguration. The reconfiguration feasibility is determined by performance models considering the application scalability and reconfiguration overheads. Different policies for implementing malleability are presented, each targeting a specific workload in terms of job size and scalability. The comprehensive evaluation shows an improvement in the slowdown up to 49% compared to the non-adaptive baseline scheduling algorithm. Njoud O. Almaaitah, David E. Singh, Taylan Özden, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2024 | Detailed parallel social modeling for the analysis of COVID-19 spreadabstractAbstract Agent-based epidemiological simulators have been proven to be one of the most successful tools for the analysis of COVID-19 propagation. The ability of these tools to reproduce the behavior and interactions of each single individual leads to accurate and detailed results, which can be used to model fine-grained health-related policies like selective vaccination campaigns or immunity waning. One characteristic of these tools is the large amount of input data and computational resources that they require. This relies on the development of parallel algorithms and methodologies for generating, accessing, and processing large volumes of data from multiple data sources. This work presents a parallel workflow for extending the social modeling of EpiGraph, an agent-based simulator. We have included two novel parallel social generation stages that generate a detailed and realistic social model and one new visualization stage. This work also presents a description of the algorithms used in each stage, different optimization techniques that permit to reduce the application convergence time, and a practical evaluation of large workloads on HPC systems. Results show that this contribution can be efficiently executed in parallel architectures and the results allow to increase the simulation detail level, representing a significant advance in the simulator scenario modeling. As a summary of results, the first contribution of this paper is the development of two models (a spatial and a social one) that assign geographical and socioeconomic indicators to each simulated individual (i.e., agents), reproducing the real social distribution of the city of Madrid. The second contribution presents an improved parallel and distributed algorithm that executes the two aforementioned models using different parallelization strategies and preserving the load balance. Aymar Cublier Martínez, Jesús Carretero 0001, David E. Singh |
J. Supercomput. | 2 |
| 2024 | Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future OpportunitiesabstractWith the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research. Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001 |
IEEE Trans. Parallel Distributed Syst. | 23 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 1 |
| 2023 | Exascale programing models for extreme data processingabstractExtreme Data is an incarnation of Big Data concept distinguished by the massive amounts of data that must be queried, communicated and analysed in (near) real-time by using a very large number of memory/storage elements and Exascale computing systems. Immediate examples are the scientific data produced at a rate of hundreds of gigabits-per-second that must be stored, filtered and analysed, the millions of images per day that must be mined/analyzed in parallel, the one billion of social data posts queried in real-time on an in-memory components database. Traditional disks or commercial storage cannot handle nowadays the extreme scale of such application data. Following the need of improvement of current concepts and technologies, ASPIDE's activities focus on data-intensive applications running on systems composed of up to millions of computing elements (Exascale systems). Practical results will include the methodology and software prototypes that will be designed and used to implement Exascale applications. Francisco Javier García Blas, Jesús Carretero 0001 |
CF | 2 |
| 2023 | eScience Serverless Data Storage Services in the Edge-Fog-Cloud ContinuumabstractMeshStore, a fault-tolerant serverless storage model for edge-fog-cloud continuum systems, enables organizations to integrate distributed heterogeneous storage resources into a single unified storage service for the sharing of data through serverless functions deployed on edge-fog-cloud environments to create continuum dataflows. This unified service automatically and transparently manages the input/output data of serverless functions by coupling storage structures, including load-balancing and data allocation/location algorithms. Organizations also can add non-functional requirement properties (e.g., either reliability or security) to the storage structures when managing sensitive data. Dante D. Sánchez-Gallegos, Diana Carrizales-Espinoza, José Luis González 0002, Jesús Carretero 0001 |
e-Science | 4 |
| 2023 | Hercules: Scalable and Network Portable In-Memory Ad-Hoc File System for Data-Centric and High-Performance Applications
Francisco Javier García Blas, Genaro Sanchez-Gallegos, Cosmin Petre, Alberto Riccardo Martinelli, Marco Aldinucci, Jesús Carretero 0001 |
Euro-Par | 6 |
| 2023 | Dynamic management of processes and communicators in malleable MPI applicationsabstractA malleable application is defined as one that can increase or decrease its resources dynamically based on workload variations. These applications often leverage a job manager to handle the resources. The MPI standard incorporates strategies to increase the number of processes connected to a running application. However, it does not clarify how to use these same features to reduce said processes. This paper presents a strategy compatible with the latest version of the MPI standard to develop malleable MPI applications. It allows the addition and removal of resources at the computing node level in a discretionary manner. The results obtained from the evaluation show that the proposed strategy has acceptable performance and scalability for the functionality it provides. Javier Fernández 0001, Alberto Cascajo, Jesús Carretero 0001 |
ICPADS | 3 |
| 2023 | Fine-grained parallel social modelling for analyzing COVID-19 propagationabstractAgent-based epidemiological simulators have been proven to be one of the most successful tools for the analysis of the COVID-19 propagation. The ability of these tools to reproduce the behavior and interactions of each single individual leads to accurate and detailed results, that can be used to model fine-grained health-related policies like selective vaccination campaigns or immunity waning. One characteristic of these tools is the large amount of input data and computational resources and that they require. This relies on the development of parallel algorithms and methodologies for generating, accessing and processing large volumes of data from multiple data sources. This work presents a parallel workflow for extending the social modelling of EpiGraph, an agent-based simulator. We have included two novel parallel social generation stages -that provide detailed and realistic social model- and one new visualization stage. The work presents a description of the algorithms used in each stage and a practical evaluation on a real platform. Results show that this contribution can be efficiently executed in parallel architectures and increases the simulation detail level, representing a significant advance in the simulator scenario modelling. Aymar Cublier Martínez, Álejandro Alvarez Isabel, Jesús Carretero 0001, David E. Singh |
PDP | 3 |
| 2023 | Blockchain-based schemes for continuous verifiability and traceability of IoT dataabstractThis paper presents a continuous delivery/continuous verifiability (CD/CV) framework for IoT dataflows in edge-fog-cloud. In this framework a CD model based on extraction, transformation, and load (ETL) mechanism as well as a directed acyclic graph (DAG) construction, enable end-users to create efficient schemes for the continuous verification and validation of the execution of applications in edge-fog-cloud infrastructures. This framework also provides tools for continuous verification and validation (CV) of predefined execution sequences and the integrity of digital assets using blockchain. CV model converts ETL and DAG into business model, smart contracts in a private blockchain for the automatic and transparent registration of transactions performed by each application in workflows/pipelines created by CD model without altering applications nor edge-fog-cloud workflows. This framework ensures that IoT dataflow delivers verifiable information for organizations to conduct critical decision-making processes with certainty. A containerized parallelism approach solves portability issues and reduces/compensates the overhead produced by CD/CV operations. The talk will also present evaluation results of the CD/CV framework based on a case study where user mobility information is used to identify interest points, patterns, and maps. The experimental evaluation results show the feasibility of CD/CV to register transactions performed in IoT dataflows through edge-fog-cloud in a private blockchain network. Cristhian Martinez-Rendon, José Luis González 0002, Dante D. Sánchez-Gallegos, Jesús Carretero 0001 |
PDP | 4 |
| 2023 | A novel approach for large-scale environmental data partitioning on cloud and on-premises storage for compute continuum applicationsabstractSummary Cloud‐based services have proved useful in several research fields, such as engineering, health science, and astrophysics, to mention a few examples. The computational environmental science community developed a strong need for cloud facilities to store, process, and manage data from observations and numerical models for simulations and forecasts. Weather forecast models and global sensor networks deal with multidimensional geo‐referenced data∖sets. However, environmental data consumer applications usually require a relatively small amount of multidimensional input data slice to analyze a specific area or time interval. Hence, reducing data dimension for information retrieval is mandatory. This paper presents a twofold solution: a technique to load and retrieve the sliced multidimensional data set on different cloud services such as Amazon Web Service (AWS), Google Cloud Platform, and Microsoft Azure. The experimental results performed on these cloud services highlight that the proposed method can significantly speed up the process of loading and retrieving the data slices compared to working with the entire data set in bulk or OPeNDAP server. Gennaro Mellone, Ciro Giuseppe De Vita, Dante D. Sánchez-Gallegos, Genaro Sanchez-Gallegos, Catherine Alessandra Torres Charles, Francisco Javier García Blas, Jesús Carretero 0001, José Luis González 0002, Giuliano Laccetti |
Concurr. Comput. Pract. Exp. | 7 |
| 2023 | Xel: A cloud-agnostic data platform for the design-driven building of high-availability data science services
Juan Armando Barron-Lugo, José Luis González 0002, Ivan López-Arévalo, Jesús Carretero 0001, José-Lázaro Martínez-Rodríguez |
Future Gener. Comput. Syst. | 4 |
| 2023 | Evaluating the spread of Omicron COVID-19 variant in SpainabstractThis work analyzes the propagation the highly transmissible COVID-19 variant Omicron across Spain via simulation by using EpiGraph. EpiGraph is an agent-based parallel simulator that reproduces the COVID-19 propagation over wide areas. In this work we consider a population of 19,574,086 individuals of the 63 most populated cities of Spain, for the time interval between May 15th 2021 and March 6th 2022. The main variants existing at the start of the simulation were the Alpha and Delta, with prevalence of 4% and 96%. Then, during the second half of November 2021, the Omicron variant appears in Spain. Due to the higher transmission of this new variant - about 2 times larger than Delta, it quickly spreads through all the cities and becomes the dominant strain in the country. In this work we analyze the propagation of this variant under different mobility restrictions and patient zero scenarios. We first define a baseline scenario which reproduces the existing conditions of the COVID-19 propagation in Spain for our period of study. We then consider alternative scenarios for different starting locations of the propagation. Finally, for each one of these scenarios, we evaluate different transportation intensities - i.e. movement of individuals between the cities. The main conclusion is that, independently of the initial location of the Omicron variant and the existing transportation conditions, the Omicron variant spreads through all the country in a short time interval. The work presented in this paper also implements and evaluates a power monitoring and optimization system aimed at reducing the energy consumption of such massive simulations as the ones performed in EpiGraph. Miguel Guzmán-Merino, Maria-Cristina V. Marinescu, Alberto Cascajo, Jesús Carretero 0001, David E. Singh |
Future Gener. Comput. Syst. | 4 |
| 2023 | On the building of efficient self-adaptable health data science services by using dynamic patterns
Genaro Sanchez-Gallegos, Dante D. Sánchez-Gallegos, José Luis González 0002, Hugo G. Reyes-Anastacio, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 5 |
| 2023 | CD/CV: Blockchain-based schemes for continuous verifiability and traceability of IoT data for edge-fog-cloudabstractThis paper presents a continuous delivery/continuous verifiability ( CD/CV ) method for IoT dataflows in edge–fog–cloud. A CD model based on extraction, transformation, and load (ETL) mechanism as well as a directed acyclic graph ( DAG ) construction, enable end-users to create efficient schemes for the continuous verification and validation of the execution of applications in edge–fog–cloud infrastructures. This scheme also verifies and validates established execution sequences and the integrity of digital assets . CV model converts ETL and DAG into business model, smart contracts in a private blockchain for the automatic and transparent registration of transactions performed by each application in workflows/pipelines created by CD model without altering applications nor edge–fog–cloud workflows. This model ensures that IoT dataflows delivers verifiable information for organizations to conduct critical decision-making processes with certainty. A containerized parallelism model solves portability issues and reduces/compensates the overhead produced by CD/CV operations. We developed and implemented a prototype to create CD/CV schemes, which were evaluated in a case study where user mobility information is used to identify interest points, patterns, and maps. The experimental evaluation revealed the efficiency of CD/CV to register the transactions performed in IoT dataflows through edge–fog–cloud in a private blockchain network in comparison with state-of-art solutions. Cristhian Martinez-Rendon, José Luis González 0002, Dante D. Sánchez-Gallegos, Jesús Carretero 0001 |
Inf. Process. Manag. | 4 |
| 2023 | PuzzleMesh: A Puzzle Model to Build Mesh of Agnostic Services for Edge-Fog-CloudabstractThis paper presents the design, development, and evaluation of PuzzleMesh, an agnostic service mesh composition model to process large volumes of data in edge-fog-cloud environments. This model is based on a puzzle metaphor where pieces, puzzles, and metapuzzles represent self-contained autonomous and reusable software artifacts encapsulated into containers and published as microservices. Apiecerepresents the integration of apps with I/O interfaces (loops/sockets), parallel processing, and management software. Apuzzlerepresents a processing structure (e.g., workflows) built coupling pieces through loops and sockets. Puzzles integrate structures with a microservice architecture, implicit continuous dataflows, and transparent data exchange management software. Ametapuzzlerepresents a recursive assemble of puzzles. A mesh represents a pool of pieces, puzzles, and metapuzzles available for designers to choose artifacts to build services. A prototype developed using PuzzleMesh model was evaluated through case studies about the automatic construction of processing services for the acquisition, pre-processing, manufacturing, preserving, and visualizing of satellite imagery. A qualitative comparison revealed that PuzzleMesh provides a flexible way to build reusable and portable services and to improve the usability of the services. The case study also revealed that PuzzleMesh yielded better performance results than other state-of-the-art tools. Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001, Heidy Marisol Marín-Castro, Andrei Tchernykh, Raffaele Montella |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | On the building of self-adaptable systems to efficiently manage medical dataabstractThe systems that meet non-functional requirements (NFRs) are key for e-health services to face up events such as service outages and violations of confidentiality. How-ever, traditional NFR systems produce overhead in execution time, which could affect critical decision-making processes. This paper presents a dynamic parallel pattern construction model to design and create efficient NFR self-adaptable sys-tems. The construction of patterns is performed in two design phases: in the first one, the designers build NFR systems by creating pipelines including as many applications as required to meet the NFRs established by healthcare organizations. In the second phase, a pipeline is converted into a worker that auto-matically is added to a dynamic pattern. In a dynamic pattern, the workers can be cloned to be executed by different parallel patterns (e.g., manager/worker, divide&conquer, etc.) to face changes in the incoming workload during execution time, which converts a worker into a self-adaptable NFR system. A proto-type was implemented to create self-adaptable NFR systems, which were used in a case study to manage spirometry studies, tomography images, and electrocardiograms. The evaluation showed the effectiveness of this dynamic pattern model to create self-adaptable systems when processing multiple types of medical data/contents. The evaluation also revealed that the self-adaptable NFR systems built by dynamic patterns yielded significant performance gain in a direct comparison with the implementation of NFR application pipelines built by a traditional framework called Jenkins. Genaro Sanchez-Gallegos, Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001 |
CCGRID | 4 |
| 2022 | Improving Congestion Control through Fine-Grain Monitoring of InfiniBand NetworksabstractCongestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms. Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001 |
HOTI | 8 |
| 2022 | SeRSS: a storage mesh architecture to build serverless reliable storage servicesabstractCloud storage has been the solution for organizations to manage the exponential growth of data observed over the past few years. However, end-users still suffer from side-effects of cloud service outages, which particularly affect edge-fog-cloud environments. This paper presents SeRSS, a storage mesh architecture to create and operate reliable, configurable, and flexible serverless storage services for heterogeneous infrastructures. A case study was conducted based on-the-fly building of storage services to manage medical imagery. The experimental evaluation revealed the efficiency of SeRSS to manage and store data in a reliable manner in heterogeneous infrastructures. Diana Carrizales-Espinoza, Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001, Ricardo Marcelín-Jiménez |
PDP | 4 |
| 2022 | Convergence of HPC and Big Data in extreme-scale data analysis through the DCEx programming modelabstractHigh-level programming models can help application developers to access and use resources without the need to manage low-level architectural entities, as a parallel programming model defines a set of programming abstractions that simplify the way by which a programmer structures and expresses her/his algorithm. Early proposals of Exascale programming tools are based on the adaptation of traditional parallel programming languages and hybrid solutions. This incremental approach is too conservative, often resulting in very complex code. This paper describes the design features, the programming constructs, and the runtime mechanisms of the Data Centric programming model for Exascale systems (DCEx). DCEx is based on structuring applications into data-parallel blocks. Blocks are units of shared-and distributed-memory parallel computation, communication, and migration in the memory/storage hierarchy. Blocks and their message queues are mapped onto processes and placed in memory/storage by the DCEx runtime. Those data-parallel blocks are orchestrated by using distributed parallel patterns that simplify the development cost. DCEx aims to reach the convergence of traditional HPC programming models, mainly based on MPI, with the emerging technologies based on the data intensive paradigms. To demonstrate the potential of DCEx, we carried out an experimental evaluation developing a real-world diffusion-weighted magnetic resonance imaging data processing application in a neuroimaging research context. Francisco Javier García Blas, Javier Fernández 0001, Jesús Carretero 0001, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Alberto Fernández-Pena, Daniel Martín de Blas |
SBAC-PAD | 3 |
| 2022 | Improving Performance and Capacity Utilization in Cloud Storage for Content Delivery and Sharing ServicesabstractContent delivery and sharing (CDS) is a popular and cost effective cloud-based service for organizations to deliver/share contents to/with end-users, partners and insider users. This type of service improves the data availability and I/O performance by producing and distributing replicas of shared contents. However, such a technique increases overhead on the storage/network resources. This article introduces a threefold methodology to improve the trade-off between I/O performance and capacity utilization of cloud storage for CDS services. This methodology includes: i) Definition of a classification model for identifying types of users and contents by analyzing their consumption/ demand and sharing patterns, ii) Usage of the classification model for defining content availability and load balancing schemes, and iii) Integration of a dynamic availability scheme into a cloud-based CDS system. Our method was implemented on both a simulator and a real-world CDS service, supporting information sharing operations performed in a cloud storage. An experimental evaluation, conducted in a private cloud through simulation and emulation of workloads, showed the feasibility of this methodology in terms of storage capacity utilization, whereas the real-world implementation revealed the efficiency of applying a classification model to information sharing patterns in terms of I/O performance. Víctor Jesús Sosa Sosa, Alfredo Barron, José Luis González 0002, Jesús Carretero 0001, Ivan López-Arévalo |
IEEE Trans. Cloud Comput. | 4 |
| 2021 | A Federated Content Distribution System to Build Health Data Synchronization ServicesabstractIn organizational environments, such as in hospitals, data have to be processed, preserved, and shared with other organizations in a cost-efficient manner. Moreover, organizations have to accomplish different mandatory non-functional requirements imposed by the laws, protocols, and norms of each country. In this context, this paper presents a Federated Content Distribution System to build infrastructure-agnostic health data synchronization services. In this federation, each hospital manages local and federated services based on a pub/sub model. The local services manage users and contents (i.e., medical imagery) inside the hospital, whereas federated services allow the cooperation of different hospitals sharing resources and data. Data preparation schemes were implemented to add non-functional requirements to data. Moreover, data published in the content distribution system are automatically synchronized to all users subscribed to the catalog where the content was published. Diana Carrizales-Espinoza, Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001 |
PDP | 4 |
| 2021 | LIMITLESS - LIght-weight MonItoring Tool for LargE Scale SystemsabstractThis work presents LIMITLESS, a HPC framework that provides new strategies for monitoring clusters. LIMITLESS is a scalable light-weight monitor that is integrated with other HPC runtimes in order to obtain an holistic view of the system that combines both platform and application monitoring. This paper presents a description of the novel components of the architecture, including new approaches for reaching a higher scalability based on a combination of in-transit processing and performance prediction. This work also includes a practical evaluation on simulated and real platforms, that shows significant monitoring scalability, retrieving data capacity and reduced overheads. Alberto Cascajo, David E. Singh, Jesús Carretero 0001 |
PDP | 3 |
| 2020 | Predictive data analysis techniques applied to dropping out of university studiesabstractStudent dropout is a major problem in university studies all around the world. To alleviate this problem, it is important to detect as soon as possible student attrition before he or she becomes a deserter. A student may be considered a deserter when she/he has not completed her academic credits or leave the studies. In this paper we present a study made at a higher education institution, by analyzing the records of 530 higher education students from 52 different careers with application date 2015 to 2018, considering factors such as academic monitoring, financial situation, personal and social information. These are some issues or mix of problems that could affect dropout rates. Analyze student behavior by implementing predictive analytics techniques reduce the gaps between professional demands and applicants' competencies. We applied predictive analytical techniques to identify the relationship of factors characterizing students who leave the university. As a result, we have elaborated a conceptual model to predict the risk of defection and applied machine learning techniques to generate preventive and corrective alerts as a student permanence strategy. This study shows that information is important, but the application of machine learning in the student's prior knowledge and its relationship to a dynamic and pre-established profile of the deserter student is essential to generate early strategies that manage to reduce the gaps between professional demands and applicants' competencies. In addition, a data model has been created to give solution to the issue get generated preventive and corrective alerts. Cindy Espinoza Aguirre, Jesús Carretero 0001 |
CLEI | 2 |
| 2020 | Exposing data locality in HPC-based systems by using the HDFS backendabstractNowadays, there are two main approaches for dealing with data-intensive applications: parallel file systems in classical High-Performance Computing (HPC) centers and Big Data like parallel file system for ensuring the data centric vision. Furthermore, there is a growing overlap between HPC and Big Data applications, given that Big Data paradigm is a growing consumer of HPC resources. HDFS is one of the most important file systems for data intensive applications while, from the parallel file systems point of view, MPI-IO is the most used interface for parallel I/O. In this paper, we propose a novel solution for taking advantage of HDFS through MPI-based parallel applications. To demonstrate its feasibility, we have included our approach in MIMIR, a MapReduce framework for MPI-based applications. We have optimized MIMIR framework by providing data locality features provided by our approach. The experimental evaluation demonstrates that our solution offers around 25% performance for map phase compared with the MIMIR baseline solution. José Rivadeneira, Félix García Carballeira, Jesús Carretero 0001, Francisco Javier García Blas |
HiPC | 3 |
| 2020 | Mapping and scheduling HPC applications for optimizing I/OabstractIn HPC platforms, concurrent applications are sharing the same file system. This can lead to conflicts, especially as applications are more and more data intensive. I/O contention can represent a performance bottleneck. The access to bandwidth can be split in two complementary yet distinct problems. The mapping problem and the scheduling problem. The mapping problem consists in selecting the set of applications that are in competition for the I/O resource. The scheduling problem consists then, given I/O requests on the same resource, in determining the order to these accesses to minimize the I/O time. In this work we propose to couple a novel bandwidth-aware mapping algorithm to I/O list-scheduling policies to develop a cross-layer optimization solution. Jesús Carretero 0001, Emmanuel Jeannot, Guillaume Pallez, David E. Singh, Nicolas Vidal 0001 |
ICS | 1 |
| 2020 | Applying big data paradigms to a large scale scientific workflow: Lessons learned and future directions
Silvina Caíno-Lores, Andrei Lapin, Jesús Carretero 0001, Peter G. Kropf |
Future Gener. Comput. Syst. | 3 |
| 2020 | New Parallel and Distributed Tools and Algorithms for Life Sciences
Jesús Carretero 0001, Dagmar Krefting |
Future Gener. Comput. Syst. | 1 |
| 2020 | Towards enhanced MRI by using a multiple back end programming framework
Francisco Javier García Blas, David del Rio Astorga, Jesús Carretero 0001, José Daniel García |
Future Gener. Comput. Syst. | 3 |
| 2020 | A gearbox model for processing large volumes of data by using pipeline systems encapsulated into virtual containers
Miguel Santiago-Duran, José Luis González 0002, André Brinkmann, Hugo G. Reyes-Anastacio, Jesús Carretero 0001, Raffaele Montella, Gregorio Toscano Pulido |
Future Gener. Comput. Syst. | 5 |
| 2020 | Accelerated iterative image reconstruction for cone-beam computed tomography through Big Data frameworks
Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Manuel Desco, Mónica Abella |
Future Gener. Comput. Syst. | 3 |
| 2020 | Kulla, a container-centric construction model for building infrastructure-agnostic distributed and parallel applications
Hugo G. Reyes-Anastacio, José Luis González 0002, Víctor Jesús Sosa Sosa, Jesús Carretero 0001, Francisco Javier García Blas |
J. Syst. Softw. | 4 |
| 2020 | CloudBench: an integrated evaluation of VM placement algorithms in clouds
Mario A. Gomez-Rodriguez, Víctor Jesús Sosa Sosa, Jesús Carretero 0001, José Luis González 0002 |
J. Supercomput. | 3 |
| 2019 | Exploiting Stream Parallelism of MRI Reconstruction Using GrPPI over Multiple Back-EndsabstractIn recent years, on-line processing of data streams has been established as a major computing paradigm. This is due mainly to two reasons: first, more and more data are generated in near real-time that need to be processed; the second reason is given by the need of efficient parallel applications. However, the above-mentioned areas expose a tough challenge over traditional data-analysis techniques, which have been forced to evolve to a stream perspective. In this work we present an comparative study of a stream-aware multi-staged application, which has been implemented using GrPPI, a generic and reusable parallel pattern interface for C++ applications. We demonstrate the benefits of using this interface in terms of programability, performance, and scalability. Francisco Javier García Blas, David del Rio Astorga, Javier Daniel Garcia, Jesús Carretero 0001 |
CCGRID | 4 |
| 2019 | A policy-based containerized filter for secure information sharing in organizational environments
José Luis González 0002, Oscar Telles-Hurtado, Ivan López-Arévalo, Miguel Morales-Sandoval, Víctor Jesús Sosa Sosa, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 6 |
| 2019 | Trends on heterogeneous and innovative hardware and software systems
Alba Cristina Magalhaes Alves de Melo, Jesús Carretero 0001, Per Stenström, Sanjay Ranka, Eduard Ayguadé |
J. Parallel Distributed Comput. | 2 |
| 2019 | Combining malleability and I/O control mechanisms to enhance the execution of multiple applications
David E. Singh, Jesús Carretero 0001 |
J. Syst. Softw. | 2 |
| 2018 | Spark-DIY: A Framework for Interoperable Spark Operations with High Performance Block-Based Data ModelsabstractToday's scientific applications are increasingly relying on a variety of data sources, storage facilities, and computing infrastructures, and there is a growing demand for data analysis and visualization for these applications. In this context, exploiting Big Data frameworks for scientific computing is an opportunity to incorporate high-level libraries, platforms, and algorithms for machine learning, graph processing, and streaming; inherit their data awareness and fault-tolerance; and increase productivity. Nevertheless, limitations exist when Big Data platforms are integrated with an HPC environment, namely poor scalability, severe memory overhead, and huge development effort. This paper focuses on a popular Big Data framework -Apache Spark- and proposes an architecture to support the integration of highly scalable MPI block-based data models and communication patterns with a map-reduce-based programming model. The resulting platform preserves the data abstraction and programming interface of Spark, without conducting any changes in the framework, but allows the user to delegate operations to the MPI layer. The evaluation of our prototype shows that our approach integrates Spark and MPI efficiently at scale, so end users can take advantage of the productivity facilitated by the rich ecosystem of high-level Big Data tools and libraries based on Spark, without compromising efficiency and scalability. Silvina Caíno-Lores, Jesús Carretero 0001, Bogdan Nicolae, Orcun Yildiz, Tom Peterka |
BDCAT | 2 |
| 2018 | GPU-accelerated iterative reconstruction for limited-data tomography in CBCT systemsabstractBACKGROUND: Standard cone-beam computed tomography (CBCT) involves the acquisition of at least 360 projections rotating through 360 degrees. Nevertheless, there are cases in which only a few projections can be taken in a limited angular span, such as during surgery, where rotation of the source-detector pair is limited to less than 180 degrees. Reconstruction of limited data with the conventional method proposed by Feldkamp, Davis and Kress (FDK) results in severe artifacts. Iterative methods may compensate for the lack of data by including additional prior information, although they imply a high computational burden and memory consumption. RESULTS: pixels) using partitioning strategies in forward- and back-projection operations. We evaluated the algorithm on small-animal data for different scenarios with different numbers of projections, angular span, and projection size. Reconstruction time varied linearly with the number of projections and quadratically with projection size but remained almost unchanged with angular span. Forward- and back-projection operations represent 60% of the total computational burden. CONCLUSION: Efficient implementation using parallel processing and large-memory management strategies together with GPU kernels enables the use of advanced reconstruction approaches which are needed in limited-data scenarios. Our GPU implementation showed a significant time reduction (up to 48 ×) compared to a CPU-only implementation, resulting in a total reconstruction time from several hours to few minutes. Claudia de Molina, Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Manuel Desco, Mónica Abella |
BMC Bioinform. | 4 |
| 2018 | New directions in mobile, hybrid, and heterogeneous clouds for cyberinfrastructures
Jesús Carretero 0001, Francisco Javier García Blas, Gabriel Antoniu, Dana Petcu |
Future Gener. Comput. Syst. | 1 |
| 2018 | Sacbe: A building block approach for constructing efficient and flexible end-to-end cloud storage
José Luis González 0002, Víctor Jesús Sosa Sosa, Arturo Díaz-Pérez, Jesús Carretero 0001, Jedidiah Yanez-Sierra |
J. Syst. Softw. | 4 |
| 2017 | Data-Aware Support for Hybrid HPC and Big Data ApplicationsabstractNowadays there is a raising interest in bridging the gap between Big Data application models and data-intensive HPC. This work explores the effects that Big Data-inspired paradigms could have in current scientific applications through the evaluation of a real-world application from the hydrology domain. This evaluation led to experience that portrayed the key aspects of the HPC and Big Data paradigms that made them successful in their respective worlds. With this information, we established a research roadmap to build a platform suitable for HPC hybrid applications, with a focus on efficient data management and fault-tolerance. Silvina Caíno-Lores, Florin Isaila, Jesús Carretero 0001 |
CCGrid | 3 |
| 2017 | Medical Imaging Processing on a Big Data platform using Python: Experiences with Heterogeneous and Homogeneous ArchitecturesabstractThe apparition of new paradigms, programming models, and languages that offer better programmability and better performance turns the implementation of current scientific applications into a less time-consuming task than years ago. One significant example of this trend is the MapReduce programming model and its implementation using Apache Spark. Nowadays, this programming model is mainly used for data analysis and machine learning applications, although it has been expanded to its usage in the HPC community. On the side of programming languages, Python has positioned itself as an alternative to other scientific programming languages, such as Matlab or Julia. In this work we explore the capabilities of Python and Apache Spark as partners in the implementation of the backprojection operator of a CT reconstruction application. We present two interesting approaches with two different types of architectures: a heterogeneous architecture including NVidia GPUs and a full performance CPU mode with the compatibility with C/C++ native source code. We experimentally demonstrate that current CPU-based implementations scale with the number of computational units. Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Mónica Abella, Manuel Desco |
CCGrid | 3 |
| 2017 | Algorithms and applications towards the convergence of high-end data-intensive and computing systemsabstractWith the increasing availability of data generated by scientific instruments and simulations, today, solving many of our most important scientific and engineering problems requires high-end computing systems (HECS)1 that may be able to process and storage a huge amount of data.2 With this landscape, many synergies between extreme-scale computing, simulations, and data intensive applications might arise.(3, 4) However, the high-performance computing and data analysis platforms, paradigms, and tools have evolved in many cases in different fields, having their own specific methodologies, tools, and techniques. We need to evolve systems and paradigms to create High-End Data-Intensive Computing Systems (HEDICS) to create high-end resources that must be powerful enough in a broad sense (computation, storage, I/O capacity, communications, etc), but at the same time have to provide utilities from the Big Data computing (BDC) space to satisfy the data management and analytics needs of near future applications. Future HECS platforms will be likely characterized by a three to four orders of magnitude, increasing in concurrency, a substantially larger storage capacity, and a deepening of the storage hierarchy. Moreover, the advent of the Big Data challenges5 has generated new initiatives closely related to ultrascale computing systems in large scale distributed systems. The current uncoordinated development model of independently applying optimizations at each layer of the system software I/O software stack will not scale to the required levels of distribution, concurrency, storage hierarchy, and capacity.6 Thus, we need reusable, modular, and scalable frameworks for designing high-end reconfigurable computers, including novel data processing building block and innovative programming models. In those aspects, many new topics are open to research: parallel and distributed algorithms for HEDICS; algorithms for aggressive management of information and knowledge from massive data sources; resource management and scheduling in high-end data and computing systems; tools and environments for parallel/distributed high-end software development; new programming models, as well as machine and application abstractions; resilience issues in HEDICS; adaptive software; architectures, networks, and systems suited for extreme-scale and Big Data; massive distributed and parallel data analytics and feature extraction; new I/O and storage systems valid for HEDICS; and novel and redesigned high-end scientific and engineering computing. This special issue is intended to provide an overview of some key topics and state-of-the-art of recent advances in subjects relevant to High-End Data-Intensive Computing Systems. The general objectives are to address, explore, and exchange information on the challenges and current state-of-the-art in HEDICS, new programming models, run-times, and data facilities design and performance, and their application in various science and engineering domains. This special issue includes research papers addressing the state-of-the-art in high-end data-intensive computing systems. A set of carefully selected works was invited based on the original presentations at the 16th International Conference on Algorithms and Architectures for Parallel Processing (ICA3PP 2016),7 which was held in Granada, Spain, December 2016 and the Third International Workshop of Sustainable Ultrascale Network (NESUS 2016),8 held in Sofia, Bulgaria, October 2016. The extended works have been thoroughly reviewed by an international technical reviewing committee, and only nine papers covering a wide range of relevant challenges in HEDICS were selected for this special issue. The manuscripts present research works showing the convergence of High-End Data and Computing Systems, including new frameworks and platforms, system software enhancements, algorithm design and optimization, programming paradigms and techniques, data processing support in high-end computing systems, and run-time support for HEDICS and performance simulations, measurement, and evaluations. The set of accepted papers can be organized under the following key subjects and subsections and are briefly described in the remaining parts of this section. Current parallel and distributed programming frameworks aid developers to a great extent in implementing applications that exploit homogeneous resources. Nevertheless, it is generally accepted that the ability to develop large-scale distributed applications has lagged seriously behind other developments in cyber-infrastructure.9 Thus, developers strongly require additional expertise to properly port and tune their applications to operate efficiently on specific parallel and distributed platforms, which is not straightforward and demands considerable efforts and specific knowledge. One important cause is the lack of high-level parallel pattern abstractions in the existing frameworks. Dolz et al,10 in their paper A Generic Parallel Pattern Interface for Stream and Data Processing, propose GRPPI, a generic and reusable parallel pattern interface for both stream processing and data-intensive C++ applications available for high-end nodes. GRPPI accommodates a layer between developers and existing parallel programming back-ends targeting multi-core processors, such as C++ threads, OpenMP and Intel TBB, and accelerators back-end like CUDA Thrust. Furthermore, thanks to its high-level C++ API and pattern composability features, GRPPI enables users to easily expose parallelism via stand-alone patterns or pattern compositions matching in sequential applications. The authors evaluate this interface using an image processing use case and demonstrate its benefits from the usability, flexibility, and performance points of views. Furthermore, they analyse the impact of using stream and data pattern compositions on CPUs, GPUs, and heterogeneous configurations. To scale to the next level, as high-end data intensive computing systems become more widespread for scientific applications, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data-intensive applications in high-performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks generally arranged in a directed acyclic graph style, which communicate through storage abstractions. Regarding the paper A Data-aware Scheduling Strategy for Workflow Execution in Clouds, Marozzo et al11 present the integration between DMCF and Hercules solutions by using a data-aware scheduling strategy for exploiting data locality in data-intensive workflows. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation, while Hercules is an in-memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. The experimental results demonstrate the performance improvements achieved using the proposed data-aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with the new proposed scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. In spite of former solutions, network traffic is always a major problem in HEDICS due to data movements. In their paper A scalable synthetic traffic model of Graph500 for computer networks analysis, Fuentes et al12 provide a simulation tool for network architects that need to evaluate the suitability of their interconnect for Big Data applications. Their development is a low computation- and memory-demanding synthetic traffic model that emulates the behaviour of the Graph500 communications and is publicly available in an open-source network simulator. The characterization of network traffic is inferred from a profile of several executions of the benchmark with different input parameters, and the equations in the model have been validated against an execution of benchmarks with a different set of parameters to measure also the impact of the node computation capabilities and network characteristics in the execution time of the model. To cope with huge jobs, some organizations use volunteer computing to get computing resources to scientific projects, so that organizations can be able to attain large computing power from volunteer clients instead of making a high investment in infrastructure. However, there are projects, like the ATLAS@Home project,13 in which the number of running jobs has reached a plateau, due to a high load on data servers and networks caused by file transfers. Alonso et al,14 in the paper A New Volunteer Computing Model for Data-Intensive Applications, provide an alternative, named ComBoS, to improve the performance of volunteer computing projects that have reached their limit due to the I/O bottleneck in data servers by having a percentage of the volunteer clients running as data servers, called data volunteers, to reduce the load on data servers. This solution also improves data locality, leveraging the network latencies of closer machines, as shown by the performance increase provides by their solution, applied to three different BOINC projects. Two current trends in Big Data processing have made the usage of GPGPUs very popular in HEDICS: information discovery and deep-learning techniques15 and collective video games.16 In both cases, there is an increasing trend to discharge client nodes by sending bulk computing to heterogeneous high-end computing nodes for data processing. Data compression is an important area in many data management applications, like training of deep learning, where data must be decompressed many times. Nakano et al,17 in their paper Adaptive Loss-Less Data Compression Method Optimized for GPU Decompression, present a novel lossless data compression method, called Adaptive LossLess (ALL) data compression, designed with the objective of performing decompression very efficiently on the GPU. Evaluations of the ALL data compression method against published lossless data compression methods implemented in GPU show improvements between 1.22 and 23.5 times running on the same GPU. Due to the massive extension of many mobile applications, such as sensors and smart phones, it is crucial for HEDICS to offload applications to high-end nodes so that low-power devices can be used as clients. One possible approach to deal with this problem is the solution proposed in the paper Accelerating Linux and Android applications on low-power devices through remote GPGPU offloading by Montella et al.18 They describe the architecture and integration of RAPID, a complete framework suite for computation offloading to help low-powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices, providing lightweight secure data transmission of the offloading operations. The proposed framework is highly modular and exposes a rich Application Programming Interface (API) to developers, making it highly versatile while hiding the complexity of the underlying networking layer. The evaluation results show that Java/Android GPGPU code offloading is possible, through a BioSurveillance application, a commercial real-time face recognition application. High-end networked scientific and engineering applications requires usually HPC for numerical computing and large storage capabilities at end nodes. As the problem grows, the scientific community, in its never-ending road of larger and more efficient computational resources, is in need of more efficient implementations that can adapt better to the current parallel platforms and in need of more new solutions for memory problems that are now memory bound. The memory problem is addressed by Valero19 in the paper Reducing Memory Requirements for Large Size LBM Simulations on GPUs, where he proposes some initiatives to minimize the memory requirements of the Lattice- Boltzmann Method for its usage on GPGPUs to run large scale simulations. The proposed approach allows the author to execute bigger simulations on the same platform without additional memory transfers, those achieving a high performance. In particular, the paper presents two new implementations, LBM-Ghost and LBM-Swap, which are deeply analysed, presenting the pros and cons of each of them. The need of parallelization at high-end nodes is addressed in the paper Parallel solvers for fractional power diffusion problems by Starikoviius et al. 20 The authors construct and investigate parallel solvers for problems described by fractional powers of elliptic operators, like fractional diffusion. Three state-of-the-art approaches are used to transform the non-local fractional-order differential problem into local partial differential equation problems formulated in a space of higher dimension. Scalability of the developed parallel algorithms is investigated, and their parallel performance is compared in the paper. Finally, the problem of accuracy and efficiency for statistic distributions is addressed by Monni et al21 in the paper Fitting Long-Tailed Distribution to Empirical Data. The authors discuss about the limits of the analysis of empirical fat-tailed distributions, which can describe a variety of evolving systems, both natural and man-made. An algorithm to fit fat-tailed distributions is presented and tested against samplings of the power law, the Yule, the log-normal, and Weibull distributions. The algorithm is general and can be applied to any numerical dataset. Thus, the authors compute the parameters defining the shape of each distribution and test the results against simulations. Their method with another state-of- the-art technique to estimate the parameters of empirical distributions. The accuracy of the estimations is discussed, and they conclude that their method based on a weighted iterated χ2 test performs better than the other. Power laws can fit a variety of distributions coming from real data, so a systematic approach to the measurement of the accuracy of fitting algorithms is essential. Articles presented in this special issue provide recent advances in some fields related to high-end data-intensive computing systems. They were selected by invitation of best ranked from two conferences and peer reviewed by journal selection. Acceptance rate for the special issue was below 50% of the invited papers. We hope that the ideas presented in this special issue can contribute to this strategically important, exciting, and fast growing research area and will be of interest for readers of the journal. As guest editors of this special issue, we would like to express our gratitude to all of the authors who submitted their papers to this special issue, and to the Reviewers that helped us with their hard work and the feedback provided to the authors. We also wish to express our gratitude to the Editor-in-Chief Geoffrey C. Fox for the opportunity to edit this special issue and his assistance during the special issue preparation. We acknowledge the following Reviewing Committee members: Pawe Czarnul (Poland), Guilherme Dinis (Sweden), Ece Guran Schmidt (Turkey), Massimiliano Ferrara (Italy), Shih-Hao Hung (Taiwan), Florin Isaila (Spain), Dingde Jiang (China), Amin Khan (Portugal), Marcin Kostur (Poland), Kenli Li (China), Francesco Longo (italy), Francesc Lordan (Spain), Najme Mansouri (Iran), Panagiotis Michailidis (Greece), Eike Mueller (UK), Tomas Potuzak (Cezch Republic), Philipp Reinecke (Germany), Francisco Rodrigo (Spain), Gopal Shyam (India), Shengen Yan (China), Wenwu Tang (USA), and Peng Zhang (USA). Jesús Carretero 0001, Francisco Javier García Blas, Koji Nakano, Peter Mueller |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | A data-aware scheduling strategy for workflow execution in cloudsabstractSummary As data intensive scientific computing systems become more widespread, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data‐intensive applications in high‐performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks arranged in a directed acyclic graph style, which communicate through storage abstractions. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation. Hercules is an in‐memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. This work improves the integration between DMCF and Hercules by using a data‐aware scheduling strategy for exploiting data locality in data‐intensive workflows. This paper presents experimental results demonstrating the performance improvements achieved using the proposed data‐aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with our scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. Fabrizio Marozzo, Francisco Rodrigo Duro, Francisco Javier García Blas, Jesús Carretero 0001, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Boosting analyses in the life sciences via clusters, grids and clouds
Sandra Gesing, Jesús Carretero 0001, Francisco Javier García Blas, Johan Montagnat |
Future Gener. Comput. Syst. | 2 |
| 2017 | Experimental evaluation of a flexible I/O architecture for accelerating workflow engines in ultrascale environments
Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001, Justin M. Wozniak, Robert B. Ross |
Parallel Comput. | 4 |
| 2017 | Virtual Environments and Advanced Interfaces
Daphne Economou, Markos Mentzelopoulos, Nektarios Georgalas, Jesús Carretero 0001, Francisco Javier García Blas |
Pers. Ubiquitous Comput. | 4 |
| 2017 | Model-based energy-aware data movement optimization in the storage I/O stack
Pablo Llopis, Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2016 | Flexible Data-Aware Scheduling for Workflows over an In-memory Object StoreabstractThis paper explores novel techniques for improving the performance of many-task workflows based on the Swift scripting language. We propose novel programmer options for automated distributed data placement and task scheduling. These options trigger a data placement mechanism used for distributing intermediate workflow data over the servers of Hercules, a distributed key-value store that can be used to cache file system data. We demonstrate that these new mechanisms can significantly improve the aggregated throughput of many-task workflows with up to 86x, reduce the contention on the shared file system, exploit the data locality, and trade off locality and load balance. Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Justin M. Wozniak, Jesús Carretero 0001, Robert B. Ross |
CCGrid | 5 |
| 2016 | CLARISSE: A Middleware for Data-Staging Coordination and Control on Large-Scale HPC PlatformsabstractOn current large-scale HPC platforms the data path from compute nodes to final storage passes through several networks interconnecting a distributed hierarchy of nodes serving as compute nodes, I/O nodes, and file system servers. Although applications compete for resources at various system levels, the current system software offers no mechanisms for globally coordinating the data flow for attaining optimal resource usage and for reacting to overload or interference. In this paper we describe CLARISSE, a middleware designed to enhance data-staging coordination and control in the HPC software storage I/O stack. CLARISSE exposes the parallel data flows to a higher-level hierarchy of controllers, thereby opening up the possibility of developing novel cross-layer optimizations, based on the run-time information. To the best of our knowledge, CLARISSE is the first middleware that decouples the policy, control, and data layers of the software I/O stack in order to simplify the task of globally coordinating the data staging on large-scale HPC platforms. To demonstrate how CLARISSE can be used for performance enhancement, we present two case studies: an elastic load-aware collective I/O and a cross-application parallel I/O scheduling policy. The evaluation illustrates how coordination can bring a significant performance benefit with low overheads by adapting to load conditions and interference. Florin Isaila, Jesús Carretero 0001, Robert B. Ross |
CCGrid | 2 |
| 2016 | QuizMonitor: A learning platform that leverages student monitoringabstractThis work presents the design, implementation, and evaluation of a learning platform that addresses two main objectives: first it provides and on-line quiz tool for students which can be used as a complementary learning approach to the classroom courses. Secondly, this tool performs a detailed analysis of learners use, considering not only the number of mistakes students have made but also the student temporal use distribution and opinion (obtained by a survey) about the difficulty of the learning contents. All this information is processed and used to provide feedback to the teachers identifying the most difficult contents of the area of study and the students with a low learning performance. We have developed and evaluated this tool in two university degree subjects. The obtained results are promising, showing that students can improve their final grades through its use and that teachers can identify the student learning problems, in order to assess where students may require more support or challenge. In addition, we analyze the impact of different metrics obtained by this tool (academic performance, intensity of study and student proactivity) in the final subject grades. Carlos Gomez, David E. Singh, Jesús Carretero 0001 |
EDUCON | 3 |
| 2016 | Porting Matlab Applications to High-Performance C++ Codes: CPU/GPU-Accelerated Spherical Deconvolution of Diffusion MRI Data
Francisco Javier García Blas, Manuel F. Dolz, José Daniel García, Jesús Carretero 0001, Alessandro Daducci, Yasser Alemán-Gómez, Erick Jorge Canales-Rodríguez |
ICA3PP | 4 |
| 2016 | Methodological Approach to Data-Centric Cloudification of Scientific Iterative Workflows
Silvina Caíno-Lores, Andrei Lapin, Peter G. Kropf, Jesús Carretero 0001 |
ICA3PP | 4 |
| 2016 | Improving the Energy Efficiency of MPI Applications by Means of MalleabilityabstractThis work presents two novel techniques for increasing the energy efficiency of parallel applications by means of malleability. These techniques are implemented as an extension of Flex-MPI, a library implemented on top of MPI, which provides performance-aware dynamic reconfiguration for MPI-based applications. During the application execution, Flex-MPI performs energy and performance monitoring by means of energy and performance counters. It leverages this information in order to adapt the program performance using two energy policies: energy minimization and performance-per-watt maximization. The evaluation results show that these new energy-aware capabilities permit MPI applications to be executed in an optimized way both in terms of performance and energy efficiency. Manuel Rodriguez-Gonzalo, David E. Singh, Francisco Javier García Blas, Jesús Carretero 0001 |
PDP | 4 |
| 2016 | HeteroPar 2014, APCIE 2014, and TASUS 2014 Special IssueabstractThese workshops were organized by members of the Nesus Cost Action IC 1305: Network for Sustainable Ultrascale Computing, which is a follow-up of COST Actions IC0804 and IC0805 1. The goal of the NESUS Action is to establish an open European research network targeting sustainable solutions for ultrascale computing aiming at cross fertilization among HPC, large-scale distributed systems, and big data management. This network aims at contributing to glue disparate researchers working across different areas and provide a meeting ground for researchers in these separate areas to exchange ideas, to identify synergies, and to pursue common activities in research topics such as sustainable software solutions (applications and system software stack), data management, energy efficiency, and resilience. The selected papers cover very important scientific issues encountered nowadays such as the following: CPU/GPU execution, system-on-chip programming, parallel algorithms taking into account various constraints (energy, communication, etc.), programming models, and so on. We really hope that the reader will enjoy this high-quality issue, and we are sure that she/he will find it highly relevant to the state-of-the-art of today's heterogeneous and parallel computing. Jesús Carretero 0001, Raimondas Ciegis, Emmanuel Jeannot, Laurent Lefèvre, Gudula Rünger, Domenico Talia, Julius Zilinskas |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Introduction to sustainable ultrascale computing systems and applications
Jesús Carretero 0001, Francisco Javier García Blas, Raimondas Ciegis |
J. Supercomput. | 1 |
| 2016 | RS-Pooling: an adaptive data distribution strategy for fault-tolerant and large-scale storage systems
Moisés Quezada Naquid, Ricardo Marcelín-Jiménez, José Luis González 0002, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2015 | Simulation Platform for X-Ray Computed Tomography Based on Low-Power Systems
Estefania Serrano, Francisco Javier García Blas, Alberto Verza, Jesús Carretero 0001 |
ICA3PP (4) | 4 |
| 2015 | A Multi-Objective Simulator for Optimal Power Dimensioning on Electric Railways using Cloud ComputingabstractPower dimensioning and energy saving have been traditionally two main issues regarding the deployment of
electric grids. Electric railways are also concerned about these issues, and simulators have been traditionally
used to test such infrastructure deployments. The main goal of this paper is to present the Railway electric
Power Consumption Simulator, a simulation model and tool for the railway energy provisioning problem. This
simulator aims to propose electric railway infrastructure deployments, optimizing the quality of the electric
flow supplied to train, as well as saving as much energy as possible. The paper describes the simulator
structure, as well as the ontology used to translate railway infrastructure elements into an electric circuit.
Because these two objectives are conflicting, a multi-objective optimization problem is formulated and solved.
Finally, a standard railway scenario is used to illustrate the capabilities of the tool, trying to find the best
electric substation placements in order to optimize such objectives. The evaluation shows how the tool can
handle hundreds of simulated scenarios using Cloud Computing techniques. Jesús Carretero 0001, Silvina Caíno-Lores, Félix García Carballeira, Alberto García Fernández |
SIMULTECH | 1 |
| 2015 | Multi-dimensional recursive routing with guaranteed delivery in Wireless Sensor Networks
Stefano Chessa, Soledad Escolar, Susanna Pelagatti, Jesús Carretero 0001 |
Comput. Commun. | 4 |
| 2015 | A comparative study of an X-ray tomography reconstruction algorithm in accelerated and cloud computing systemsabstractSummary With the increase of resolution in medical image scanners and the need of faster reconstruction methods, new ways of exploiting the inherent parallelism of reconstruction algorithms have arisen. In this paper, we present Mangoose++, an application to perform X‐ray computed tomography that supports multiple grades of parallelism. This parallelism is tackled with two different approaches: the usage of parallel nodes with multicore CPUs in a cloud environment and the usage of high‐performance computing (HPC)‐based parallel architectures such as general‐purpose computing on graphics processing unit (GPGPU) or Intel Xeon Phi. In this paper, we show the design and implementation of the application in three types of platforms related to the previous mentioned approaches, comparing and analyzing the performance, resource utilization, and scalability of each platform. Accelerators offer high performance for data sizes that fit inside the accelerator memory. This is the main advantage of Intel Xeon Phi that, in this work, obtains similar performance results than compute unified device architecture (CUDA)‐based GPGPU versions, comparing with cards with less memory capacity. In our evaluation experiments, we additionally analyze and discuss the costs and efficiency of Mangoose++ over Amazon Compute Cloud platform, demonstrating that lower times can be achieved in a reasonable price compared with owned HPC‐based hardware. A comparison between distinct hardware configurations is provided for emphasizing on the advantages and disadvantages of each one. Copyright © 2015 John Wiley & Sons, Ltd. Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Towards efficient large scale epidemiological simulations in EpiGraphabstractThe work we present in this paper focuses on understanding the propagation of flu-like infectious outbreaks between geographically distant regions due to the movement of people outside their base location. Our approach incorporates geographic location and a transportation model into our existing region-based, closed-world EpiGraph simulator to model a more realistic movement of the virus between different geographic areas. This paper describes the MPI-based implementation of this simulator, including several optimization techniques such as a novel approach for mapping processes onto available processing elements based on the temporal distribution of process loads. We present an extensive evaluation of EpiGraph in terms of its ability to simulate large-scale scenarios, as well as from a performance perspective. Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
Parallel Comput. | 4 |
| 2015 | Enhancing the performance of malleable MPI applications by using performance-aware dynamic reconfiguration
Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
Parallel Comput. | 4 |
| 2014 | Evaluation of the Feasibility of Making Large-Scale X-Ray Tomography Reconstructions on CloudsabstractThis work focuses on the evaluation of the suitability of Mangoose++, a medical image application for reconstruction of 3D volumes, by means of Cloud Computing. Due to the increasing resolution of panel detectors in computed tomography and the need of lower execution times, the use of parallel implementations for clusters and accelerators have been generalized. Anyhow, the renewal and maintenance of hardware is expensive which makes Cloud Computing a valuable alternative. In our evaluation, we analyze and discuss the costs and efficiency of the Mangosee++ application over Amazon EC2 platform, demonstrating that lower times can be achieved in a reasonable price compared with owned HPC-based hardware. We also provide a comparison between distinct hardware configurations so that we can infer the advantages and disadvantages of each one. Estefania Serrano, Guzmán Bermejo, Francisco Javier García Blas, Jesús Carretero 0001 |
CCGRID | 4 |
| 2014 | High-performance X-ray tomography reconstruction algorithm based on heterogeneous accelerated computing systemsabstractMany medical image processing applications need high processing speed to achieve almost real-time image reconstruction features. Due to that, massively parallel architectures based on accelerators have become very popular in the area, specially GPGPUs. In this paper we show Mangoose++, an application to perform X-Ray Computed Tomography (CT) from medical image based on a new implementation of the FDK algorithm. Mangoose++ have been designed and implemented to exploit the parallelism existing on several hardware accelerators platforms, as GPGPUs and Intel Xeon Phi accelerators. In this paper we show the design and implementation of the application in three types of platforms, multi-core CPU, GPGPU, and Intel Xeon Phi, and the evaluation made to test the performance, resource utilization, and scalability of each platform. Moreover, to avoid hardware dependencies, we have also implemented the application using the OpenACC runtime to check portability and the overhead incurred when using runtimes. The evaluation results show that our solution is faster than recent related works and that, in terms of computation, Intel Xeon Phi and the CUDA-based GPU versions obtain similar results as the problem size increases. Moreover, the evaluation shows that using OpenACC, we have enhanced programmability because there is a single version of the source code. But it also shows that using OpenACC heavily affects performance of Mangoose++, which is reduced in a 50% when compared with the many-core versions, even when it is not so drastical when compared to the CPU version. Estefania Serrano, Guzmán Bermejo, Francisco Javier García Blas, Jesús Carretero 0001 |
CLUSTER | 4 |
| 2014 | Routing with virtual coordinates in mobile sensor networksabstractThe realization of smart cities relies on the availability of large amount of data about occurring phenomena/events that can be guaranteed by very large deployments of Wireless Sensor Networks (WSN). This poses a great challenge to the scalability of current routing protocols for WSN due to the size and density of the network and the presence of mobile sensors.We tackle this problem by proposing a solution based on virtual coordinate systems combined with mechanisms that renew the virtual coordinates and suitable routing schemes. The simulation results show that this approach is actually suited to this context and that it guarantees high delivery rate and low path length. Stefano Chessa, Soledad Escolar, Susanna Pelagatti, Jesús Carretero 0001 |
ISCC | 4 |
| 2014 | A holistic approach to railway engineering design using a simulation framework
Jesús Carretero 0001, Carlos Gomez, Alberto García Fernández, Félix García Carballeira |
SIMULTECH | 1 |
| 2014 | Survey of Energy-Efficient and Power-Proportional Storage SystemsabstractIncreasingly, large-scale computing systems are consuming more power each passing year. As power consumption is on the rise, concern has been raised over the growing implications on power bills, carbon emissions and power supply limitations for data centers. Computing systems use hardware components that are not power proportional and machines comprising large systems tend to be underutilized, resulting in great wastage of energy. Motivated by this fact, researchers aim to improve the energy efficiency of these systems by increasing resource utilization and exploiting system characteristics in order to achieve power proportionality. However, many challenges burden the task of producing a power-aware, energy-efficient large-scale system that provides the same performance as today's systems. Particularly, storage systems are of great concern since storage consumes a large amount of power in the data center. This paper outlines the main problems and challenges of delivering power-efficient storage solutions, and proposes a taxonomy of power-aware techniques. This work aims to provide a detailed exposition of current trends in power-aware storage systems in a comprehensive and organized manner, outlining and comparing current solutions and challenges. Pablo Llopis, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001 |
Comput. J. | 4 |
| 2014 | Surfing the optimization space of a multiple-GPU parallel implementation of a X-ray tomography reconstruction algorithm
Francisco Javier García Blas, Mónica Abella, Florin Isaila, Jesús Carretero 0001, Manuel Desco |
J. Syst. Softw. | 4 |
| 2014 | The Internet of Things: connecting the world
Jesús Carretero 0001, José Daniel García |
Pers. Ubiquitous Comput. | 1 |
| 2014 | Energy management in solar cells powered wireless sensor networks for quality of service optimization
Soledad Escolar, Stefano Chessa, Jesús Carretero 0001 |
Pers. Ubiquitous Comput. | 3 |
| 2014 | CONDESA: A Framework for Controlling Data Distribution on Elastic Server ArchitecturesabstractApplications running in today's data centers show high workload variability. While seasonal patterns, trends and expected events may help building proactive resource allocation policies, this approach has to be complemented with adaptive strategies which should address unexpected events such as flash crowds and volume spikes. Additionally, the limitations of current I/O infrastructures in the face of dramatic increase of data generation require, the ability to build novel abstractions and models for robust decision making regarding data layout and data locality. In this work, we present CONDESA (CONtrolling Data distribution on Elastic Server Architectures), a framework for exploring adaptive data distribution strategies for elastic server architectures. To the best of our knowledge CONDESA is the first platform that permits to systematically study the interplay between five data related strategies: workload prediction, adaptive control of data distribution and server provisioning, adaptive data grouping, adaptive data placement, and adaptive system sizing. We demonstrate how CONDESA can be used for browsing the design space of adaptive data distribution policies. We show how prediction models can be compared in terms of overhead and accuracy. We evaluate the impact of change detection on prediction accuracy and how CONDESA can be used for choosing an adequate prediction horizon. We demonstrate how adaptive prediction can be used for sizing a server system. Finally, we show how prediction models, change detection strategies, and data placement policies can be combined and compared based on server utilization, load balance, data locality, over- and underprovisioning. Juan Manuel Tirado, Daniel Higuero, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2013 | FLEX-MPI: An MPI Extension for Supporting Dynamic Load Balancing on Heterogeneous Non-dedicated Systems
Gonzalo Martín 0001, Maria-Cristina V. Marinescu, David E. Singh, Jesús Carretero 0001 |
Euro-Par | 4 |
| 2013 | Energy management of networked, solar cells powered, wireless sensorsabstractSolar cells combined with power management algorithms enable the dynamic scheduling of Wireless Sensor Networks applications in a reference period, where the objective of the scheduling is to maximize the application quality level while conserving an energy level sufficient to constantly maintain the sensor operation. In this paper we consider networked, solar cells powered wireless sensors and we propose an algorithm aims to find a global, (sub)optimal scheduling that maximizes the overall quality of service in the sensors and keeps the system energy neutral, thus ensuring that the system works uninterruptedly. Soledad Escolar, Stefano Chessa, Jesús Carretero 0001 |
MSWiM | 3 |
| 2013 | Parallel implementation of a X-ray tomography reconstruction algorithm based on MPI and CUDAabstractMost small-animal X-ray computed tomography (CT) scanners are based on cone-beam geometry with a flat-panel detector orbiting in a circular trajectory. Image reconstruction in these systems is usually performed by approximate methods based on the algorithm proposed by Feldkamp, Davis and Kress (FDK). Currently there is a strong need to speedup the reconstruction of X-Ray CT data in order to extend its clinical applications. The evolution of the semiconductor detector panels has resulted in an increase of detector elements density, which produces a higher amount of data to process. This work focuses on future high-resolution studies (density up to 4096 pixeles), in which multiple level of parallelism will be needed in the reconstruction. In addition, this paper addresses the future challenges of processing high-resolution images in many-core and distributed architectures. In our evaluation section we demonstrate that our solution is 17% faster than recent related works. Francisco Javier García Blas, Florin Isaila, Mónica Abella, Jesús Carretero 0001, Ernesto Liria, Manuel Desco |
EuroMPI | 4 |
| 2013 | Improving MPI applications with a new MPI_Info and the use of the memoizationabstractThe MPI forum is actively working for a better MPI standard. The results are the new version 3 of the MPI standard, and the efforts for the incoming MPI 3.1/4.0. The technological changes provide many opportunities for improvements and new ideas. This paper introduces two main contributions in this direction: (1) how to improve the MPI_Info object implementation, and (2) a new way of using the former improved MPI_Info object as a storage solution. Alejandro Calderón 0001, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Daniel Higuero, Borja Bergua |
EuroMPI | 2 |
| 2013 | A hierarchical parallel storage system based on distributed memory for large scale systemsabstractThis paper presents the design and implementation of a storage system for high performance systems based on a multiple level I/O caching architecture. The solution relies on Memcached as a parallel storage system, preserving its powerful capacities such as transparency, quick deployment, and scalability. The designed parallel storage system targets to reduce the I/O latency in data-intensive high performance applications. The proposed solution consists of a user-level library and extended Memcached servers. The solution aims to be hierarchical by deploying Memcached-based I/O servers across all the infrastructure data path. Our experiments demonstrate that our solution is up to 40% faster than PVFS2. Francisco Rodrigo Duro, Francisco Javier García Blas, Jesús Carretero 0001 |
EuroMPI | 3 |
| 2013 | Parallel algorithm for simulating the spatial transmission of influenza in EpiGraphabstractThis paper introduces an approach to modeling and simulating the propagation of flu-like infectious diseases over large, widely spread urban areas connected by transportation networks. We incorporate geographic location and a transportation model into our region-based, closed-world EpiGraph simulator to realistically model the movement of the virus between different geographic regions. The resulting simulator can assist in understanding how outbreaks propagate between far apart regions due to the movement of people outside their base location. This paper describes the MPI-based implementation of EpiGraph and its performance evaluation when simulating large-scale scenarios. We evaluate the simulator both on a distributed memory system and on a shared memory system. Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
EuroMPI | 4 |
| 2013 | An approach for constructing private storage services as a unified fault-tolerant system
José Luis González 0002, Jesús Carretero 0001, Víctor Jesús Sosa Sosa, Juan F. Rodriguez Cardoso, Ricardo Marcelín-Jiménez |
J. Syst. Softw. | 2 |
| 2013 | An open framework for translating portable applications into operating system-specific wireless sensor networks applicationsabstractSUMMARY Wireless sensor networks (WSNs) are distributed systems integrated by tiny devices, called sensor nodes, with capabilities to monitor the environment and forward their measurements to a special node, the sink, where the results can be collected and further processed. The trend in WSN is moving towards heterogeneous networks that will contain different sensor nodes running different instances of custom operating systems. Given the growing demand of new hardware platforms and operating systems specifically designed for sensor nodes, the applications programming for sensor nodes is becoming a challenging process that needs to be alleviated. Currently, application programming for sensor nodes is a complex, ad hoc, and error‐prone process where the portability among different platforms has been sacrificed. In this paper, we propose an open framework aimed to achieve application portability in heterogeneous sensor networks. Our approach provides the programming abstractions needed to support the application development process for sensor nodes. We have implemented an open framework that provides a set of tools on top of the most popular WSN operating systems to translate portable applications to the native operating system in an automatic, simple, and transparent way for developers. We have also evaluated the applications thus generated in terms of productivity and overhead, by comparing their footprint to those originally developed in each specific operating system. The results show that the overhead is minimal—4% in the worst case—and in some cases, it was even possible to reduce the footprint by using code optimizations. Copyright © 2012 John Wiley & Sons, Ltd. Soledad Escolar, Jesús Carretero 0001 |
Softw. Pract. Exp. | 2 |
| 2012 | Geology: Modular Georecommendation in Gossip-Based Social NetworksabstractGeolocated social networks, combining traditional social networking features with geolocation information, have grown tremendously over the last few years. Yet, very few works have looked at implementing geolocated social networks in a fully distributed manner, a promising avenue to handle the growing scalability challenges of these systems. In this paper, we focus on georecommendation, and show that existing decentralized recommendation mechanisms perform in fact poorly on geodata. We propose a set of novel gossip-based mechanisms to address this problem, in a modular similarity framework called GEOLOGY. The resulting platform is lightweight, efficient, and scalable, and we demonstrate its advantages in terms of recommendation quality and communication overhead on a real dataset of 15,694 users from Foursquare, a leading geolocated social network. Jesús Carretero 0001, Florin Isaila, Anne-Marie Kermarrec, François Taïani, Juan Manuel Tirado |
ICDCS | 1 |
| 2012 | Reconciling Dynamic System Sizing and Content Locality through Hierarchical Workload ForecastingabstractThe cloud has recently surged as a promising paradigm for hosting scalable Web systems serving a large number of users with large workload variations. It makes possible to dynamically add and remove resources to horizontally scalable architectures in order to save costs, while maintaining the quality of service. However, in order to achieve these goals the resource management of a platform must include policies and mechanisms for dynamically resizing the system, redistributing content and redirecting user requests. In this work, we address the problem of reconciling dynamic system sizing and content locality. There are three main contributions of our study. First, we address the problem of determining the system size by employing a hierarchical prediction framework that proactively provisions resources based on statistical models of the incoming workload. Second, we show how to employ the hierarchical prediction framework for designing a dispatching mechanism which can be used with any request distribution policy. Third, we propose two novel prediction-based locality-aware request distribution policies: Oblivious Locality-Aware Request Distribution (OLARD) and Affinity-Based Locality-Aware Request Distribution (ABLARD). We demonstrate the advantages of using our hierarchical prediction framework and how our approach achieves a high content locality, while adapting to unexpected workload changes. Juan Manuel Tirado, Daniel Higuero, Florin Isaila, Jesús Carretero 0001 |
ICPADS | 4 |
| 2012 | Guaranteed-delivery in arbitrary dimensional Wireless Sensor Networks by means of recursive virtual coordinatesabstractDue to limitations in sensors' hardware and communication, routing in Wireless Sensor Networks (WSN) has required different approaches as compared to more powerful and conventional networks. One successful approach to the routing problem in WSN is based on geographic protocols that, however, must rely on coordinates and that are limited to two-dimensional networks. We propose here an approach that guarantees packet delivery in networks of any dimensionality. This approach combines a recursive coordinate assignment protocol based on the network topology and a specific routing protocol. Both protocols are simple, and the path lengths produced by the routing protocol are slightly larger than the shortest paths. Stefano Chessa, Soledad Escolar, Susanna Pelagatti, Paolo Baronti, Jesús Carretero 0001 |
ISCC | 5 |
| 2012 | iCanCloud: A Brief Architecture OverviewabstractDuring the last years the use of cloud computing environments for both researchers and enterprises is a major trend. On the one hand, cloud computing is a flexible and scalable paradigm that offers hardware and software services to be purchased by users. On the other hand, users have extra costs in several procedures, like calculating how many resources are needed by the applications. These tasks could be expensive and time-consuming. In this work we present iCanCloud, a simulation platform for modeling and simulating cloud environments. iCanCloud provides a complete set of modules for modeling with high level detail, a cloud computing environments and their underlying architecture. Gabriel G. Castañé, Alberto Nuñez, Jesús Carretero 0001 |
ISPA | 3 |
| 2012 | Optimization of Quality of Service in Wireless Sensor Networks Powered by Solar CellsabstractSensors equipped with solar cells and rechargeable batteries are useful in many outdoor, long-lasting applications. In these sensors the cycles of energy harvesting and battery recharge need to be managed appropriately in order to avoid sensor unavailability due to energy shortages. We suggest that adapting the sensor duty cycle (and specifically its environment sampling frequency) to the expected energy production and residual battery charge is very useful to avoid sensors unavailability. To this purpose we introduce a novel concept of QoS, which is measured in terms of sampling frequency, and we provide a mean to maximize the QoS, i.e. the extent of the period in which the sensor operates at the user's desired sampling frequency. Soledad Escolar, Stefano Chessa, Jesús Carretero 0001 |
ISPA | 3 |
| 2012 | Virtual I/O Forwarding for Cloud-based HPC ApplicationsabstractIn this work we present a flexible I/O virtualization solution, which is evaluated and compared to similar I/O forwarding technologies. Furthermore, we present a distributed I/O forwarding backend architecture which allows for complex, multi-host setups that can aim for different I/O strategies such as power proportionality and higher throughput. Pablo Llopis, Gonzalo Martín 0001, Borja Bergua, Jesús Carretero 0001 |
ISPA | 4 |
| 2012 | A Black Box Model for Storage Devices Based on Probability DistributionsabstractTraditional approaches for storage devices simulation have been based on detailed analytical models. However, detailed models require detailed computations which may be not affordable for large scale simulations. Moreover, highly detailed models cannot be easily generalized. A different approach is the black-box statistical modeling, where the storage device, its interface, and the interconnection mechanisms are modeled as a single stochastic process, defining the request response time as a random variable with an unknown distribution. A random variate generator can be built and integrated into a simulation model. This approach allows to generate a simulation model for both real and synthetic workloads. This article describes a method suitable for building fast simulation models for storage devices. Our method uses as starting point a workload and produces a random variate generator which can be easily integrated into large scale simulation models. A comparison between our variate generator and the widely known simulation tool DiskSim, shows that our variate generator is faster, and can be as accurate as DiskSim. Laura Prada, Alejandro Calderón 0001, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001 |
ISPA | 5 |
| 2012 | Enhancing File Transfer Scheduling and Server Utilization in Data Distribution InfrastructuresabstractThis paper presents a methodology for efficiently solving the file transfer scheduling problem in a distributed environment. Our solution is based on the relaxation of an objective-based time-indexed formulation of a linear programming problem. The main contributions of this paper are the following. First, we introduce a novel approach to the relaxation of the time-indexed formulation of the transfer scheduling problem in multi-server and multi-user environments. Our solution consists of reducing the complexity of the optimization by transforming it into an approximation problem, whose proximity to the optimal solution can be controlled depending on practical and computational needs. Second, we present a distributed deployment of our methodology, which leverages the inherent parallelism of the divide-and-conquer approach in order to speed-up the solving process. Third, we demonstrate that our methodology is able to considerably reduce the schedule length and idle time in a computationally tractable way. Daniel Higuero, Juan Manuel Tirado, Florin Isaila, Jesús Carretero 0001 |
MASCOTS | 4 |
| 2012 | Runtime Support for Adaptive Resource Provisioning in MPI Applications
Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
EuroMPI | 4 |
| 2012 | An ontology-driven decision support system for high-performance and cost-optimized design of complex railway portal frames
Ruben Saa, Alberto García Fernández, Carlos Gomez, Jesús Carretero 0001, Félix García Carballeira |
Expert Syst. Appl. | 4 |
| 2012 | Expanding the volunteer computing scenario: A novel approach to use parallel applications on volunteer computing
Alejandro Calderón 0001, Félix García Carballeira, Borja Bergua, Luis Miguel Sánchez, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 5 |
| 2012 | iCanCloud: A Flexible and Scalable Cloud Infrastructure Simulator
Alberto Nuñez, José Luis Vázquez-Poletti, Agustín C. Caminero, Gabriel G. Castañé, Jesús Carretero 0001, Ignacio Martín Llorente |
J. Grid Comput. | 5 |
| 2012 | Dynamic-CoMPI: dynamic optimization techniques for MPI parallel applications
Rosa Filgueira, Jesús Carretero 0001, David E. Singh, Alejandro Calderón 0001, Alberto Nuñez |
J. Supercomput. | 2 |
| 2011 | Predictive Data Grouping and Placement for Cloud-Based Elastic Server InfrastructuresabstractWorkload variations on Internet platforms such as YouTube, Flickr, LastFM require novel approaches to dynamic resource provisioning in order to meet QoS requirements, while reducing the Total Cost of Ownership (TCO) of the infrastructures. The economy of scale promise of cloud computing is a great opportunity to approach this problem, by developing elastic large scale server infrastructures. However, a proactive approach to dynamic resource provisioning requires prediction models forecasting future load patterns. On the other hand, unexpected volume and data spikes require reactive provisioning for serving unexpected surges in workloads. When workload can not be predicted, adequate data grouping and placement algorithms may facilitate agile scaling up and down of an infrastructure. In this paper, we analyze a dynamic workload of an on-line music portal and present an elastic Web infrastructure that adapts to workload variations by dynamically scaling up and down servers. The workload is predicted by an autoregressive model capturing trends and seasonal patterns. Further, for enhancing data locality, we propose a predictive data grouping based on the history of content access of a user community. Finally, in order to facilitate agile elasticity, we present a data placement based on workload and access pattern prediction. The experimental results demonstrate that our forecasting model predicts workload with a high precision. Further, the predictive data grouping and placement methods provide high locality, load balance and high utilization of resources, allowing a server infrastructure to scale up and down depending on workload. Juan Manuel Tirado, Daniel Higuero, Florin Isaila, Jesús Carretero 0001 |
CCGRID | 4 |
| 2011 | Multi-model prediction for enhancing content locality in elastic server infrastructuresabstractInfrastructures serving on-line applications experience dynamic workload variations depending on diverse factors such as popularity, marketing, periodic patterns, fads, trends, events, etc. Some predictable factors such as trends, periodicity or scheduled events allow for proactive resource provisioning in order to meet fluctuations in workloads. However, proactive resource provisioning requires prediction models forecasting future workload patterns. This paper proposes a multi-model prediction approach, in which data are grouped into bins based on content locality, and an autoregressive prediction model is assigned to each locality-preserving bin. The prediction models are shown to be identified and fitted in a computationally efficient way. We demonstrate experimentally that our multi-model approach improves locality over the uni-model approach, while achieving efficient resource provisioning and preserving a high resource utilization and load balance. Juan Manuel Tirado, Daniel Higuero, Florin Isaila, Jesús Carretero 0001 |
HiPC | 4 |
| 2011 | Optimizing Distributed Architectures to Improve Performance on Checkpointing ApplicationsabstractNowadays, satisfying the global throughput targets of each application in High Performance Computing systems is a difficult task because of the high number of architectural configurations having a considerable impact on the overall system performance, such as the number of storage servers, features of the communication links, number of CPU cores per node, etc. In this paper we have performed a thorough study of the compared performance of scaling up HPC cluster architectures using a checkpointing application model. This study is specifically focused on multi-core HPC clusters and the scaling process is oriented towards the three main resources: computing power, communications and storage. The main goal of this work is to evaluate and analyze how evolves both scalability and bottlenecks existent on different HPC multi-core architectures using different architectural configurations. In order to achieve this goal, a set of simulation experiments has been achieved using a simulation framework, called SIMCAN, specifically designed for modeling and simulating HPC architectures. The results obtained show that the computing power is well suited thanks to the multi-core processors, while the problems are found on the storage and on the communications channels, being the storage network the main bottleneck. Alberto Nuñez, Javier Fernández 0001, Jesús Carretero 0001, Laura Prada, Mario Blaum |
HPCC | 3 |
| 2011 | A Power-Aware Based Storage Architecture for High Performance ComputingabstractThe energy crisis of the last years and the ever increasing conscience about the negative effects of energy waste on the climate change have brought the sustainability both into public attention, industry, and scientific scrutiny. Energy demand has been increasing in many subsystems, specially in data centers and supercomputers. This paper considers the problem of saving energy on storage systems taking advantage of SSD devices. SSDs and magnetic disk devices offer different power characteristics, being SSD devices much less power consuming than conventional magnetic disk devices. We propose a novel power saving solution based on SSD devices, namely SSD-PASS. Our storage system obtains benefits of permanent caching on SSDs in storage nodes. Nowadays we can find solutions that do not consider the viability and feasibility of the SSD-based storage systems, in terms of monetary cost. We present a cost analysis and evaluate our proposed architecture, in terms of saved energy and performance. Our cost model takes into account magnetic disk and SSD devices reliability metrics, current energy prices, and replacement costs. The experimental results demonstrate that our solution achieves a significant reduction in energy consumption and subsequent monetary savings by up to 66\%. We have evaluated the proposed approach with realistic workloads of three well-known HPC applications. Laura Prada, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001, Alberto Nuñez |
HPCC | 4 |
| 2011 | Design of a New Cloud Computing Simulation Platform
Alberto Nuñez, José Luis Vázquez-Poletti, Agustín C. Caminero, Jesús Carretero 0001, Ignacio Martín Llorente |
ICCSA (3) | 4 |
| 2011 | Cross-layer optimization of Low Power Listening MAC protocols for Wireless Sensor NetworksabstractMAC protocols for Wireless Sensor Networks based on channel checking to detect incoming packets (such as Low Power Listening, LPL) require some form of synchronization among sensors, in order to schedule the times in which the sender and the receiver should turn on their radios, and to ensure correct packet delivery. However, while travelling along multihop paths, packets may accumulate delays that force the sensors either to incur packet losses and retransmissions or to readjust the planned scheduling. In both cases this fact impacts on the energy budget of the sensors. In this paper we propose a cross-layer optimization for MAC protocols that use LPL. This optimization takes into account high-level information of the application in order to compute adaptive delays in every sensor along a multihop path, with the goal of adjusting precisely the activity time window of the sensor along the path. We validate our delay-based model by evaluating different scenarios, and we compare it against the LPL model. The simulation results confirm the validity of our approach and demonstrate that a delay-based model can improve the synchronization achieved through the LPL strategy. Soledad Escolar, Stefano Chessa, Jesús Carretero 0001 |
ISCC | 3 |
| 2011 | Power saving-aware prefetching for SSD-based systems
Laura Prada, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2011 | Design and Evaluation of Multiple-Level Data Staging for Blue Gene SystemsabstractParallel applications currently suffer from a significant imbalance between computational power and available I/O bandwidth. Additionally, the hierarchical organization of current Petascale systems contributes to an increase of the I/O subsystem latency. In these hierarchies, file access involves pipelining data through several networks with incremental latencies and higher probability of congestion. Future Exascale systems are likely to share this trait. This paper presents a scalable parallel I/O software system designed to transparently hide the latency of file system accesses to applications on these platforms. Our solution takes advantage of the hierarchy of networks involved in file accesses, to maximize the degree of overlap between computation, file I/O-related communication, and file system access. We describe and evaluate a two-level hierarchy for Blue Gene systems consisting of client-side and I/O node-side caching. Our file cache management modules coordinate the data staging between application and storage through the Blue Gene networks. The experimental results demonstrate that our architecture achieves significant performance improvements through a high degree of overlap between computation, communication, and file I/O. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Robert Latham, Robert B. Ross |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | New Contributions for Simulating Large Distributed SystemsabstractNowadays, simulation of large distributed environments is a very important research field. Due to the large number of components to be simulated, the execution of those simulations requires a high level of resources such as CPU and memory. Thus, in order to increase the performance of those simulations, a feasible solution consists on splitting the complete model in sub-domains, where each sub-domain is executed in a single machine. In this paper we propose a strategy to automatically accomplish the parallelization of those environments, which has been implemented and tested in the SIMCAN simulation platform. Alberto Nuñez, Javier Fernández 0001, Jesús Carretero 0001 |
DS-RT | 3 |
| 2010 | Affinity P2P: A self-organizing content-based locality-aware collaborative peer-to-peer network
Juan Manuel Tirado, Daniel Higuero, Florin Isaila, Jesús Carretero 0001, Adriana Iamnitchi |
Comput. Networks | 4 |
| 2010 | Branch replication scheme: A new model for data replication in large scale data grids
José María Pérez, Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, Javier Fernández 0001 |
Future Gener. Comput. Syst. | 3 |
| 2010 | Scalable Storage Systems and High-Perfomance Applications
Jesús Carretero 0001, José Daniel García |
J. Supercomput. | 1 |
| 2010 | New techniques for simulating high performance MPI applications on large storage networks
Alberto Nuñez, Javier Fernández 0001, José Daniel García, Félix García Carballeira, Jesús Carretero 0001 |
J. Supercomput. | 5 |
| 2009 | Latency Hiding File I/O for Blue Gene SystemsabstractThis paper presents the design and implementation of a novel file I/O solution for Blue Gene systems. We propose a hierarchical I/O cache architecture based on open source software. Our solution is based on an asynchronous data staging strategy, which hides the latency of file system access from compute nodes. The performance results demonstrate the high scalability and significant performance improvements of our architecture over existing solutions. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Robert Latham, Samuel Lang, Robert B. Ross |
CCGRID | 3 |
| 2009 | Resource selection for fast large-scale Virtual Appliances PropagationabstractThe increase of Dynamic Virtual Infrastructures usage brings up some problems. One of them is the efficient deployment, over large-scaled distributed systems, of the Virtual Appliances images. To address this problem two points needs to be faced, which nodes to select and how to transfer the VA images to those nodes. In this paper we propose a function for efficient node selection. This customizable function can be tailored to prioritize distribution time or node performance. We study how to tailor the function in order to balance both factors. We evaluate the performance of this function selection in conjunction with a deployment algorithm called Geometric Propagation obtaining exceptional results. We propose a mechanism that allows deploying a VA image over a large number of nodes within a reasonable period of time. Alejandra Rodríguez, Jesús Carretero 0001, Borja Bergua, Félix García Carballeira |
ISCC | 2 |
| 2009 | Saving power in flash and disk hybrid storage systemabstractThis paper considers the question of saving energy in the disk drive making advantage of diverse devices in a hybrid storage system employing flash and disk drives. The flash and disk offer different power characteristics, being flash much less power consuming than the disk drive. We propose a technique that uses a flash device as a cache for a single disk device. We examine various options for managing the flash and disk devices in such a hybrid system and show that the proposed method saves energy in diverse scenarios. We implemented a simulator composed of disk and flash devices. This paper gives an overview of the design and evaluation of the proposed approach with the help of realistic workloads. Laura Prada, José Daniel García, Jesús Carretero 0001, Félix García Carballeira |
MASCOTS | 3 |
| 2009 | Scalability in data management
Jesús Carretero 0001, José Daniel García |
J. Supercomput. | 1 |
| 2009 | A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2009 | A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001 |
J. Supercomput. | 4 |
| 2008 | View-Based Collective I/O for MPI-IOabstractThis paper presents the design and implementation of a new file system independent collective I/O optimization based on file views: view-based collective I/O. View-based collective I/O has been implemented and evaluated inside ROMIO implementation of MPI-IO standard. The evaluation section shows that view-based I/O outperforms the original two-phase collective I/O from ROMIO in most of the cases for three well-known parallel I/O benchmarks. This is especially due to a smaller cost of scatter/gather operations, a reduction of the metadata overhead, and a smaller number of collective communication and synchronization primitives used in the implementation. Francisco Javier García Blas, Florin Isaila, David E. Singh, Jesús Carretero 0001 |
CCGRID | 4 |
| 2008 | Exploiting data compression in collective I/O techniquesabstractThis paper presents Two-Phase Compressed I/O (TPC I/O,) an optimization of the Two-Phase collective I/O technique from ROMIO, the most popular MPI-IO implementation. In order to reduce network traffic, TPC I/O employs LZO algorithm to compress and decompress exchanged data in the inter-node communication operations. The compression algorithm has been fully implemented in the MPI collective technique, allowing to dynamically use (or not) compression. Compared with Two-Phase I/O, Two-Phase Compressed I/O obtains important improvements in the overall execution time for many of the considered scenarios. Rosa Filgueira, David E. Singh, Juan Carlos Pichel, Jesús Carretero 0001 |
CLUSTER | 4 |
| 2008 | New techniques for simulating high performance MPI applications on large storage networksabstractIn this paper we present new techniques for simulating high performance MPI applications on large storage networks. Performance analysis of high performance application on large storage networks is a very complex and time-consuming task. However, modelling and studying the behaviour of any application on complex network architectures is crucial to obtain good performance. The goal of this work is to predict both scalability degree and performance of any high computing applications on any network architecture. A very interesting feature of this work is that our approach does not require to modify the application in order to simulate its behaviour. Also, there is no need to modify the simulator code to test different architectures. It can be done just creating a new configuration file. In order to perform those analyses we have used SIMCAN, a simulation tool to analyzing high-performance I/O architectures, developed at University Carlos III de Madrid. To validate this work we have used the BIPS3D application on several hardware-based architectures and on our simulator. The comparative results of those environments are presented to show the accuracy and efficiency of our approach. Alberto Nuñez, Javier Fernández 0001, José Daniel García, Jesús Carretero 0001 |
CLUSTER | 4 |
| 2008 | Reordering Algorithms for Increasing Locality on Multicore ProcessorsabstractIn order to efficiently exploit available parallelism, multicore processors must address contention for shared resources as cache hierarchy. This fact becomes even more important when irregular codes are executed on them, which is the case for sparse matrix ones. In this paper a technique for increasing locality of sparse matrix codes on multicore platforms is presented. The technique consists on reorganizing the data guided by a locality model which introduces the concept of windows of locality. The evaluation of the reordering technique has been performed on two different leading multicore platforms: Intel Core2Duo and Intel Xeon. Experimental results show important performance improvements when using our reordered matrices with respect to original ones. In particular, an average execution time reduction of about 30% is achieved considering different number of running threads. These results are due to an improved overall cache behavior. Likewise, a comparison of our proposal with some standard reordering techniques is included in the paper. Results point out that the reordering technique always outperforms standard algorithms and is effective for matrices with any structure. Juan Carlos Pichel, David E. Singh, Jesús Carretero 0001 |
HPCC | 3 |
| 2008 | M-PLAT: Multi-Programming Language Adaptive TutorabstractIn this paper we introduce M-PLAT, an intelligent tutoring system for helping students to learn the basics of programming languages. In fact, the M-PLAT system represents a full collection of intelligent tutoring systems, and due to its modular and hierarchical architecture it can be upgraded to deal with a new programming language that is not yet included in the system. Thus, this tutoring system is not limited to a unique programming language, making M-PLAT a very scalable system. The best important feature of our system is that M-PLAT dynamically adapts itself to the learning style of each student, optimizing the learning time to each student. Alberto Nuñez, Javier Fernández 0001, José Daniel García, Laura Prada, Jesús Carretero 0001 |
ICALT | 5 |
| 2008 | AHPIOS: An MPI-Based Ad Hoc Parallel I/O SystemabstractThis paper presents the design and implementation of a portable ad-hoc parallel I/O system (AHPIOS). AHPIOS virtualizes on-demand available distributed storage resources and allows the files to be striped over several storage devices. Additionally, the design unifies the configuration of the MPI-IO library and the AHPIOS data servers. By a strong integration of the application, MPI-IO library and file system, a significant performance improvement can be achieved. The experimental section shows that the full MPI-IO integrated AHPIOS implementation of file access operations outperforms the existing MPI-IO implementation by as much as 495% for file writes and 522% for file reads. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Wei-keng Liao, Alok N. Choudhary |
ICPADS | 3 |
| 2008 | Model for on-demand virtual computing architectures - OVCAabstractHigh performance computers are becoming popular for non scientific areas. Growth of computer capacity, virtualization techniques, and network capacity, give an opportunity to exploit the potential of those systems and to integrate different concepts for improving resource utilization. In this paper, we propose an architecture which manages groups of distributed resources, deploying infrastructures on-demand. Resources, no matter if they are local or remote, physical or virtual, are managed homogeneously. Highly customized environments are dynamically created to fulfill clientpsilas requirements, by means of virtual machines. The proposed architecture is based on distributed systems not linked to any specific implementation, which offers high flexibility and applicability. A functional prototype has been implemented and tested at our lab. Evaluation results demonstrate that our initial assumptions are correct. Alejandra Rodríguez, Javier Fernández 0001, Jesús Carretero 0001 |
ISCC | 3 |
| 2008 | Comparing Grid Data Transfer Technologies in the Expand Parallel File SystemabstractData management is one of the most important problems in grid environments. One important challenge facing grid computing is the design of a grid file system. The Global Grid Forum defines a grid file system as a human-readable resource namespace for management of heterogeneous distributed data resources, that can span across multiple autonomous administrative domains. This paper evaluates Expand, a new grid file system according to the Global Grid Forum recommendations that integrates heterogeneous data storage resources in grids using standard grid technologies: GridFTP and the OGSA ByteIO interface defined by the Open Grid Forum. Borja Bergua, Félix García Carballeira, Alejandro Calderón 0001, Luis Miguel Sánchez, Jesús Carretero 0001 |
PDP | 5 |
| 2007 | Optimization and evaluation of parallel I/O in BIPS3D parallel irregular applicationabstractThis paper presents the optimization and evaluation of parallel I/O for the BIPS3D parallel irregular application, a 3-dimensional simulation of BJT and HBT bipolar devices. The parallel version of BIPS3D employs Metis, a library for partitioning graphs, finite element meshes, or sparse matrices. First, we show how the partitioning information provided by Metis can be used in order to improve the performance of parallel I/O. Second, we propose a novel technique, called Interval Data Grouping (IDG), which exploits the data replication of mesh nodes for optimizing the scheduling of the parallel file operations. Finally, we evaluate the parallel I/O version of BIPS3D for various existing parallel I/O techniques and present an in-depth analysis of the IDG performance. Rosa Filgueira, David E. Singh, Florin Isaila, Jesús Carretero 0001, Antonio J. García-Loureiro |
IPDPS | 4 |
| 2007 | Multiple-Phase Collective I/O Technique for Improving Data Access LocalityabstractThis paper presents multiple-phase collective I/O, a novel collective I/O technique for distributed memory multiprocessors. Multiple-phase collective I/O is a refinement of two-phase collective I/O technique. The communication phase is structured into several steps, which progressively increase the locality of the data to be written to a file system. Besides the description of multiple-phase collective I/O, our paper addresses two additional objectives. First, the authors target to improve the efficiency of the sulphur transport Eurelian model 2 (STEM-II) application. STEM-II is an air quality model that simulates transport, chemical transformations, emission and deposition processes in a unified framework. Due to the large amount of processed data, I/O becomes a critical factor for the application performance. Multiple-phase collective I/O, considerably enhances the performance of the I/O stage in particular and, consequently, of the whole application in general. Second objective consists of evaluating and comparing the performance of multiple-phase collective I/O with that of other well known parallel I/O techniques David E. Singh, Florin Isaila, Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001 |
PDP | 5 |
| 2007 | Dispatching Requests in Partially Replicated Web Clusters - An Adaptation of the LARD Algorithm
José Daniel García, Laura Prada, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Luis Miguel Sánchez |
WEBIST (1) | 3 |
| 2007 | A global and parallel file system for grids
Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, José Daniel García, Luis Miguel Sánchez |
Future Gener. Comput. Syst. | 2 |
| 2006 | On the Reliability of Web Clusters with Partial Replication of ContentsabstractTraditionally, distributed Web servers have used two strategies for allocating files on server nodes: full replication and full distribution. While full replication provides a highly reliable solution, it limits storage capacity to the capacity of the smallest node. On the other hand, full distribution provides higher storage capacity at the cost of lower reliability. A hybrid solution is partial replication where every file is allocated to a small number of nodes. The most promising architecture for a partial replication strategy is the Web cluster architecture. However, Web clusters present a big flaw from reliability perspective as they contain a single point of failure. To correct this flaw, in this paper we present a modified architecture: the Web cluster with distributed Web switch. Reliability of Web clusters is evaluated for different replication strategies. System evaluations show that our proposal leads to a highly reliable solution with high scalability. José Daniel García, Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, David E. Singh, Alejandro Calderón 0001 |
ARES | 2 |
| 2006 | Integrating Logical and Physical File Models in the MPI-IO Implementation for "Clusterfile"abstractThis paper presents the design and implementation of the MPI-IO interface for the Clusterfile parallel file system. The approach offers the opportunity of achieving a high correlation between the file access patterns of parallel applications and the physical file distribution. First, any physical file distribution can be expressed by means of MPI data types. Second, mechanisms such as views and collective I/O operations are portably implemented inside the file system, unifying the I/O scheduling strategies of the MPI-IO library and the file system. The experimental section demonstrates performance benefits of more than one order of magnitude. Florin Isaila, David E. Singh, Jesús Carretero 0001, Félix García Carballeira, Gabor Szeder, Thomas Moschny |
CCGRID | 3 |
| 2006 | A Quantitative Justification to Partial Replication of Web Contents
José Daniel García, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Alejandro Calderón 0001, David E. Singh |
ICCSA (4) | 2 |
| 2006 | A New I/O Architecture for Improving the Performance in Large Scale Clusters
Luis Miguel Sánchez, Florin Isaila, Félix García Carballeira, Jesús Carretero 0001, Rolf Rabenseifner, Panagiotis A. Adamidis |
ICCSA (5) | 4 |
| 2006 | MAPFS: A flexible multiagent parallel file system for clusters
María S. Pérez 0001, Jesús Carretero 0001, Félix García Carballeira, José M. Peña 0002, Víctor Robles |
Future Gener. Comput. Syst. | 2 |
| 2005 | High Performance Java Input/Output for Heterogeneous Distributed ComputingabstractCurrently there is a growing interest in using Java for high performance computing. Java has many advantages for high performance computing: it is based on a high-level and object-oriented programming model with support for multithreading and distributed computing. Furthermore, Java 's virtual machine allows applications to run on multiple heterogeneous platforms. A major problem with the use of Java for high performance computing is the I/O. This problem has been solved traditionally in clusters using parallel file systems and parallel I/O libraries, however there is a lack of parallel file systems for Java applications. In this paper, we present a Java parallel I/O library called jExpand. It provides high performance I/O by using several NFS servers in parallel, as NFS can be found in multiple platforms (Linux, Solaris, Windows 2000, etc), we provide a universal parallel file system that can be used everywhere. jExpand requires no changes in the NFS server as it uses RPC operations to provide parallel access to the same file. The paper describes the design, implementation and evaluation of jExpand. José María Pérez, Luis Miguel Sánchez, Félix García Carballeira, Alejandro Calderón 0001, Jesús Carretero 0001 |
ISCC | 5 |
| 2004 | A Model for Use Case Priorization Using Criticality Analysis
José Daniel García, Jesús Carretero 0001, José María Pérez, Félix García Carballeira |
ICCSA (4) | 2 |
| 2004 | An Adaptive Cache Coherence Protocol Specification for Parallel Input/Output SystemsabstractCaching has been intensively used in memory and traditional file systems to improve system performance. However, the use of caching in parallel file systems and I/O libraries has been limited to I/O nodes to avoid cache coherence problems. We specify an adaptive cache coherence protocol that is very suitable for parallel file systems and parallel I/O libraries. This model exploits the use of caching, both at processing and I/O nodes, providing performance improvement mechanisms such as aggressive prefetching and delayed-write techniques. The cache coherence problem is solved by using a dynamic scheme of cache coherence protocols with different sizes and shapes of granularity. The proposed model is very appropriate for parallel I/O interfaces, such as MPI-IO. Performance results, obtained on an IBM SP2, are presented to demonstrate the advantages offered by the cache management methods proposed. Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, José María Pérez, José Daniel García |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2003 | Data Allocation and Load Balancing for Heterogeneous Cluster Storage SystemsabstractDistributed filesystems are a typical solution in networked environments as clusters and grids. Parallel filesystems are a typical solution in order to reach high performance I/O distributed environment, but those filesystems have some limitations in heterogeneous storage systems. Usually in distributed systems, load balancing is used as a solution to improve the performance, but typically the distribution is made between peer-to-peer computational resources and from the processor point of view. In heterogeneous systems, like heterogeneous clusters of workstations, the existing solutions do not work so well. However, the utilization of those systems is more extended every day, having an extreme example in the grid environment. In this paper we bring attention to those aspects of heterogeneous distributed data systems presenting a parallel file system that take into account heterogeneity of storage nodes, the dynamic addition of new storage nodes, and an algorithm to group requests in heterogeneous systems. José María Pérez, Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, Luis Miguel Sánchez |
CCGRID | 3 |
| 2003 | Video Forwarding Techniques for Mixed Wired and Wireless NetworksabstractDuring the last years, Internet video streaming has experiences a phenomenal growth. This is happening despite the notorious difficulties of transmitting data packets with a deadline over the Internet, due to variability in throughput, delays and losses. These problems arise significantly when using wireless networks where the available bandwidth is low and the losses are important due to its error prone transmission nature. In this paper we propose a fast-forwarding technique that is based on segmenting the movie on different files. Normal movie reproduction requires all the files, but fast-forwarding reproduction only requires one file. Those files can me merged by the client or by the server. The segmentation is frame based, grouping all the frames that can be independently decoded together. The resulting file can be showed with any existing player. This group of frames would be the ones to use in a fast-forward reproduction. Our techniques can also be useful in adaptive environments, like wireless networks, because there is no problem for the fast-forward file to use the same optimizations that exist for full movie files. This method also reduces the storage bandwidth and the storage size needed (there is no extra data for fast-forwarding). We also propose a video server architecture that takes advantage of this technique to achieve full interactive video reproduction. The evaluation results shown in this paper demonstrates that our technique enhances video fast-forwarding operations. Javier Fernández 0001, Jesús Carretero 0001, Félix García Carballeira, José María Pérez, Alejandro Calderón 0001, José J. Muñoz |
ISCC | 2 |
| 2003 | A hierarchical disk scheduler for multimedia systems
Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, Alok N. Choudhary |
Future Gener. Comput. Syst. | 1 |
| 2002 | MAPFS_MAS: A Model of Interaction among Information Retrieval AgentsabstractMAPFS is a parallel file system integrated with a multiagent system responsible for the information retrieval [2]. The use of a multiagent system implies coordination among the agents such system consists of. The principal María S. Pérez 0001, Félix García Carballeira, Jesús Carretero 0001 |
CCGRID | 3 |
| 2002 | Design and Implementation of a Parallel I/O Runtime System for Irregular Applications
Jaechun No, Sung-Soon Park 0001, Jesús Carretero 0001, Alok N. Choudhary |
J. Parallel Distributed Comput. | 3 |
| 2001 | New Techniques for Collective Communications in Clusters: A Case Study with MPIabstractThe paper describes new techniques to increase the performance of collective communication operations in clusters. These techniqnes are based in multithreading operations and on-line data compression. The techniques proposed have been implemented in MiMPI, a thread-safe implementation of MPI. We have evaluated, and compared, the performance of MiMPI with other implementations of MPI available for clusters with Linux and Windows 2000. The benchmark used has been MPBench, a flexible and portable framework to allow benchmarking of MPI implementations. Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001, Javier Fernández 0001, Oscar Pérez |
ICPP | 3 |
| 1997 | Evaluating ParFiSys: A high-performance parallel and distributed file system
Jesús Carretero 0001, Francisco García 0001, Pedro de Miguel |
J. Syst. Archit. | 2 |
| 1997 | Performance Increase Mechanisms for Parallel and Distributed File Systems
Jesús Carretero 0001, Pedro de Miguel, Francisco García 0001 |
Parallel Comput. | 1 |
| 1993 | A fault-tolerant server on MACH
S Arévalo, Jesús Carretero 0001, J. L. Castellanos, F. Barco |
Microprocess. Microprogramming | 2 |