EDBT 2026 Demo / reviewers in the wild / expert
Vicente Blanco 0001
dblp:02/3794 · also Vicente Blanco Pérez
· DBLP profile ↗
41ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-1166-6310ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 2:4 Pruning on Edge Devices: Performance, Energy Efficiency and Accuracy
Nicolás Hernández, Pedro A. Toledo, Vicente Blanco 0001, Francisco Almeida |
Euro-Par (2) | 3 |
| 2025 | Correction: Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator (VASP)
Isidoro Nieves-Pírez, Alfonso Muñoz, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 4 |
| 2025 | KubePipe: a container-based high-level parallelization tool for scalable machine learning pipelinesabstractAbstract As the complexity and scale of machine learning applications continue to grow, the need for efficient training methodologies becomes increasingly critical. Traditional training processes can be time-intensive, often limiting rapid development and deployment. In response to this challenge, we present KubePipe, a high-level tool that abstracts parallelism and containerization from the user, allowing non-expert users to leverage advanced parallel architectures without requiring deep knowledge of parallel computing or container orchestration. KubePipe enables the concurrent execution of multiple machine learning workflows within a Kubernetes cluster, optimizing computational resources and significantly reducing training times. By leveraging containerized environments, KubePipe ensures a high degree of modularity, scalability, and portability, making it adaptable to various machine learning frameworks and tasks. Our experimental results demonstrate substantial performance improvements when using KubePipe compared to conventional pipeline implementations. This paper explores the architecture and functionality of KubePipe, providing insights into its integration with existing machine learning systems and highlighting its potential to streamline the training process in high-performance computing environments. Daniel Suárez, Francisco Almeida, Vicente Blanco 0001, Pedro A. Toledo |
J. Supercomput. | 3 |
| 2025 | Singularity to deploy HPC applications: a study case with WRFabstractAbstract This study evaluates the performance and portability of the Weather Research and Forecasting model in high-performance computing environments, comparing traditional baremetal deployments with containerized executions using Singularity. Experiments were conducted on the TeideHPC system, focusing on execution time, the impact of different compilers (Intel vs. GCC), and parallelization strategies (MPI and OpenMP). The results indicate that while Singularity introduces a performance overhead of 11 to 15%, it offers significant advantages in portability and reproducibility. Additionally, the study highlights the importance of compiler choice and the influence of container image size on startup times, emphasizing the need for careful optimization in containerized HPC workflows. Pierre-Simon Callist Yannick Tondreau, Juan Carlos Pérez, Juan Pedro Díaz, Vicente Blanco 0001, Jonatán Felipe |
J. Supercomput. | 4 |
| 2024 | AI4WATER: A Digital Twin for Irrigated AgricultureabstractThis study presents a Digital Twin (DT) that is being created to optimize the use of the available hydric resources, and mitigate the effects of the increasing water shortage in irrigated agriculture in fields in the Urgell channel region (Lleida). A DT is "a virtual representation of an object or system that spans its lifecycle, it is updated from real-time data, and uses simulation, machine learning and reasoning to help decision-making." It will model the water fluxes using the knowledge of the amounts of water taken in, used, and returned to the environment, and other parameters that impact the water budget, such as atmospheric variables (temperature, water vapor deficit, relative humidity, solar radiance…), surface soil moisture, and evapotranspiration maps, etc. Satellite Earth Observation (EO) data, collocated with in-situ data from a network of 20 soil moisture probes and 2 meteo stations will be used to train the DT. Additionally, a rover-based ground penetrating radar will be used for cross-calibration. Adriano Camps, Carlos López-Martínez, Amadeu Gonga, Guillem Gracia-Sola, Adrián Pérez 0001, Alberto Alonso-González, Mercè Vall-Llossera, Hyuk Park 0001, Vicente Blanco 0001, Oriol Caselles, Carles Domenech, Paul Catala, Joan Adrià Ruiz-de-Azua, Montserrat Solsona |
IGARSS | 9 |
| 2024 | Parallel programming in mobile devices with FancyJCLabstractAbstract Mobile devices and handheld systems, such as the smartphones and tablets universally extended, are becoming increasingly powerful. Their basic hardware configuration is usually state-of-the-art heterogeneous architectures consisting of multi-core processors and some kind of accelerator such as GPUs or DSPs. Specific code adapted to the architecture is mandatory if high-performance computation is required and low-level libraries and parallelism are needed, which constitutes an important barrier for the usual developer in such devices. In this context, we propose the FancyJCL framework. It provides a high-level abstraction layer that hides implementation details and allows to develop parallel programs for mobile devices. The target platform for FancyJCL is mainly Android and Java developers due to their high market penetration. A very simple, seemingly sequential encoding results in parallel efficient OpenCL code. FancyJCL is itself based on the Fancier framework, which enables optimal memory management across memory spaces on unified memory systems. Benchmarks of FancyJCL code developed for a wide range of image processing algorithms show good performance with low development effort. Sergio Afonso, Óscar Gómez-Cárdenes, Paula Expósito, Vicente Blanco 0001, Francisco Almeida |
J. Supercomput. | 4 |
| 2024 | Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniquesabstractAbstract This document addresses some inherent problems in Machine Learning (ML), such as the high computational and energy costs associated with their implementation on IoT devices. It aims to study and analyze the performance and efficiency of quantization as an optimization method, as well as the possibility of training ML models directly on an IoT device. Quantization involves reducing the precision of model weights and activations while still maintaining acceptable levels of accuracy. Using representative networks for facial recognition developed with TensorFlow and TensorRT, Post-Training Quantization and Quantization-Aware Training are employed to reduce computational load and improve energy efficiency. The computational experience was conducted on a general-purpose computer featuring an Intel i7-1260P processor and an NVIDIA RTX 3080 graphics card used as an accelerator. Additionally, a NVIDIA Jetson AGX Orin was used as an example of an IoT device. We analyze the feasibility of training on an IoT device, the impact of quantization optimization on knowledge transfer-trained models and evaluate the differences between Post-Training Quantization and Quantization-Aware Training in such networks on different devices. Furthermore, the performance and efficiency of NVIDIA’s inference accelerator (Deep Learning Accelerator - DLA, in its 2.0 version) available at the Jetson Orin architecture are studied. We concluded that the Jetson device is capable of performing training on its own. The IoT device can achieve inference performance similar to that of the more powerful processor, thanks to the optimization process, with better energy efficiency. Post-Training Quantization has shown better performance, while Quantization-Aware Training has demonstrated higher energy efficiency. However, since the accelerator cannot execute certain layers of the models, the use of DLA worsens both the performance and efficiency results. Nicolás Hernández, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2024 | Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator (VASP)abstractAbstract This work tackles the performance and energy consumption analysis of a legacy scientific application, the VASP (Vienna Ab-initio Simulation Package), an application commonly used by physicists and chemists for modeling materials at the atomic scale. Many of these scientific applications have been implemented in Fortran, where energy metrics instrumentation is not straightforward. We obtained performance figures (execution time and energy consumption) by instrumenting the source code using EML. This energy measurement library has been modified to introduce Fortran interfaces for these metrics. The analysis was carried out using different matrix algebra libraries, parallelization techniques, and hardware platforms, emphasizing on the MPI, OpenMP, and CUDA parallel implementations of the algorithms used in VASP. We employ various material specifications (atomic structures) and molecular sizes of a silicon-based crystal to create a set of benchmarks for these specifications, leading to some recommendations for final users regarding performance improvements. The proposed benchmarking technique assists the user in selecting the right combination of problem size, compilers, and parallelization options available in VASP. For a given system platform, the user will be able to determine not only the architecture to use (GPU or multicore processors), but also the appropriate library and parallelization according to the atomic structure and molecular size. Isidoro Nieves-Pírez, Alfonso Muñoz, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 4 |
| 2024 | Comprehensive analysis of energy efficiency and performance of ARM and RISC-V SoCsabstractAbstract Over the past few years, ARM has been the dominant player in embedded systems and System-on-Chips (SoCs). With the emergence of hardware platforms based on the RISC-V architecture, a practical comparison focusing on their energy efficiency and performance is needed. In this study, our goal is to comprehensively evaluate the energy efficiency and performance of ARM and RISC-V SoCs in three different systems. We will conduct benchmark tests to measure power consumption and overall system performance. The results of our study are valuable to developers and researchers looking for the most appropriate hardware platform for energy-efficient computing applications. Our observations suggest that RISC-V Instruction Set Architecture (ISA) implementations may demonstrate lower average power consumption than ARM, but this does not automatically imply a superior performance per watt ratio for RISC-V. The primary focus of the study is to evaluate and compare these ISA implementations, aiming to identify potential areas for enhancing their energy efficiency. Furthermore, to ensure the practical applicability of our findings, we will use the Computational Fluid Dynamics software OpenFOAM. This step serves to validate the relevance of our results in real-world scenarios. It allows us to fine-tune execution parameters based on the insights gained from our initial study. By doing so, we aim not only to provide meaningful conclusions but also to investigate the transferability of our results to practical applications. Our analysis will also scrutinize the capabilities of these SoCs when handling nonsynthetic software workloads, thereby broadening the scope of our evaluation. Daniel Suárez, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2020 | A Dynamic Multi-Objective Approach for Dynamic Load Balancing in Heterogeneous SystemsabstractModern standards in High Performance Computing (HPC) have started to consider energy consumption and power draw as a limiting factor. New and more complex architectures have been introduced in HPC systems to afford these new restrictions, and include coprocessors such as GPGPUs for intensive computational tasks. As systems increase in heterogeneity, workload distribution becomes a more core problem to achieve the maximum efficiency in every computational component. We present a Multi-Objective Dynamic Load Balancing (DLB) approach where several objectives can be applied to tune an application. These objectives can be dynamically exchanged during the execution of an algorithm to better adapt to the resources available in a system. We have implemented the Multi-Objective DLB together with a generic heuristic engine, designed to perform multiple strategies for DLB in iterative problems. We also present Ull Multiobjective Framework (UllMF), an open-source tool that implements the Multi-Objective generic approach. UllMF separates metric gathering, objective functions to be optimized and load balancing algorithms, and improves code portability using a simple interface to reduce the costs of new implementations. We illustrate how performance and energy consumption are improved for the implemented techniques, and analyze their quality using different DLB techniques from the literature. Alberto Cabrera Pérez, Alejandro Acosta, Francisco Almeida, Vicente Blanco 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | A heuristic technique to improve energy efficiency with dynamic load balancing
Alberto Cabrera Pérez, Alejandro Acosta, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 4 |
| 2015 | Energy Measurement Library (EML) Usage and Overhead AnalysisabstractEnergy consumption and efficiency analysis has raised as an interesting topic to address for high performance computing researchers. Exascale computers, with current capabilities, would require a huge amount of power to operate at full capacity. This has attracted the interest of many researchers in the field contributing to elaborate multiple libraries and tools to measure energy and power consumption. In spite of the efforts, there is no current standard for energy measurements. It is known that the accuracy of the measurements is highly dependent on measurement tools (power meters, performance counters, ). A standard library should offer independence between hardware and software and portability for the instrumented code between different architectures. Additionally, the overhead of the library should be low enough to achieve accurate measurements. We proposed EML (Energy Measurement Library) and now we present insights on the usage and the overhead of the library, so that it could contribute to the establishment of an standard for energy measurement. Alberto Cabrera Pérez, Francisco Almeida, Javier Arteaga, Vicente Blanco 0001 |
PDP | 4 |
| 2014 | Modeling the performance of parallel applications using model selection techniquesabstractSUMMARY Nowadays, parallel architectures are changing so fast that there is a need for scalable and efficient tools to analyze and predict the performance of parallel applications. Analytical models are proved to be a useful approximation for characterizing parallel algorithms, but developing accurate analytical models is a hard issue, and, in general, they provide coarse performance predictions due to their intrinsic lack of accuracy. In this paper, we describe in detail the Tools for Instrumentation and Analysis (TIA) framework, an easy‐to‐use tool that automatically obtains accurate performance models by means of analytical expressions. This framework automatizes most of its internal tasks, reducing opportunities for human error, and it only requires the user to focus on the metrics and execution parameters that might influence the performance, those that should be considered in the modeling process. Its main advantage over other tools is that TIA uses model selection techniques that allow the automation of the modeling process. As a case of study, the use of TIA to obtain analytical models of different implementations of the broadcast collective communication in a cluster of multicores is shown. The results obtained by TIA are evaluated and compared with theoretical approaches based on the LogGP model. Copyright © 2013 John Wiley & Sons, Ltd. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Analytical Energy Models for MPI Communications on a Sandy-Bridge ArchitectureabstractComplexity in parallel computation for current HPC systems has being modeled in order to estimate time and energy consumption, which have been conduced to evaluate real costs. Communication analytical models for time and energy comsumption are an attractive issue as part of this cost analysis in parallel computation. In this paper, we present time and energy analytical models for MPI communications using OSU micro benchmarks under Sandy Bridge-EP computer architecture. RAPL (Running Average Power Limit) interface has been setup to provide mechanisms to enforce real energy consumption measurements for OSU micro benchmarks on this architecture. Our applied methodology defines an experimental framework based on several core configurations and different transmission rates. The RAPL interface allows us to obtain different energy values for the OSU benchmark executions. The measures obtained are used to obtain time and energy consumption parameters of our models, establishing average errors below 1% for almost all configurations in real executions under this computer architecture for an ad-hoc model. Alternative and simpler time and energy models are also proposed with errors up to 3.1% for the worst case. Francisco Almeida, Vicente Blanco 0001, Isidro Gonzalez, Alberto Cabrera Pérez, Domingo Giménez |
ICPP | 2 |
| 2013 | Analytical Modeling of the Energy Consumption for the High Performance LinpackabstractComparable to time performance models, it is now possible to estimate performance based upon energy consumption for HPC systems. The predictive ability of the analytical modeling is an interesting feature that motivates us to approach this methodology for the case of energy consumption. In this paper, we present an analytical model for predicting the energy consumption for the High Performance Linpack (HPL). The derived model can be used to know in advance the energy consumed by the HPL over a target architecture, and can be integrated into the schedulers of operating systems or queue managers. We established an experimental setup using a standard metered PDU that allowed us to measure the energy consumption for the HPL benchmark on our cluster. With the monitoring system in place, we can obtain the architectural and algorithmic parameters associated for both performance and energy analytical models. Also this has made possible watts and gflops-per-watt prediction when we execute Linpack executions with concrete algorithm parameters in our cluster. Alberto Cabrera Pérez, Francisco Almeida, Vicente Blanco 0001, Domingo Giménez |
PDP | 3 |
| 2013 | Modeling energy consumption for master-slave applications
Francisco Almeida, Vicente Blanco 0001, Alberto Cabrera Pérez, J. Ruiz |
J. Supercomput. | 2 |
| 2012 | OpenCF-R: R in the Cloud
Juan Carlos Castillo, Francisco Almeida, Vicente Blanco 0001, Adrián Santos |
CLOSER | 3 |
| 2012 | Towards the Dynamic Load Balancing on Heterogeneous Multi-GPU SystemsabstractThe advent of multicore systems, joined to the potential acceleration of the graphics processing units, alleviates some well known important architectural problems at the expense of a considerable increment of the programmability wall. The heterogeneity, both at architectural and programming level at the same time, raises the programming difficulties. Adapting existing code and libraries to such systems is a fundamental problem. The performance of this code is affected by the large interdependence between the code and the parallel architecture. We have developed a dynamic load balancing library that allows parallel code to be adapted to a wide variety of heterogeneous systems. The overhead introduced by our system is minimal and the cost to the programmer negligible. This system has been applied to solve load imbalance problems appearing in homogeneous and heterogeneous multi-GPU platforms. As case studies, we consider matrix multiply and resource allocation problems, in different heterogeneous scenarios in multi-GPU systems. The unbalanced nature of these algorithms and heterogeneous systems allowed us to test the success of our load balancing approach. Alejandro Acosta, Vicente Blanco 0001, Francisco Almeida |
ISPA | 2 |
| 2012 | Model Selection to Characterize Performance Using Genetic AlgorithmsabstractThe TIA modeling framework provides analytical models of the performance of parallel applications. The resulting models are obtained using model selection techniques and are accurate enough for various purposes. Its main drawback is that the completion time depends on the number of candidate models and, in some situations, it becomes critical. In this work, a genetic algorithm is proposed for reducing the time for searching of the best candidate model. The use of this genetic algorithm to obtain the performance model of the linear implementation of the broadcast collective communication in a cluster of multicores is shown. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
ISPA | 5 |
| 2011 | Estimating the effect of cache misses on the performance of parallel applications using analytical modelsabstractIn this paper a methodology to characterize the influence of cache misses on the performance of parallel applications is presented. This methodology is based on analytical models provided by the TIA framework. This framework obtains analytical models of given observable quantities by instrumenting the source code and applying model selection techniques. In particular, two metrics related with the performance are considered in this work: the number of cache misses and the elapsed time. Based on both models, the influence in terms of execution time due to the cache misses can be inferred. Two different versions of the parallel product of dense matrices are used as case of study. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
AICCSA | 2 |
| 2011 | PETransWS: Web Service Computing Platform for Logistics and TransportationabstractIn large organizations and small firms in transportation, there is a growing need to use and analyze spatial data. Transportation system analysis and planning as well as mobility studies frequently use Geographic Information Systems (GIS). In this paper we propose the development of a web services platform dedicated to transportation and logistics. Taking advantage of the web services development framework PyOpenCF we integrate in the same environment services oriented to geolocalization, logistical optimization, etc. We develop a PyOpenCF client whose graphic interface allows processes with spatial data to be launched and provides a visualization of the results. Francisco Almeida, Vicente Blanco 0001, Julio Brito, Andres Crespo, José A. Moreno-Pérez, Adrián Santos |
PDP | 2 |
| 2011 | Adaptive load balancing of iterative computation on heterogeneous nondedicated systems
José Antonio Martínez, Francisco Almeida, Ester M. Garzón, Alejandro Acosta, Vicente Blanco 0001 |
J. Supercomput. | 5 |
| 2011 | Using accurate AIC-based performance models to improve the scheduling of parallel applications
Diego Rodríguez Martínez, Julio L. Albín, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera, Vicente Blanco 0001 |
J. Supercomput. | 6 |
| 2011 | Web services based scheduling in OpenCF
Adrián Santos, Francisco Almeida, Vicente Blanco 0001, Juan Carlos Castillo |
J. Supercomput. | 3 |
| 2010 | CoEDApplets - Collaborating in the Development of Teaching-oriented Applets
Francisco Almeida, Vicente Blanco 0001, J. Regalado, Adrián Santos |
CSEDU (1) | 2 |
| 2010 | Performance Modeling of MPI Applications Using Model Selection TechniquesabstractA new method for obtaining models of the performance of parallel applications based on statistical analysis is presented in this paper. This method is based on the Akaike's information criterion (AIC) that provides an objective mechanism to rank different models by means of an experimental data fit. The input of the modeling process is a set of variables and parameters that can a priori influence the performance of the application. This set can be provided by the user. Using this information, the method automatically generates a set of candidate models. These models are fit to the experimental data and the AIC score of each model is calculated. The model with the best AIC score is selected as the best model. Also, using the AIC scores of all candidate models, useful statistical information is provided to help the user to evaluate the quality of the selected model, as well as indications of how to interactively improve this modeling process. As a first case of study, statistical models obtained for different implementations of the broadcast collective communication in Open MPI are shown. These models are very accurate, exceeding its adjustment to theoretical approaches based on the LogGP model. Finally, the NAS Parallel Benchmark is also characterized using this new method with good results in terms of accuracy. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
PDP | 5 |
| 2009 | EDApplets: A Web Tool for Teaching Data Structures and Algorithmic Techniques
Francisco Almeida, Vicente Blanco 0001, J. Regalado, Adrián Santos |
CSEDU (2) | 2 |
| 2009 | Accurate analytical performance model of communications in MPI applicationsabstractThis paper presents a new LogP-based model, called LoOgGP, which allows an accurate characterization of MPI applications based on microbenchmark measurements. This new model is an extension of LogP for long messages in which both overhead and gap parameters perform a linear dependency with message size. The LoOgGP model has been fully integrated into a modelling framework to obtain statistical models of parallel applications, providing the analyst with an easy and automatic tool for LoOgGP parameter set assessment to characterize communications. The use of LoOgGP model to obtain a statistical performance model of an image deconvolution application is illustrated as a case of study. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
IPDPS | 5 |
| 2009 | Using Web Services for Performance Monitoring and SchedulingabstractThe adoption of Web Service standards provides us with an increased level of manageability, extensibility and interoperability between loosely coupled services.The adoption of Web Services technologies atop sites for performance monitoring and scheduling will improve the efficient use of the computational resources. Web Services provide the ability to decompose HPC resources and functionality into a set of discoverable and loosely coupled services, which are capable of interaction in heterogeneous environments. At the same time, Web Services can address many of the interoperability issues that can be encountered in large scale systems.End users can access to these services to decide which system will be the most suitable for their needs. Other tools like schedulers can use the resources available as services to optimize the HPC resources and minimize jobs waiting time. Adrián Santos, Francisco Almeida, Vicente Blanco 0001, David Diez, Jonás Regueira, Esaú Sicilia |
PDP | 3 |
| 2009 | IDEWEP: Web service for astronomical parallel image deconvolution
Francisco Almeida, Vicente Blanco 0001, Carlos Delgado, Francisco de Sande, Adrián Santos |
J. Netw. Comput. Appl. | 2 |
| 2009 | Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
J. Supercomput. | 5 |
| 2008 | Towards the Automatic Service Generation and Scheduling in the OpenCF ProjectabstractWeb services-based technologies have emerged as a technological alternative for computational web portals. Facilitating access to distributed resources through web interfaces while simultaneously ensuring security is one of the main goals in most of the currently existing manifold tools and frameworks. OpenCF, the open source computational framework that we have developed, shares these objectives and adds others, like enforced portability, genericity, modularity and compatibility with a wide range of high performance computing systems. Adrián Santos, Francisco Almeida, Vicente Blanco 0001, David Diez, Jonás Regueira, Esaú Sicilia |
CISIS | 3 |
| 2007 | Lightweight Web Services for High Performace Computing
Adrián Santos, Francisco Almeida, Vicente Blanco 0001 |
ECSA | 3 |
| 2007 | Software Tools for Performance Modeling of Parallel ProgramsabstractThis paper presents a framework based on a user driven methodology to obtain analytical models of MPI applications on parallel systems in a systematic and easy to use way. This methodology consists of two stages. In the first one, instrumentation of the source code is performed using CALL, which is a profiling tool for interacting with the code in an easy, simple and direct way. New features are added to CALL to obtain different performance metrics and store the performance information in XML files. Using this information, an analytical model of the performance behavior is obtained in the second stage by means of R, a language and environment for statistical analysis. The structure of the whole framework is detailed in this paper, and some selected examples are used to show its practical use. Diego Rodríguez Martínez, Vicente Blanco 0001, Marcos Boullón-Magán, José Carlos Cabaleiro, Casiano Rodríguez, Francisco F. Rivera |
IPDPS | 2 |
| 2006 | An Open Source Web Service Based Platform for Heterogeneous Clusters
Francisco Almeida, Sergio Barrachina 0001, Vicente Blanco 0001, Enrique S. Quintana-Ortí, Adrián Santos |
ISPA | 3 |
| 2006 | Parallelization of GSL: The Web Service InterfaceabstractWe present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements. José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
PDP | 9 |
| 2004 | Predicting the performance of parallel programs
Vicente Blanco 0001, Jesus A. González, Coromoto León, Casiano Rodríguez, A. Marcela Printista |
Parallel Comput. | 1 |
| 2004 | Performance Prediction for Parallel Iterative Solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
J. Supercomput. | 1 |
| 2003 | From Complexity Analysis to Performance Analysis
Vicente Blanco 0001, Jesus A. González, Coromoto León, Casiano Rodríguez |
Euro-Par | 1 |
| 2003 | AVISPA: visualizing the performance prediction of parallel iterative solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 1 |
| 2001 | Modeling and improving locality for the sparse-matrix-vector product on cache memories
Dora Blanco Heras, Vicente Blanco 0001, José Carlos Cabaleiro, Francisco F. Rivera |
Future Gener. Comput. Syst. | 2 |