EDBT 2026 Demo / reviewers in the wild / expert
Francisco Almeida
dblp:00/2671 · also Francisco Almeida Rodriguez
· DBLP profile ↗
62ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-1279-9636ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 45 · 10 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorArtificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 2:4 Pruning on Edge Devices: Performance, Energy Efficiency and Accuracy
Nicolás Hernández, Pedro A. Toledo, Vicente Blanco 0001, Francisco Almeida |
Euro-Par (2) | 4 |
| 2025 | Correction: Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator (VASP)
Isidoro Nieves-Pírez, Alfonso Muñoz, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2025 | KubePipe: a container-based high-level parallelization tool for scalable machine learning pipelinesabstractAbstract As the complexity and scale of machine learning applications continue to grow, the need for efficient training methodologies becomes increasingly critical. Traditional training processes can be time-intensive, often limiting rapid development and deployment. In response to this challenge, we present KubePipe, a high-level tool that abstracts parallelism and containerization from the user, allowing non-expert users to leverage advanced parallel architectures without requiring deep knowledge of parallel computing or container orchestration. KubePipe enables the concurrent execution of multiple machine learning workflows within a Kubernetes cluster, optimizing computational resources and significantly reducing training times. By leveraging containerized environments, KubePipe ensures a high degree of modularity, scalability, and portability, making it adaptable to various machine learning frameworks and tasks. Our experimental results demonstrate substantial performance improvements when using KubePipe compared to conventional pipeline implementations. This paper explores the architecture and functionality of KubePipe, providing insights into its integration with existing machine learning systems and highlighting its potential to streamline the training process in high-performance computing environments. Daniel Suárez, Francisco Almeida, Vicente Blanco 0001, Pedro A. Toledo |
J. Supercomput. | 2 |
| 2024 | Parallel programming in mobile devices with FancyJCLabstractAbstract Mobile devices and handheld systems, such as the smartphones and tablets universally extended, are becoming increasingly powerful. Their basic hardware configuration is usually state-of-the-art heterogeneous architectures consisting of multi-core processors and some kind of accelerator such as GPUs or DSPs. Specific code adapted to the architecture is mandatory if high-performance computation is required and low-level libraries and parallelism are needed, which constitutes an important barrier for the usual developer in such devices. In this context, we propose the FancyJCL framework. It provides a high-level abstraction layer that hides implementation details and allows to develop parallel programs for mobile devices. The target platform for FancyJCL is mainly Android and Java developers due to their high market penetration. A very simple, seemingly sequential encoding results in parallel efficient OpenCL code. FancyJCL is itself based on the Fancier framework, which enables optimal memory management across memory spaces on unified memory systems. Benchmarks of FancyJCL code developed for a wide range of image processing algorithms show good performance with low development effort. Sergio Afonso, Óscar Gómez-Cárdenes, Paula Expósito, Vicente Blanco 0001, Francisco Almeida |
J. Supercomput. | 5 |
| 2024 | Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniquesabstractAbstract This document addresses some inherent problems in Machine Learning (ML), such as the high computational and energy costs associated with their implementation on IoT devices. It aims to study and analyze the performance and efficiency of quantization as an optimization method, as well as the possibility of training ML models directly on an IoT device. Quantization involves reducing the precision of model weights and activations while still maintaining acceptable levels of accuracy. Using representative networks for facial recognition developed with TensorFlow and TensorRT, Post-Training Quantization and Quantization-Aware Training are employed to reduce computational load and improve energy efficiency. The computational experience was conducted on a general-purpose computer featuring an Intel i7-1260P processor and an NVIDIA RTX 3080 graphics card used as an accelerator. Additionally, a NVIDIA Jetson AGX Orin was used as an example of an IoT device. We analyze the feasibility of training on an IoT device, the impact of quantization optimization on knowledge transfer-trained models and evaluate the differences between Post-Training Quantization and Quantization-Aware Training in such networks on different devices. Furthermore, the performance and efficiency of NVIDIA’s inference accelerator (Deep Learning Accelerator - DLA, in its 2.0 version) available at the Jetson Orin architecture are studied. We concluded that the Jetson device is capable of performing training on its own. The IoT device can achieve inference performance similar to that of the more powerful processor, thanks to the optimization process, with better energy efficiency. Post-Training Quantization has shown better performance, while Quantization-Aware Training has demonstrated higher energy efficiency. However, since the accelerator cannot execute certain layers of the models, the use of DLA worsens both the performance and efficiency results. Nicolás Hernández, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 2 |
| 2024 | Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator (VASP)abstractAbstract This work tackles the performance and energy consumption analysis of a legacy scientific application, the VASP (Vienna Ab-initio Simulation Package), an application commonly used by physicists and chemists for modeling materials at the atomic scale. Many of these scientific applications have been implemented in Fortran, where energy metrics instrumentation is not straightforward. We obtained performance figures (execution time and energy consumption) by instrumenting the source code using EML. This energy measurement library has been modified to introduce Fortran interfaces for these metrics. The analysis was carried out using different matrix algebra libraries, parallelization techniques, and hardware platforms, emphasizing on the MPI, OpenMP, and CUDA parallel implementations of the algorithms used in VASP. We employ various material specifications (atomic structures) and molecular sizes of a silicon-based crystal to create a set of benchmarks for these specifications, leading to some recommendations for final users regarding performance improvements. The proposed benchmarking technique assists the user in selecting the right combination of problem size, compilers, and parallelization options available in VASP. For a given system platform, the user will be able to determine not only the architecture to use (GPU or multicore processors), but also the appropriate library and parallelization according to the atomic structure and molecular size. Isidoro Nieves-Pírez, Alfonso Muñoz, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2024 | Comprehensive analysis of energy efficiency and performance of ARM and RISC-V SoCsabstractAbstract Over the past few years, ARM has been the dominant player in embedded systems and System-on-Chips (SoCs). With the emergence of hardware platforms based on the RISC-V architecture, a practical comparison focusing on their energy efficiency and performance is needed. In this study, our goal is to comprehensively evaluate the energy efficiency and performance of ARM and RISC-V SoCs in three different systems. We will conduct benchmark tests to measure power consumption and overall system performance. The results of our study are valuable to developers and researchers looking for the most appropriate hardware platform for energy-efficient computing applications. Our observations suggest that RISC-V Instruction Set Architecture (ISA) implementations may demonstrate lower average power consumption than ARM, but this does not automatically imply a superior performance per watt ratio for RISC-V. The primary focus of the study is to evaluate and compare these ISA implementations, aiming to identify potential areas for enhancing their energy efficiency. Furthermore, to ensure the practical applicability of our findings, we will use the Computational Fluid Dynamics software OpenFOAM. This step serves to validate the relevance of our results in real-world scenarios. It allows us to fine-tune execution parameters based on the insights gained from our initial study. By doing so, we aim not only to provide meaningful conclusions but also to investigate the transferability of our results to practical applications. Our analysis will also scrutinize the capabilities of these SoCs when handling nonsynthetic software workloads, thereby broadening the scope of our evaluation. Daniel Suárez, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 2 |
| 2020 | A Dynamic Multi-Objective Approach for Dynamic Load Balancing in Heterogeneous SystemsabstractModern standards in High Performance Computing (HPC) have started to consider energy consumption and power draw as a limiting factor. New and more complex architectures have been introduced in HPC systems to afford these new restrictions, and include coprocessors such as GPGPUs for intensive computational tasks. As systems increase in heterogeneity, workload distribution becomes a more core problem to achieve the maximum efficiency in every computational component. We present a Multi-Objective Dynamic Load Balancing (DLB) approach where several objectives can be applied to tune an application. These objectives can be dynamically exchanged during the execution of an algorithm to better adapt to the resources available in a system. We have implemented the Multi-Objective DLB together with a generic heuristic engine, designed to perform multiple strategies for DLB in iterative problems. We also present Ull Multiobjective Framework (UllMF), an open-source tool that implements the Multi-Objective generic approach. UllMF separates metric gathering, objective functions to be optimized and load balancing algorithms, and improves code portability using a simple interface to reduce the costs of new implementations. We illustrate how performance and energy consumption are improved for the implemented techniques, and analyze their quality using different DLB techniques from the literature. Alberto Cabrera Pérez, Alejandro Acosta, Francisco Almeida, Vicente Blanco 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | High-performance code optimizations for mobile devices
Sergio Afonso, Alejandro Acosta, Francisco Almeida |
J. Supercomput. | 3 |
| 2019 | A heuristic technique to improve energy efficiency with dynamic load balancing
Alberto Cabrera Pérez, Alejandro Acosta, Francisco Almeida, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2017 | Automatic Acceleration of Stencil Codes in Android Devices
Sergio Afonso, Alejandro Acosta, Francisco Almeida |
ICA3PP | 3 |
| 2016 | The particle filter algorithm: parallel implementations and performance analysis over Android mobile devicesabstractSummary The advent of emergent system on chip and multiprocessor system on chip opens a new era on the small mobile devices (smartphones, tablets, etc.) in terms of computing capabilities and applications to be addressed. Given the ability of these devices to interact with the real world through the camera, the development of efficient algorithms related to image processing and computer vision is mandatory. The particle filter algorithm is an algorithm frequently used in image and video processing; it constitutes the baseline algorithm in many applications: feature tracking, facial recognition, tracking of vehicles in traffic, video compression and so on. We propose a parallel implementation for the particle filter algorithm oriented to mobile Android devices. Three different versions of this algorithm are presented: a Java sequential implementation and two Renderscript parallel versions, an ad hoc implementation and a parallel implementation generated automatically with Paralldroid. The results obtained by the parallel versions over different Android platforms present high accuracy with a high processing rate of frame per second and a high speedup at the same time. Copyright © 2015 John Wiley & Sons, Ltd. Alejandro Acosta, Francisco Almeida |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Extending Paralldroid with object oriented annotations
Alejandro Acosta, Sergio Afonso, Francisco Almeida |
Parallel Comput. | 3 |
| 2015 | Hyperheuristics Based on Parametrized Metaheuristic SchemesabstractThe use of a unified parametrized scheme for metaheuristics facilitates the development of metaheuristics and their application. The unified scheme can also be used to implement hyperheuristics on top of parametrized metaheuristics, selecting appropriate values for the metaheuristic parameters, and consequently the metaheuristic itself. The applicability of hyperheuristics to efficiently solve computational search problems is tested with the application of local and global search methods (GRASP, Tabu Search, Genetic algorithms and Scatter Search) and their combinations to three problems: a problem of optimization of power consumption in operation of wells,the determination of the kinetic constants of a chemical reaction and the maximum diversity problem. The hyperheuristic approach provides satisfactory values for the metaheuristic parameters and consequently satisfactory metaheuristics. José-Matías Cutillas-Lozano, Domingo Giménez, Francisco Almeida |
GECCO | 3 |
| 2015 | Parallel Implementations of the Particle Filter Algorithm for Android Mobile DevicesabstractThe advent of emergent System-on-Chip (SoCs) and multiprocessor System-on-Chip (MPSocs) opens a new era on the small mobile devices (Smartphones, Tablets, ) in terms of computing capabilities and applications to be addressed. Given the ability of these devices to interact with the real world through the camera, is mandatory the development of efficient algorithms related to image processing and computer vision. We present a parallel implementation on mobile Android devices of the Particle Filter algorithm. We developed three different version of this algorithm. A Java sequential implementation and two Render script parallel versions, an ad-hoc implementation and a Paralldroid generated implementation. The results obtained by the parallel versions present an important speedup and a high accurate with a high processing rate of frame per seconds. Alejandro Acosta, Francisco Almeida |
PDP | 2 |
| 2015 | Energy Measurement Library (EML) Usage and Overhead AnalysisabstractEnergy consumption and efficiency analysis has raised as an interesting topic to address for high performance computing researchers. Exascale computers, with current capabilities, would require a huge amount of power to operate at full capacity. This has attracted the interest of many researchers in the field contributing to elaborate multiple libraries and tools to measure energy and power consumption. In spite of the efforts, there is no current standard for energy measurements. It is known that the accuracy of the measurements is highly dependent on measurement tools (power meters, performance counters, ). A standard library should offer independence between hardware and software and portability for the instrumented code between different architectures. Additionally, the overhead of the library should be low enough to achieve accurate measurements. We proposed EML (Energy Measurement Library) and now we present insights on the usage and the overhead of the library, so that it could contribute to the establishment of an standard for energy measurement. Alberto Cabrera Pérez, Francisco Almeida, Javier Arteaga, Vicente Blanco 0001 |
PDP | 2 |
| 2014 | Performance Analysis of Paralldroid Generated ProgramsabstractThe advent of emergent System-on-Chip (SoCs) and multiprocessor System-on-Chip (MPSocs) opens a new era on the small mobile devices (Smartphones, Tablets, ...) in terms of computing capabilities and applications to be addressed. The efficient use of such devices, including the parallel power, is still a challenge for general purpose programmers due to the very high learning curve demanding very specific knowledge of the devices. While some efforts are currently being made, mainly in the scientific scope, the scenario is still quite far from being the desirable for non-scientific applications where very few of them take advantage of the parallel capabilities of the devices. We develop a performance analysis in several SoCs using Paralldroid. Paralldroid (Framework for Parallelism in Android), is a parallel development framework oriented to general purpose programmers for standard mobile devices. Paralldroid presents a programming model that unifies the different programming models of Android. The user just implements a Java application and introduces a set of Paralldroid annotations in the sections of code to be optimized. The Paralldroid system automatically generates the native C, OpenCL or Renderscript code for the annotated section. The Paralldroid transformation model involves source-to-source transformations and skeletal programming. Alejandro Acosta, Francisco Almeida |
PDP | 2 |
| 2014 | Android $$^\mathrm{TM}$$ TM development and performance analysis
Alejandro Acosta, Francisco Almeida |
J. Supercomput. | 2 |
| 2013 | Analytical Energy Models for MPI Communications on a Sandy-Bridge ArchitectureabstractComplexity in parallel computation for current HPC systems has being modeled in order to estimate time and energy consumption, which have been conduced to evaluate real costs. Communication analytical models for time and energy comsumption are an attractive issue as part of this cost analysis in parallel computation. In this paper, we present time and energy analytical models for MPI communications using OSU micro benchmarks under Sandy Bridge-EP computer architecture. RAPL (Running Average Power Limit) interface has been setup to provide mechanisms to enforce real energy consumption measurements for OSU micro benchmarks on this architecture. Our applied methodology defines an experimental framework based on several core configurations and different transmission rates. The RAPL interface allows us to obtain different energy values for the OSU benchmark executions. The measures obtained are used to obtain time and energy consumption parameters of our models, establishing average errors below 1% for almost all configurations in real executions under this computer architecture for an ad-hoc model. Alternative and simpler time and energy models are also proposed with errors up to 3.1% for the worst case. Francisco Almeida, Vicente Blanco 0001, Isidro Gonzalez, Alberto Cabrera Pérez, Domingo Giménez |
ICPP | 1 |
| 2013 | Analytical Modeling of the Energy Consumption for the High Performance LinpackabstractComparable to time performance models, it is now possible to estimate performance based upon energy consumption for HPC systems. The predictive ability of the analytical modeling is an interesting feature that motivates us to approach this methodology for the case of energy consumption. In this paper, we present an analytical model for predicting the energy consumption for the High Performance Linpack (HPL). The derived model can be used to know in advance the energy consumed by the HPL over a target architecture, and can be integrated into the schedulers of operating systems or queue managers. We established an experimental setup using a standard metered PDU that allowed us to measure the energy consumption for the HPL benchmark on our cluster. With the monitoring system in place, we can obtain the architectural and algorithmic parameters associated for both performance and energy analytical models. Also this has made possible watts and gflops-per-watt prediction when we execute Linpack executions with concrete algorithm parameters in our cluster. Alberto Cabrera Pérez, Francisco Almeida, Vicente Blanco 0001, Domingo Giménez |
PDP | 2 |
| 2013 | High-level specifications for automatically generating parallel codeabstractSUMMARY The arrival of multicore systems, along with the speed‐up potential available in graphics processing units, has given us unprecedented low‐cost computing power. These systems address some of the known architecture problems but at the expense of considerably increased programming complexity. Heterogeneity, at both the architectural and programming levels, poses a great challenge to programmers. Many proposals have been put forth to facilitate the job of programmers. Leaving aside proposals based on the development of new programming languages because of the effort this represents for the user (effort to learn and reuse code), the remaining proposals are based on transforming sequential code into parallel code, or on transforming parallel code designed for one architecture into parallel code designed for another. A different approach relies on the use of skeletons. The programmer has available set of parallel standards that comprise the basis for developing parallel code while programming sequential code. In this context, we propose a methodology for developing an automatic source‐to‐source transformation in a specific domain. This methodology is instantiated in a framework aimed at solving dynamic programming problems. Using this framework, the final user (a physician, mathematician, biologist, etc.) can express her problem using an equation in Latex, and the system will automatically generate the optimal parallel code for homogeneous or heterogeneous architectures. This approach allows for great portability toward these new emerging architectures and for great productivity, as evidenced by the computational results.Copyright © 2012 John Wiley & Sons, Ltd. Alejandro Acosta, Francisco Almeida, Ignacio Peláez |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Skeletal based programming for dynamic programming on MultiGPU systems
Alejandro Acosta, Francisco Almeida |
J. Supercomput. | 2 |
| 2013 | From latex specifications to parallel codes
Alejandro Acosta, Francisco Almeida, Ignacio Peláez |
J. Supercomput. | 2 |
| 2013 | Modeling energy consumption for master-slave applications
Francisco Almeida, Vicente Blanco 0001, Alberto Cabrera Pérez, J. Ruiz |
J. Supercomput. | 1 |
| 2013 | High performance computing tools in science and engineering
Francisco Almeida, Jesús Vigo-Aguiar |
J. Supercomput. | 1 |
| 2013 | Parameterized Schemes of Metaheuristics: Basic Ideas and Applications With Genetic Algorithms, Scatter Search, and GRASPabstractSome optimization problems can be tackled only with metaheuristic methods, and to obtain a satisfactory metaheuristic, it is necessary to develop and experiment with various methods and to tune them for each particular problem. The use of a unified scheme for metaheuristics facilitates the development of metaheuristics by reutilizing the basic functions. In our proposal, the unified scheme is improved by adding transitional parameters. Those parameters are included in each of the functions, in such a way that different values of the parameters provide different metaheuristics or combinations of metaheuristics. Thus, the unified parameterized scheme eases the development of metaheuristics and their application. In this paper, we expose the basic ideas of the parameterization of metaheuristics. This methodology is tested with the application of local and global search methods (greedy randomized adaptive search procedure [GRASP], genetic algorithms, and scatter search), and their combinations, to three scientific problems: obtaining satisfactory simultaneous equation models from a set of values of the variables, a task-to-processor assignment problem with independent tasks and memory constrains, and thep-hub median location-allocation problem. Francisco Almeida, Domingo Giménez, Jose-Juan López-Espín, Melquíades Pérez Pérez |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2012 | OpenCF-R: R in the Cloud
Juan Carlos Castillo, Francisco Almeida, Vicente Blanco 0001, Adrián Santos |
CLOSER | 2 |
| 2012 | Towards the Dynamic Load Balancing on Heterogeneous Multi-GPU SystemsabstractThe advent of multicore systems, joined to the potential acceleration of the graphics processing units, alleviates some well known important architectural problems at the expense of a considerable increment of the programmability wall. The heterogeneity, both at architectural and programming level at the same time, raises the programming difficulties. Adapting existing code and libraries to such systems is a fundamental problem. The performance of this code is affected by the large interdependence between the code and the parallel architecture. We have developed a dynamic load balancing library that allows parallel code to be adapted to a wide variety of heterogeneous systems. The overhead introduced by our system is minimal and the cost to the programmer negligible. This system has been applied to solve load imbalance problems appearing in homogeneous and heterogeneous multi-GPU platforms. As case studies, we consider matrix multiply and resource allocation problems, in different heterogeneous scenarios in multi-GPU systems. The unbalanced nature of these algorithms and heterogeneous systems allowed us to test the success of our load balancing approach. Alejandro Acosta, Vicente Blanco 0001, Francisco Almeida |
ISPA | 3 |
| 2012 | A framework for the application of metaheuristics to tasks-to-processors assignation problems
Francisco Almeida, Javier Cuenca 0001, Domingo Giménez, Antonio Llanes, Juan-Pedro Martínez-Gallar |
J. Supercomput. | 1 |
| 2011 | PETransWS: Web Service Computing Platform for Logistics and TransportationabstractIn large organizations and small firms in transportation, there is a growing need to use and analyze spatial data. Transportation system analysis and planning as well as mobility studies frequently use Geographic Information Systems (GIS). In this paper we propose the development of a web services platform dedicated to transportation and logistics. Taking advantage of the web services development framework PyOpenCF we integrate in the same environment services oriented to geolocalization, logistical optimization, etc. We develop a PyOpenCF client whose graphic interface allows processes with spatial data to be launched and provides a visualization of the results. Francisco Almeida, Vicente Blanco 0001, Julio Brito, Andres Crespo, José A. Moreno-Pérez, Adrián Santos |
PDP | 1 |
| 2011 | A parameterized shared-memory scheme for parameterized metaheuristics
Francisco Almeida, Domingo Giménez, Jose-Juan López-Espín |
J. Supercomput. | 1 |
| 2011 | Adaptive load balancing of iterative computation on heterogeneous nondedicated systems
José Antonio Martínez, Francisco Almeida, Ester M. Garzón, Alejandro Acosta, Vicente Blanco 0001 |
J. Supercomput. | 2 |
| 2011 | Web services based scheduling in OpenCF
Adrián Santos, Francisco Almeida, Vicente Blanco 0001, Juan Carlos Castillo |
J. Supercomput. | 2 |
| 2010 | CoEDApplets - Collaborating in the Development of Teaching-oriented Applets
Francisco Almeida, Vicente Blanco 0001, J. Regalado, Adrián Santos |
CSEDU (1) | 1 |
| 2009 | EDApplets: A Web Tool for Teaching Data Structures and Algorithmic Techniques
Francisco Almeida, Vicente Blanco 0001, J. Regalado, Adrián Santos |
CSEDU (2) | 1 |
| 2009 | Using Web Services for Performance Monitoring and SchedulingabstractThe adoption of Web Service standards provides us with an increased level of manageability, extensibility and interoperability between loosely coupled services.The adoption of Web Services technologies atop sites for performance monitoring and scheduling will improve the efficient use of the computational resources. Web Services provide the ability to decompose HPC resources and functionality into a set of discoverable and loosely coupled services, which are capable of interaction in heterogeneous environments. At the same time, Web Services can address many of the interoperability issues that can be encountered in large scale systems.End users can access to these services to decide which system will be the most suitable for their needs. Other tools like schedulers can use the resources available as services to optimize the HPC resources and minimize jobs waiting time. Adrián Santos, Francisco Almeida, Vicente Blanco 0001, David Diez, Jonás Regueira, Esaú Sicilia |
PDP | 2 |
| 2009 | IDEWEP: Web service for astronomical parallel image deconvolution
Francisco Almeida, Vicente Blanco 0001, Carlos Delgado, Francisco de Sande, Adrián Santos |
J. Netw. Comput. Appl. | 1 |
| 2009 | Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
J. Supercomput. | 2 |
| 2008 | Towards the Automatic Service Generation and Scheduling in the OpenCF ProjectabstractWeb services-based technologies have emerged as a technological alternative for computational web portals. Facilitating access to distributed resources through web interfaces while simultaneously ensuring security is one of the main goals in most of the currently existing manifold tools and frameworks. OpenCF, the open source computational framework that we have developed, shares these objectives and adds others, like enforced portability, genericity, modularity and compatibility with a wide range of high performance computing systems. Adrián Santos, Francisco Almeida, Vicente Blanco 0001, David Diez, Jonás Regueira, Esaú Sicilia |
CISIS | 2 |
| 2008 | Topic 2: Performance Prediction and Evaluation
Francisco Almeida, Michael Gerndt, Adolfy Hoisie, Martin Schulz 0001 |
Euro-Par | 1 |
| 2007 | Lightweight Web Services for High Performace Computing
Adrián Santos, Francisco Almeida, Vicente Blanco 0001 |
ECSA | 2 |
| 2007 | Performance analysis for clusters of symmetric multiprocessorsabstractIn this article we analyze and model the performance of a symmetrical multiprocessor cluster. The obtained model takes into account the heterogeneity of the architecture's communications, which allows for a better adjustment and predictive ability of the performance. We used diverse applications with different parallel schemes to check the quality of the adjustments. The adjustment of the theoretical model and the experimental results is clearly improved when we take into account the combination of various mechanisms of communication Francisco Almeida, Juan A. Gómez, José M. Badía |
PDP | 1 |
| 2006 | An Open Source Web Service Based Platform for Heterogeneous Clusters
Francisco Almeida, Sergio Barrachina 0001, Vicente Blanco 0001, Enrique S. Quintana-Ortí, Adrián Santos |
ISPA | 1 |
| 2006 | From XML Specifications to Parallel Programs
Ignacio Peláez, Francisco Almeida, Daniel González |
ISPA | 2 |
| 2006 | Parallelization of GSL: The Web Service InterfaceabstractWe present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements. José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
PDP | 8 |
| 2006 | The master-slave paradigm on heterogeneous systems: A dynamic programming approach for the optimal mapping
Francisco Almeida, Daniel González, Luz Marina Moreno |
J. Syst. Archit. | 1 |
| 2006 | Efficient parallel LAN/WAN algorithms for optimization. The mallba project
Enrique Alba 0001, Francisco Almeida, Maria J. Blesa, Carlos Cotta, Manuel Díaz, Isabel Dorta, Joaquim Gabarró, Coromoto León, Gabriel Luque, Jordi Petit |
Parallel Comput. | 2 |
| 2005 | Pipelines on heterogeneous systems: models and toolsabstractWe study the performance of pipeline algorithms in heterogeneous networks. The concept of heterogeneity is not only restricted to the differences in computational power of the nodes, but also refers to the network capabilities. We develop a skeleton tool that allows us an efficient block-cyclic mapping of pipelines on heterogeneous systems. The tool supports pipelines with a number of stages much larger than the number of physical processors available. We derive an analytical formula that allows us to predict the performance of pipelines in heterogeneous systems. According to the analytical complexity formula, numerical strategies to solve the optimal mapping problem are proposed. The computational results prove the accuracy of the predictions and effectiveness of the approach. Copyright © 2005 John Wiley & Sons, Ltd. Francisco Almeida, Daniel González, Luz Marina Moreno, Casiano Rodríguez |
Concurr. Pract. Exp. | 1 |
| 2004 | On the Use of Path Relinking for the rho-Hub Median Problem
Melquíades Pérez Pérez, Francisco Almeida, J. Marcos Moreno-Vega |
EvoCOP | 2 |
| 2003 | Towards the automatic optimal mapping of pipeline algorithms
Daniel González, Francisco Almeida, Luz Marina Moreno, Casiano Rodríguez |
Parallel Comput. | 2 |
| 2002 | MALLBA: A Library of Skeletons for Combinatorial Optimisation (Research Note)
Enrique Alba 0001, Francisco Almeida, Maria J. Blesa, J. Cabeza, Carlos Cotta, Manuel Díaz, Isabel Dorta, Joaquim Gabarró, Coromoto León, J. Luna, Luz Marina Moreno, C. Pablos, Jordi Petit, Angélica Rojas, Fatos Xhafa |
Euro-Par | 2 |
| 2002 | Optimal tiling for the RNA base pairing problemabstractDynamic programming is an important combinatorial optimization technique that has been widely used in various fields such as control theory, operations research, computational biology and computer science. Many authors have described parallel dynamic programming algorithms for the family of multistage problems. More scarce is the literature for the more general class of problems where dependences appear between non-consecutive stages. Among the important problems falling in this class is the RNA base pairing problem. In this study we propose a new parallel scheme for a large class of recurrences with triangular iteration space and nonuniform dependences that includes the RNA base pairing problem. We derive two different instances of this scheme that correspond to an horizontal and a vertical traverse of the iteration domain. We develop and extend the tiling approach for this particular class. We formulate and analytically solve the optimization problem determining the tile size that minimizes the total execution time of the tiled program on a distributed memory parallel machine. Our analyze is based on the BSP model, which assures the portability of the obtained results. The computational experiments carried out on the CRAY T3E behave according to the predictions of our theoretical model. Francisco Almeida, Rumen Andonov, Daniel González, Luz Marina Moreno, Vincent Poirriez, Casiano Rodríguez |
SPAA | 1 |
| 2001 | The Tuning Problem on Pipelines
Luz Marina Moreno, Francisco Almeida, Daniel González, Casiano Rodríguez |
Euro-Par | 2 |
| 2000 | Optimal Mapping of Pipeline Algorithms (Research Note)
Daniel González, Francisco Almeida, Luz Marina Moreno, Casiano Rodríguez |
Euro-Par | 2 |
| 2000 | From the Theory to the Tools: Parallel Dynamic ProgrammingabstractDynamic programming is an important paradigm that has been widely used to solve problems in various areas such as control theory, operation research, biology and computer science. We generalize the finite automaton formal model for dynamic programming deriving pipeline parallel algorithms. The optimality of these algorithms is established for the new class of non-decreasing finite automata. As an intermediate step for the construction of a skeleton for the automatic parallelization of dynamic programming, we have developed a tool for the implementation of pipeline algorithms. The tool maps the processes in the pipeline in the target architecture following a mix of block and cyclic policies adapted to the grain of the machine. Based on the former tool, the automatic parallelization of dynamic programming is straightforward. The use of the model and its associated tools is illustrated with the Single Resource Allocation Problem. The performance and portability of these tools is compared with specific ‘hand made’ code written by experienced programmers. The experimental results on distributed memory and shared distributed memory architectures prove the scalability of the proposed paradigm and its associated tools. Copyright © 2000 John Wiley & Sons, Ltd. Daniel González, Francisco Almeida, José Luis Roda García, Casiano Rodríguez |
Concurr. Pract. Exp. | 2 |
| 2000 | Parallel dynamic programming and automata theory
Daniel González-Morales, Francisco Almeida, Casiano Rodríguez, José Luis Roda García, I. Coloma, A. Delgado |
Parallel Comput. | 2 |
| 2000 | A new parallel model for the analysis of asynchronous algorithms
Casiano Rodríguez, José Luis Roda García, Francisco de Sande, Daniel González-Morales, Francisco Almeida |
Parallel Comput. | 5 |
| 1999 | A Skeleton for Parallel Dynamic Programming
Daniel González-Morales, Francisco Almeida, José Luis Roda García, Casiano Rodríguez |
Euro-Par | 2 |
| 1999 | Predicting the execution time of message passing modelsabstractRecent publications prove that runtime systems oriented to the Bulk Synchronous Parallel Model usually achieve remarkable accuracy in their predictions. That accuracy can be seen in the capacity of the software for packing the messages generated during the superstep and their capability to find a rearrangement of the messages sent at the end of the superstep. Unfortunately, barrier synchronisation imposes some limits both in the range of available algorithms and in their performance. The asynchronous nature of many MPI/PVM programs makes their expression difficult or infeasible using a BSP oriented library. Through the generalisation of the concept of superstep we propose two extensions of the BSP model: the BSP Without Barriers (BSPWB) and the Message Passing Machine (MPM) models. These new models are oriented to MPI/PVM parallel programming. The parameters of the models and their quality are evaluated on four standard parallel platforms. The use of these BSP extensions is illustrated using the Fast Fourier Transform and the Parallel Sorting by Regular Sampling algorithms. Copyright © 1999 John Wiley & Sons, Ltd. José Luis Roda García, Casiano Rodríguez, Daniel González-Morales, Francisco Almeida |
Concurr. Pract. Exp. | 4 |
| 1998 | h-Relation Models for Current Standard Parallel Platforms
Casiano Rodríguez, José Luis Roda García, Daniel González-Morales, Francisco Almeida |
Euro-Par | 4 |
| 1996 | A parallel algorithm for the integer knapsack problemabstractA sequential algorithm with complexity O(M2+n) for the integer knapsack problem is presented. M is the capacity of the knapsack, and n the number of objects. The algorithm admits an efficient parallelization on a p-processor ring machine. The corresponding parallel algorithm is O(M2/p+n). The parallel algorithm is compared with a version of the well-known Lee algorithm adapted to the integer knapsack problem. Computational results on both a local area network and a transputer are reported. Domingo Morales, José Luis Roda García, Casiano Rodríguez, Francisco Almeida |
Concurr. Pract. Exp. | 4 |
| 1995 | Integral Knapsack Problems: Parallel Algorithms and Their Implementations on Distributed SystemsabstractArticle Integral knapsack problems: parallel algorithms and their implementations on distributed systems Share on Authors: D. Morales Dpto. Estadística,Investigación Operativa y Computación, Universidad de La Laguna Dpto. Estadística,Investigación Operativa y Computación, Universidad de La LagunaView Profile , J. Roda Dpto. Estadística,Investigación Operativa y Computación, Universidad de La Laguna Dpto. Estadística,Investigación Operativa y Computación, Universidad de La LagunaView Profile , F. Almeida Dpto. Estadística,Investigación Operativa y Computación, Universidad de La Laguna Dpto. Estadística,Investigación Operativa y Computación, Universidad de La LagunaView Profile , C. Rodríguez Dpto. Estadística,Investigación Operativa y Computación, Universidad de La Laguna Dpto. Estadística,Investigación Operativa y Computación, Universidad de La LagunaView Profile , F. García Dpto. Estadística,Investigación Operativa y Computación, Universidad de La Laguna Dpto. Estadística,Investigación Operativa y Computación, Universidad de La LagunaView Profile Authors Info & Claims ICS '95: Proceedings of the 9th international conference on SupercomputingJuly 1995 Pages 218–226https://doi.org/10.1145/224538.224564Online:03 July 1995Publication History 10citation522DownloadsMetricsTotal Citations10Total Downloads522Last 12 Months8Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Daniel González-Morales, José Luis Roda García, Francisco Almeida, Casiano Rodríguez |
International Conference on Supercomputing | 3 |