VLDB 2026 Research / reviewers in the wild / expert
Dolores Rexachs
dblp:93/1849 · also Dolores Isabel Rexachs del Rosario
· DBLP profile ↗
51ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-5500-850XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Application of a sampling and clustering-based heuristic search algorithm to find an efficient staff configuration in an emergency departmentabstractEmergency Departments (EDs) are among the most complex areas in healthcare, requiring immediate medical attention for acute and urgent conditions. Optimizing staff configurations to reduce patient Length of Stay (LoS) and improve operational efficiency poses a significant challenge due to the combinatorial and high-dimensional nature of the problem. To identify the most effective staff configuration, we propose a heuristic optimization strategy that is based on the Montecarlo Clustering Search Algorithm (MCSA), which efficiently explores the multidimensional solution space. MCSA leverages an agent-based simulation (ABM) model that evaluates each proposed staff configuration under realistic operational conditions, providing Key Performance Indicator (KPI) feedback values related to each proposed staff configuration. Through this strategy, we explore staff configurations capable of handling patient volumes with varying acuity levels in an ED to optimize the LoS KPI. Results demonstrate that our methodology is capable to find a solution as a staff configuration that reduces LoS compared to a baseline, offering a computationally efficient and practical tool for decision-makers. We identified solutions by exploring less than 1% of the total search space, demonstrating the efficiency of the proposed approach in addressing complex optimization problems. This approach supports informed planning in healthcare environments while maintaining system feasibility and scalability. Maria Harita, Alvaro Wong, Dolores Rexachs, Emilio Luque, Eva Bruballa, Francisco Epelde |
Expert Syst. Appl. | 3 |
| 2025 | Parallel I/O analysis in distributed deep learning applications on high-performance computingabstractAbstract Distributed deep learning (DDL) applications generate heavy input/output (I/O) workloads that can create bottlenecks in high-performance computing (HPC) systems. Their optimal I/O configuration depends on factors such as access patterns, storage hardware, dataset size, and execution scale. This study proposes a systematic methodology for characterizing and optimizing I/O behavior in DDL applications, represented through the deep learning I/O benchmark (DLIO), and validated with the real DeepGalaxy application. We evaluate access modes, file formats, and Lustre file system configurations, demonstrating that stripe counts optimized for the access pattern and application scale can reduce I/O and execution times, achieving up to 18 GiB/s of bandwidth and a 5X increase in IOPS. HDF5 provides balanced performance, while TFRecord stands out in bandwidth-intensive scenarios. Shared access minimizes contention and improves scalability in multi-node executions. The results are consolidated into configuration guidelines that offer practical recommendations for practitioners to tune DDL applications for efficient execution in HPC environments. Edixon Párraga, Betzabeth León, Sandra Méndez, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 4 |
| 2024 | Exploring energy saving opportunities in fault tolerant HPC systems
Marina Morán, Javier Aldo Balladini, Dolores Rexachs, Enzo Rucci |
J. Parallel Distributed Comput. | 3 |
| 2023 | A computational methodology applied to optimize the performance of a river model under uncertainty conditions
Adriana Gaudiani, Alvaro Wong, Emilio Luque, Dolores Rexachs |
J. Supercomput. | 4 |
| 2022 | A model of checkpoint behavior for applications that have I/OabstractAbstract Due to the increase and complexity of computer systems, reducing the overhead of fault tolerance techniques has become important in recent years. One technique in fault tolerance is checkpointing, which saves a snapshot with the information that has been computed up to a specific moment, suspending the execution of the application, consuming I/O resources and network bandwidth. Characterizing the files that are generated when performing the checkpoint of a parallel application is useful to determine the resources consumed and their impact on the I/O system. It is also important to characterize the application that performs checkpoints, and one of these characteristics is whether the application does I/O. In this paper, we present a model of checkpoint behavior for parallel applications that performs I/O; this depends on the application and on other factors such as the number of processes, the mapping of processes and the type of I/O used. These characteristics will also influence scalability, the resources consumed and their impact on the IO system. Our model describes the behavior of the checkpoint size based on the characteristics of the system and the type (or model) of I/O used, such as the number I/O aggregator processes, the buffering size utilized by the two-phase I/O optimization technique and components of collective file I/O operations. The BT benchmark and FLASH I/O are analyzed under different configurations of aggregator processes and buffer size to explain our approach. The model can be useful when selecting what type of checkpoint configuration is more appropriate according to the applications’ characteristics and resources available. Thus, the user will be able to know how much storage space the checkpoint consumes and how much the application consumes, in order to establish policies that help improve the distribution of resources. Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 4 |
| 2022 | Correction to: A model of checkpoint behavior for applications that have I/O
Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 4 |
| 2022 | Scalable performance analysis method for SPMD applicationsabstractAbstract The analysis of parallel scientific applications allows us to understand their computational and communication behavior. One way of obtaining performance information is through performance tools. One such tool is parallel application signatures for performance prediction (PAS2P), based on parallel application repeatability, focusing on performance analysis and prediction. The same resources that execute the parallel application are used to perform its analysis, creating a machine independent model of the application and identifying its common patterns. However, the analysis is costly in terms of execution time due to the high number of synchronization communications performed by PAS2P, degrading performance as the number of processes increases. To solve this problem, we propose a model that reduces data dependency between processes, reducing the number of communications performed by PAS2P in the analysis stage and taking advantage of the characteristics of single program, multiple sata applications. Our analysis proposal allows us to decrease the analysis time by 29 times when the application scales to 256 processes, while keeping error levels below 11% in the runtime prediction. It is important to mention that the analysis time is not considerably affected by increasing the number of application processes. Felipe Tirado, Alvaro Wong, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 3 |
| 2021 | Analysis of parallel application checkpoint storage for system configuration
Betzabeth León, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 3 |
| 2021 | Middleware to Manage Fault Tolerance Using Semi-Coordinated CheckpointsabstractCompute node failures are becoming a normal event for many long-running and scalable MPI applications. Keeping within the MPI standards and applying some of the methods developed so far in terms of fault tolerance, we developed a methodology that allows applications to tolerate failures through the creation of semi-coordinated checkpoints within the RADIC architecture. To do this, we developed the ULSC2-RADIC middleware that divides the application into independent MPI worlds where each MPI world would correspond to a compute node and make use of the DMTCP checkpoint library in a semi-coordinated environment. We performed experimental results using scientific applications and the NAS Parallel Benchmarks to assess the overhead and also the functionality in case of a node failure. We evaluated the computational cost of the semi-coordinated checkpoints compared with the coordinated checkpoints. Alvaro Wong, Elisa Heymann, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Soft errors detection and automatic recovery based on replication combined with different levels of checkpointing
Diego Montezanti, Enzo Rucci, Armando De Giusti, Marcelo R. Naiouf, Dolores Rexachs, Emilio Luque |
Future Gener. Comput. Syst. | 5 |
| 2020 | A Method for Projections of the Emergency Department Behaviour by Non-Communicable Diseases From 2019 to 2039abstractIn this paper, a new method for prediction of future performance and demand on emergency department (ED) in Spain is presented. Increased life expediency and population aging in Spain, along with their corresponding health conditions such as non-communicable diseases (NCDs), have been suggested to contribute to higher demands on ED. These lead to inferior performance of the department and cause longer ED length of stay (LoS). Prediction and quantification of behavior of ED is, however, challenging as ED is one of the most complex parts of hospitals. Using detailed computational approaches integrated with clinical data behavior of Spain's ED in future years was predicted. First, statistical models were developed to predict how the population and age distribution of patients with non-communicable diseases change in Spain in future years. Then, an agent-based modeling approach was used for simulation of the emergency department to predict impacts of the changes in population and age distribution of patients with NCDs on the performance of ED, reflected in ED LoS, between years 2019 and 2039. Results from different projection scenarios indicated that Spain would experience a continuous increase in total ED LoS from 5.7 million hours in 2019 to 6.2 million hours in 2039 if same human and physical resources, as well as same ED configuration, are used. The results from this study can provide health care provider with quantitative information on required staff and physical resources in the future and allow health care policymakers to improve modifiable factors contributing to the demand and performance of ED. Elham Shojaei, Alvaro Wong, Dolores Rexachs, Francisco Epelde, Emilio Luque |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | RaaS: Resilience as a ServiceabstractCloud computing is continuously increasing its popularity as key features such as scalability, pay-per-use and availability continue to evolve. It is also becoming a competitive platform for running high performance computing (HPC) and parallel applications due to the increasing performance of virtualized, highly-available instances. However, migrating HPC applications to cloud still requires native fault-tolerant solutions to fully leverage cloud features and maximize the resource utilization at the best cost - particularly for long-running parallel applications where faults can cause invalid states or data loss. This requires re-executing applications which increases completion time and cost. We propose Resilience as a Service (RaaS), a fault tolerant framework for HPC applications running in cloud. In this paper RADIC architecture (Redundant Array of Distributed Independent Fault Tolerance Controllers) is used to provide clouds with a highly available, distributed and scalable fault-tolerant service. The paper explores how traditional HPC protection and recovery mechanisms must be redesigned to natively leverage cloud properties and its multiple alternatives for implementing rollback recovery protocols using virtual machines, containers, object and block storage or database services. Results show that RaaS restores and completes the application execution using available resources while reducing overhead up to 8% for different fault-tolerant configuration alternatives. Jorge Villamayor, Dolores Rexachs, Emilio Luque, Diego Lugones |
CCGrid | 2 |
| 2018 | P3S: A Methodology to Analyze and Predict Application ScalabilityabstractExecuting message-passing parallel applications on a large number of resources in an efficient way is not a trivial task. Due to the complex interaction between the parallel applications and the HPC system, many applications may suffer performance inefficiencies when they scale. To achieve an efficient use of these large-scale systems using thousands of cores, a point to consider before executing an application is to know its behavior in the system. In this work, we propose a novel methodology called P3S (Prediction of Parallel Program Scalability), which allows us to analyze and predict the scalability of message-passing applications on a given system. The methodology strives to use a bounded analysis time, and a reduced set of resources to predict the application behavior for large-scale. The experimental validation proves that the P3S is able to predict the application scalability with an average accuracy greater than 95 percent using a reduced set of resources. Javier Panadero, Alvaro Wong, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Analyzing the Parallel I/O Severity of MPI ApplicationsabstractPerformance evaluation of parallel applications plays an important role in High Performance Computing (HPC). This is also applied to parallel I/O performance evaluation, which requires understanding the I/O pattern of the application and having knowledge about the performance capacity of the HPC I/O system. In this paper, we present a methodology to evaluate the I/O performance of parallel applications based on the I/O severity degree. We define the I/O severity concept taking into account the I/O requirements of a parallel application, the mapping of I/O processes and the configuration of the I/O subsystem. Requirements are expressed in units denominated I/O phases, which are defined using the temporal and spatial pattern of different files of the application. Our approach is applied to the I/O kernels of scientific applications such as S3DIO, FLASH-IO and BT-IO on the SuperMUC supercomputer. Experimental results show that our methodology allows us to identify if a parallel application is limited by the I/O subsystem and identifying possible root causes of the I/O problems. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CCGrid | 2 |
| 2017 | An approach for an efficient execution of SPMD applications on Multi-core environments
Ronal Muresano, Hugo Meyer, Dolores Rexachs, Emilio Luque |
Future Gener. Comput. Syst. | 3 |
| 2017 | Hybrid Message Pessimistic Logging. Improving current pessimistic message logging protocols
Hugo Meyer, Ronal Muresano, Marcela Castro-León, Dolores Rexachs, Emilio Luque |
J. Parallel Distributed Comput. | 4 |
| 2015 | Fault tolerance at system level based on RADIC architectureabstractThe increasing failure rate in High Performance Computing encourages the investigation of fault tolerance mechanisms to guarantee the execution of an application in spite of node faults. This paper presents an automatic and scalable fault tolerant model designed to be transparent for applications and for message passing libraries. The model consists of detecting failures in the communication socket caused by a faulty node. In those cases, the affected processes are recovered in a healthy node and the connections are reestablished without losing data. The Redundant Array of Distributed Independent Controllers architecture proposes a decentralized model for all the tasks required in a fault tolerance system: protection, detection, recovery and masking. Decentralized algorithms allow the application to scale, which is a key property for current HPC system. Three different rollback recovery protocols are defined and discussed with the aim of offering alternatives to reduce overhead when multicore systems are used. A prototype has been implemented to carry out an exhaustive experimental evaluation through Master/Worker and Single Program Multiple Data execution models. Multiple workloads and an increasing number of processes have been taken into account to compare the above mentioned protocols. The executions take place in two multicore Linux clusters with different socket communications libraries. Marcela Castro-León, Hugo Meyer, Dolores Rexachs, Emilio Luque |
J. Parallel Distributed Comput. | 3 |
| 2015 | Parallel Application Signature for Performance Analysis and PredictionabstractPredicting the performance of parallel scientific applications is becoming increasingly complex. Our goal was to characterize the behavior of message-passing applications on different target machines. To achieve this goal, we developed a method called parallel application signature for performance prediction (PAS2P), which strives to describe an application based on its behavior. Based on the application's message-passing activity, we identified and extracted representative phases, with which we created a parallel application signature that enabled us to predict the application's performance. We experimented with using different scientific applications on different clusters. We were able to predict execution times with an average accuracy greater than 97 percent. Alvaro Wong, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | "Analysis of scalability: A parallel application model approach"abstractIn this paper we propose a methodology that allows us to predict the application scalability behavior in a specific system, providing information to select the most appropriate resources to run the application. We explain the general methodology, focusing on the presentation of a novel method to model the logical application trace for a large number of processes. This method is based on the projection of a set of executions of the application signature for a small number of processes. The generated traces are validated by comparing them with the real traces obtained with PAS2P tool. We present the experimental validation for the BT Nas Parallel Benchmark. The signatures for 16, 36, 64, 81 and 100 processes were executed and used to model and project the logical trace for 1024 processes. The results obtained show the accuracy of the method. The communication pattern was predicted without error, while the predicted error is less than 10% for the communication volume and less than 5% for the number of instructions. Javier Panadero, Alvaro Wong, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2012 | A Fault-Tolerant Cache Service for Web Search Engines: RADIC Evaluation
Carlos Gómez-Pantoja, Dolores Rexachs, Mauricio Marín, Emilio Luque |
Euro-Par | 2 |
| 2012 | Transparent Fault Tolerance Solution at Socket Level Based on RADICabstractWe present a transparent middleware for fault tolerance based on RADIC, Redundant Array of Distributed Independent Controllers, a transparent and scalable fault tolerant architecture for parallel applications. It is designed at socket level and makes a secure tunnel connection able to keep the tcp sessions established by the application in spite of node failures. It is located at user level and is independent of the message-passing communication library being used. The protection gets through uncoordinated checkpoints and log message and the recovery are done in a automatic way so in case of node failures there is no need of intervention of the administrator. We have tested our fault tolerance system by executing a master-worker (M/W) and SPMD applications that follow different communication patterns. Marcela Castro-León, Dolores Rexachs, Emilio Luque |
ISPA | 2 |
| 2012 | A Fault-Tolerant Cache Service for Web Search EnginesabstractLarge Web search engines are constructed as a collection of services that are deployed on dedicated clusters of distributed-memory processors. In particular, efficient user query throughput heavily relies on using result cache services devoted to maintaining the answers to most frequent queries. Load balancing and fault tolerance are critical to this service. This paper proposes the design of a result cache service based on consistent hashing and a strategy for enabling fault tolerance. Performance evaluation is performed by using actual queries from a commercial search engine. The results show that the proposed cache service outperforms baseline approaches, decreases the average query response time, increases query throughput and efficiently recovers performance after processor failures. Carlos Gómez-Pantoja, Veronica Gil-Costa, Dolores Rexachs, Mauricio Marín, Emilio Luque |
ISPA | 3 |
| 2011 | Impact of parallel programming models and CPUs clock frequency on energy consumption of HPC systemsabstractEnergy consumption has become one of the greatest challenges in the field of high performance computing (HPC). The energy cost produced by supercomputers during the lifetime of the installation is similar to acquisition. Thus, besides its impact on the environment, energy is a limiting factor for the HPC. Our research aims to reduce the energy consumption of computer systems to run parallel HPC applications. In this article we analyse the possible influence on the energy consumption of parallel programming paradigms of shared memory (OpenMP) and message passing (MPI), and the behaviour of systems at different clock frequencies of CPUs. The results show that the programming model has a major impact on the energy consumption of computer systems. It was found that the impact of reduced clock frequencies on the execution time, energy efficiency, and maximum power consumption depends not only on the type of application but also on its implementation in a specific programming model. We believe that another criteria to consider when choosing a parallel programming model is the impact on energy consumption. Javier Aldo Balladini, Remo Suppi, Dolores Rexachs, Emilio Luque |
AICCSA | 3 |
| 2011 | Predicting parallel applications performance using signatures: The workload effectabstractBeing able to accurately estimate how an application will perform in a specific computational system provides many useful benefits and can result in smarter decisions. In this work we present a novel approach to model the behavior of message passing parallel applications. Based in the concept of signatures, which are the most relevant parts of an application (phases), we are able to build a model that allows us to predict the application execution time in different systems with variable input data size. Executing these signatures with different input data sizes defines a program's behavior partial function. Using regression we can generalize this behavior function to predict an application performance in a target system with other input data size within a predefined range. We explain our methodology and in order to validate the proposal present results using a synthetic program and well known applications. J. Martinez Canillas, Alvaro Wong, Dolores Rexachs, Emilio Luque |
AICCSA | 3 |
| 2011 | Performance Behavior Prediction Scheme for Shared-Memory Parallel ApplicationsabstractA current challenge in computing centers with different clusters to run applications is which multicore systems must we choose to run a given shared-memory parallel application. Our proposal is to generate a node performance profile database (NPPDB), composed by performance profiles given by distinct micro benchmark-target node combination. Then, applications are executed on a base node to identify different execution phases and their weights, and to collect performance and functional data for each phase. For similarity, the information to compare behavior is always obtained on the same node. When we want to project performance behavior, we look for similarity using the information from the performance profiles database with the phase characterization, in order to select the appropriate node for running the application. John Corredor, Juan C. Moure, Dolores Rexachs, Daniel Franco 0002, Emilio Luque |
CLUSTER | 3 |
| 2011 | Methodology for Performance Evaluation of the Input/Output System on Computer ClustersabstractThe increase of processing units, speed and computational power, and the complexity of scientific applications that use high performance computing require more efficient Input/Output (I/O) systems. In order to efficiently use the I/O it is necessary to know its performance capacity to determine if it fulfills applications I/O requirements. This paper proposes a methodology to evaluate I/O performance on computer clusters under different I/O configurations. This evaluation is useful to study how different I/O subsystem configurations will affect the application performance. This approach encompasses the characterization of the I/O system at three different levels: application, I/O system and I/O devices. We select different system configuration and/or I/O operation parameters and we evaluate the impact on performance by considering both the application and the I/O architecture. During I/O configuration analysis we identify configurable factors that have an impact on the performance of the I/O system. In addition, we extract information in order to select the most suitable configuration for the application. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CLUSTER | 2 |
| 2011 | Including the Workload Effect in the Parallel Program SignatureabstractPerformance prediction and application behavior modeling have been the subject of extensive research that aims to estimate applications performance with acceptable precision. In this paper we present a novel approach to model the behavior of message passing parallel applications. There are many dimensions to consider while predicting a deterministic application behavior. Two dimensions that affect an application performance are the computational resources available and the size of its input data used in the computation. Based on the concept of signatures, we are able to build a model that allows us to predict applications execution time in different systems with variable input data size within a predefined range. Our approach generates signatures, which consist of the most relevant parts of an application (phases). Executing these phases for different workloads partially defines a program's behavior function. By using regression analysis we are able to generalize this behavior function to predict an application performance in a target system with any input data size within a predefined range. We explain our methodology and in order to validate the proposal, we present results using a synthetic program and well-known applications. We were able to estimate the total execution time for a input data size range with an average error of 4 % executing, at most, three signatures that represent less than the 10 % of the total application execution time. J. Martinez Canillas, Alvaro Wong, Dolores Rexachs, Emilio Luque |
HPCC | 3 |
| 2011 | What is Missing in Current Checkpoint Interval Models?abstractThe growth in the number of components that compose parallel computers increases their fault frequency. Currently, in such systems faults are no longer a rare event but a common problem, thus some sort of fault tolerance should be provided. In general, fault tolerance protocols rely on checkpoints. A common question surrounding check pointing is the definition of the checkpoint interval. In this paper we propose the modelling of the relationship established between the parallel applications processes due to the messages exchange in order to incorporate this relationship into current checkpoint interval models. The experimental evaluation shows that the use of our checkpoint interval model based on the definition of the parallel application inter-process dependency factor is effective to calculate the checkpoint interval for parallel applications. Our results demonstrate that the overhead prediction error is smaller than 4% in comparison with the application execution. Leonardo Fialho, Dolores Rexachs, Emilio Luque |
ICDCS | 2 |
| 2010 | Methodology for Efficient Execution of SPMD Applications on Multicore EnvironmentsabstractThe need to efficiently execute applications in heterogeneous environments is a current challenge for parallel computing programmers. The communication heterogeneities found in multicore clusters need to be addressed to improve efficiency and speedup. This work presents a methodology developed for SPMD applications, which is centered on managing communication heterogeneities and improving system efficiency on multicore clusters. The methodology is composed of three phases: characterization, mapping strategy, and scheduling policy. We focus on SPMD applications which are designed through a message-passing library for communication, and selected according to their synchronicity and communications volume. The novel contribution of this methodology is it determines the approximate number of cores necessary to achieve a suitable solution with a good execution time, while the efficiency level is maintained over a threshold defined by users. Applying this methodology gave results showing a maximum improvement in efficiency of around 43% in the SPMD applications tested. Ronal Muresano, Dolores Rexachs, Emilio Luque |
CCGRID | 2 |
| 2010 | A reconfigurable cache memory with heterogeneous banksabstractThe optimal size of a large on-chip cache can be different for different programs: at some point, the reduction of cache misses achieved when increasing cache size hits diminishing returns, while the higher cache latency hurts performance. This paper presents the Amorphous Cache (AC), a reconfigurable L2 on-chip cache aimed at improving performance as well as reducing energy consumption. AC is composed of heterogeneous sub-caches as opposed to common caches using homogenous sub-caches. The sub-caches are turned off depending on the application workload to conserve power and minimize latencies. A novel reconfiguration algorithm based on Basic Block Vectors is proposed to recognize program phases, and a learning mechanism is used to select the appropriate cache configuration for each program phase. We compare our reconfigurable cache with existing proposals of adaptive and non-adaptive caches. Our results show that the combination of AC and the novel reconfiguration algorithm provides the best power consumption and performance. For example, on average, it reduces the cache access latency by 55.8%, the cache dynamic energy by 46.5%, and the cache leakage power by 49.3% with respect to a non-adaptive cache. Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
DATE | 3 |
| 2010 | Extraction of Parallel Application Signatures for Performance PredictionabstractPredicting performance of parallel applications is becoming increasingly complex and the best performance predictor is the application itself, but the time required to run it thoroughly is a onerous requirement. We seek to characterize the behavior of message-passing applications on different systems by extracting a signature which will allow us to predict what system will allow the application to perform best. To achieve this goal, we have developed a method we called Parallel Application Signatures for Performance Prediction (PAS2P) that strives to describe an application based on its behavior. Based on the application's message-passing activity, we have been able to identify and extract representative phases, with which we created a Parallel Application Signature that has allowed us to predict the application's performance. We have experimented with different signature-extraction algorithms and found a reduction in the prediction error using different scientific applications on different clusters. We were able to predict execution times with an average accuracy of over 98%. Alvaro Wong, Dolores Rexachs, Emilio Luque |
HPCC | 2 |
| 2009 | An assessment of multi-core for a performance prediction model of tomographic reconstructionabstractThree-dimensional (3D) reconstruction of structures from projection data is essential for helping people in a wide range of areas. Algebraic reconstruction techniques (ART) are iterative procedures for recovering the structure of the 3D objects from projection images. During the seventies, the ART were dismissed due to high-demanding computing requirements. Interesting recent research aims at acquiring experience with parallelization strategies and at demonstrating the effectiveness of the massively parallel processing approach in 3D reconstructions. Multi-core (MC) technology provides new levels of performance and, therefore, it is of paramount importance to make performance predictions. The objectives of this work are both adapting an analytical performance prediction model for the iterative reconstruction techniques (IRT) to a MC environment and finding a process's CPU affinity that produces the best overall performance. BPTomo+is a parallel distributed application for tomographic reconstruction that uses IRT. Besides, it includes a process's CPU affinity mask. The analytical performance prediction model is validated by comparison of the estimated times for representative datasets against BPTomo+computation times measured on a MC server. The analytical model is shown to be quite accurate. The percentage of deviation between estimated and measured times is less than 6%. Paula Cecilia Fritzsche, Ronal Muresano, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2009 | How SPMD applications could be efficiently executed on multicore environments?abstractA challenge for programmers of parallel programming environments is to execute applications efficiently. For this reason, applications with high levels of synchronism and communications such as SPMD (single program multiple data) create a challenge regarding how to distribute tasks between PE (processing element) in a multicore cluster; this kind of environment presents high heterogeneity in communication parameters due to different communication paths present. For this reason, this work is centered around developing a methodology to distribute SPMD tasks between PEs in a multicore cluster. The task assignment process is realized through mapping and scheduling strategies based on controlling the communications heterogeneities. Finally, the objective is to obtain a good execution time while maintaining the efficiency level over a threshold. The results obtained show an improvement around 40% of efficiency in a heat transfer application, when our methodology is applied. Ronal Muresano, Dolores Rexachs, Emilio Luque |
CLUSTER | 2 |
| 2009 | Increasing the availability provided by RADIC with low overheadabstractFor machines composed of a large number of processing units, fault probability tends to increase linearly with this number. This makes the use of a fault tolerant solution a major issue. A fault tolerant solution provides certain level of availability, which is usually influenced by time overhead, performance degradation, resources or cost. In the rollback-recovery protocol, the availability increase is usually achieved by increasing the checkpoint frequency or by making several replicas of checkpoints and/or logs. Such a replication allows the solution to tolerate concurrent correlated faults, i.e., a fault in a computing node and in the stable storage. These faults are theoretically less probable, however recent studies have shown that faults are temporally and spatially correlated, consequently increasing the concurrent fault probability. The major concern replicating the checkpoints and logs is the overhead caused by storing these replicas over various repositories, which may disallow its use. In this paper we present how we increased the availability provided by RADIC, without significantly increase of its overhead. Our approach consists of parallelizing the storing of these replicas using the pipeline technique. Such a technique allows us to make low-overhead copies of checkpoints and logs over N protectors. Furthermore, as secondary benefit, the pipelining between observer and protector reduces more than four times (in the best case) the pessimistic message logging overhead. Guna Santos, Leonardo Fialho, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2009 | Parallel application signatureabstractWe seek to achieve characterization or application signature from a parallel application that will allow us, through the execution of this signature, to evaluate its performance in different computers. Sequential applications behavior can be understood by means of tools such as SimPoint. This tool can identify and select significant phases describing the applications behavior. Our proposal is to extend those concepts towards parallel applications, with the goal of modeling and predicting the parallel application. To achieve this, we developed a methodology, enabling us to identify and extract repetitive behavior to create the application signature. We have validated our proposal using scientific applications such as the NAS Parallel Benchmarks, Sweep3D. We could predict the execution time of the entire application. Alvaro Wong, Dolores Rexachs, Emilio Luque |
CLUSTER | 2 |
| 2008 | Increasing the Performability of Computer Clusters Using RADIC IIabstractPerformance and availability form an undissociable binomial for some kind of applications. Therefore, the fault tolerant solutions must take into consideration these two constraints when it has been designed. Our previous work, called RADIC, implemented a basic level protection allowing to recover from faults just using the active cluster resources, changing the system configuration. However, Such approach may genenerate some performance degradation in some cases. In this paper, we present RADIC II, which incorporates a new protection level using dynamic redundancy, allowing to mitigate or avoid the recovery side-effects. Such functionality allows restoring a changed system configuration and it can avoid the configuration changes. The results has shown that RADIC-II operates correctly and becomes itself as a good approach to provide high availability to the parallel applications without suffer a system degradation in post-recovery execution. Guna Santos, Angelo Duarte, Dolores Rexachs, Emilio Luque |
ARES | 3 |
| 2008 | Providing Non-stop Service for Message-Passing Based Parallel Applications with RADIC
Guna Santos, Angelo Duarte, Dolores Rexachs, Emilio Luque |
Euro-Par | 3 |
| 2008 | A General Approach to Predict the Performance Order of TSP Family Problems
Paula Cecilia Fritzsche, Dolores Rexachs, Emilio Luque |
ICA3PP | 2 |
| 2008 | Software Probes: Towards a Quick Method for Machine Characterization and Application Performance PredictionabstractComputers perform different applications in different ways. To characterize an application performance into a machine, the usual method is a throughout execution of it. This work is a step into a synthetic probe able to characterize a master-worker application's performance in a fraction of the time required to run it entirely. This is specially important for CPU-intensive scientific applications, who runs for very long, as it makes sense that it runs as efficiently (and fast) as possible. To know how, and for how long a master-worker application is going to run can guide the decision to use this machine or not. Our software probe takes into account only the performance-relevant parts of the application, discovering a program's relevant phases. Running solely these significant phases is a powerful way to quickly characterize the application's performance on a machine. It can help to select the best computing nodes in a grid or in a multi-cluster to run this application, and even quickly predict the total execution time for this application/data set in the machine analyzed. We also present ongoing work on a fully synthetic probe generated from programs' phases. Alexandre Otto Strube, Dolores Rexachs, Emilio Luque |
ISPDC | 2 |
| 2007 | Applying Data Mining to Define TSP Asymptotic Time ComplexityabstractComputational science is often referred to as the third science,complementing both theoretical and laboratory science. In this field, new challenges are continuously arising. The asymptotic time complexity definition of both deterministic and non-deterministic algorithms to solve all kinds of problems is one of the key points in computer science. Knowing the limit of the execution time of an algorithm when the size of the problem goes to infinity is essential. In particular, data-dependent applications is an extremely challenging problem because for a specific issue the input data sets may cause variability in execution times. The development of an entire approach to define the asymptotic time complexity of a hard data-dependent parallel application that solves the traveling salesman problem (TSP) is the focus of this study. Two different parallel TSP algorithms are presented. One of these is used to show the usefulness and the profits of the proposed approach, and the other one is used as witness. The experimental results are quite promising. Paula Cecilia Fritzsche, Dolores Rexachs, Emilio Luque |
ICTAI (2) | 2 |
| 2007 | Functional Tests of the RADIC Fault Tolerance ArchitectureabstractClusters with thousand of nodes are a reality and the current trend indicates that they are becoming larger. Such large clusters are subject to a relatively high fault frequency so a fault-tolerance scheme is mandatory to assure the correct application completion. Message passing is the programming model often used in large clusters and the current implementations used to achieve fault tolerance in message passing systems do not focus in an architecture that simultaneously attends to scalability, transparency and independence of stable/central elements. The RADIC architecture was proposed and design as a fully distributed structure in order to achieve such requirements. Such architecture defines a fully distributed fault tolerance controller implemented by a set of system processes, which collaborate in order to perform all the basic functions of a fault tolerance protocol. This paper presents the test methodology used to verify the functionality of the RADIC architecture using RADICMPI, a prototype on the MPI semantic Angelo Duarte, Dolores Rexachs, Emilio Luque |
PDP | 2 |
| 2006 | Increasing the cluster availability using RADICabstractThe redundant array of distributed independent checkpoints (RADIC) is a fault tolerant architecture based on a fully distributed array of dedicated process. These processes collaborate to create a fault tolerance controller which transparently manages all fault tolerance activities. The architecture is designed as a software layer between the application and the cluster structure and it was developed to attend to the requirements of scalability, user transparency and independency of dedicated/stable cluster resources. RADIC only requires the resources already available in the nodes used by the parallel application and it uses a pessimistic message-log rollback-recovery protocol in order to operate without any global synchronization. Such protocol, together with the independence of central elements, makes RADIC a scalable architecture that works transparently to the user. We tested the functionality and performance of the architecture in a real scenario using a prototype based on the MPI standard (RADICMPI) Angelo Duarte, Dolores Rexachs, Emilio Luque |
CLUSTER | 2 |
| 2006 | Tuning Application in a Multi-cluster Environment
Eduardo Argollo, Adriana Gaudiani, Dolores Rexachs, Emilio Luque |
Euro-Par | 3 |
| 2006 | Wide and efficient trace prediction using the local trace predictorabstractHigh prediction bandwidth enables performance improvements and power reduction techniques. This paper explores a mechanism to increase prediction width (instructions per prediction) by predicting instruction traces. Our analysis shows that predicting traces including multiple branches is not significantly less accurate than predicting single branches. A novel Local Trace Predictor organization is proposed. It increases prediction width without reducing the ratio of prediction accuracy versus memory resources with respect to a Basic Block Predictor.Compared to the previously proposed Next-Trace Predictor, the Local Trace Predictor reduces memory requirements by codifying trace predictions, and by limiting the number of traces starting at the same instruction to 2 or 4. The limit lessens prediction width only slightly, and does not affect prediction accuracy. The overall result is that the Local Trace Predictor outperforms the Next-Trace Predictor for sizes higher than 12 KBytes. Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
ICS | 3 |
| 2005 | Target Encoding for Efficient Indirect Jump Prediction
Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
Euro-Par | 3 |
| 2005 | Performance and Power Evaluation of an Intelligently Adaptive Data Cache
Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
HiPC | 3 |
| 2004 | Graduate students learning strategies through research collaborationabstractIt is already known that the learning process can be accelerated with the mixture of theoretical classes and experimental work. This paper describes an interesting experiment with that combination in the teaching of computer architecture for Ph.D. students in collaboration with a researcher in a real design investigation. As the work progressed, a simple cyclical methodology arose as reference for future works. Eduardo Argollo, Mauricio Hanzich, Diego Mostaccio, Germán Bianchini, Paula Cecilia Fritzsche, Ferran Bonàs, Emilio Luque, Juan C. Moure, Dolores Rexachs |
ITiCSE | 9 |
| 2003 | Optimizing a Decoupled Front-End Architecture: The Indexed Fetch Target Buffer (iFTB)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 2 |
| 2002 | Speeding Up Target Address Generation Using a Self-indexed FTB (Research Note)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 2 |
| 2002 | The KScalar simulatorabstractModern processors increase their performance with complex microarchitectural mechanisms, which makes them more and more difficult to understand and evaluate. KScalar is a graphical simulation tool that facilitates the study of such processors. It allows students to analyze the performance behavior of a wide range of processor microarchitectures: from a very simple in-order, scalar pipeline, to a detailed out-of-order, superscalar pipeline with non-blocking caches, speculative execution, and complex branch prediction. The simulator interprets executables for the Alpha AXP instruction set: from very short program fragments to large applications. The object's program execution may be simulated in varying levels of detail: either cycle-by-cycle, observing all the pipeline events that determine processor performance, or million cycles at once, taking statistics of the main performance issues.Instructors may use KScalar in several ways. First, it may be used to provide demonstrations in lectures or online learning environments. Second, it allows students to investigate the characteristics of specific processor microarchitectures as practical short assignments associated to a lecture course. Third, students may undertake major projects involving the optimization of real programs at the software-hardware interface, or involving the optimization of a processor microarchitecture for a given application workload.A preliminary version of KScalar has been successfully used in several lecture courses during the last two years in the University Autónoma of Barcelona. It runs on a x86/Linux/KDE system. The graphical interface has been developed using the KDE and QT libraries. The simulator engine running behind the graphical interface is a heavily-modified version of SimpleScalar. KScalar code is available under the terms of the GNU and SimpleScalar General Public License Juan C. Moure, Dolores Rexachs, Emilio Luque |
ACM J. Educ. Resour. Comput. | 2 |
| 2001 | Improving Single-Thread Fetch Performance on a Multithreaded ProcessorabstractMultithreaded processors, by simultaneously using both the thread-level parallelism and the instruction-level parallelism of applications, achieve larger instruction per cycle rate than single-thread processors. On a multi-thread workload, a clustered organization maximizes performances. On a single-thread workload, however, all but one of the clusters are idle, degrading single-thread performance significantly. Using a clustered multi-thread performance as a baseline, we propose and analyze several mechanisms and policies to improve single-thread execution exploiting the existing hardware without a significant multi-thread performance loss. We focus on the fetch unit, which is maybe the most performance-critical stage. Essentially, we analyze three ways of exploiting the idle fetch clusters: allowing a single thread accessing its neighbor clusters, use the idle fetch clusters to provide multiple-path execution, or use them to widen the effective single-three fetch block. Juan C. Moure, R. B. García, Dolores Rexachs, Emilio Luque |
DSD | 3 |