EDBT 2026 Demo / reviewers in the wild / expert
Emilio Luque
dblp:92/4942 · also Emilio Luque Fadon
· DBLP profile ↗
146ranked-venue papers
19as first author
8since 2021 · last 2026
0000-0002-2884-3232ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 112 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6Human-computer interaction and ubiquitous computing · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-authorTheory of computation · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Application of a sampling and clustering-based heuristic search algorithm to find an efficient staff configuration in an emergency departmentabstractEmergency Departments (EDs) are among the most complex areas in healthcare, requiring immediate medical attention for acute and urgent conditions. Optimizing staff configurations to reduce patient Length of Stay (LoS) and improve operational efficiency poses a significant challenge due to the combinatorial and high-dimensional nature of the problem. To identify the most effective staff configuration, we propose a heuristic optimization strategy that is based on the Montecarlo Clustering Search Algorithm (MCSA), which efficiently explores the multidimensional solution space. MCSA leverages an agent-based simulation (ABM) model that evaluates each proposed staff configuration under realistic operational conditions, providing Key Performance Indicator (KPI) feedback values related to each proposed staff configuration. Through this strategy, we explore staff configurations capable of handling patient volumes with varying acuity levels in an ED to optimize the LoS KPI. Results demonstrate that our methodology is capable to find a solution as a staff configuration that reduces LoS compared to a baseline, offering a computationally efficient and practical tool for decision-makers. We identified solutions by exploring less than 1% of the total search space, demonstrating the efficiency of the proposed approach in addressing complex optimization problems. This approach supports informed planning in healthcare environments while maintaining system feasibility and scalability. Maria Harita, Alvaro Wong, Dolores Rexachs, Emilio Luque, Eva Bruballa, Francisco Epelde |
Expert Syst. Appl. | 4 |
| 2025 | Parallel I/O analysis in distributed deep learning applications on high-performance computingabstractAbstract Distributed deep learning (DDL) applications generate heavy input/output (I/O) workloads that can create bottlenecks in high-performance computing (HPC) systems. Their optimal I/O configuration depends on factors such as access patterns, storage hardware, dataset size, and execution scale. This study proposes a systematic methodology for characterizing and optimizing I/O behavior in DDL applications, represented through the deep learning I/O benchmark (DLIO), and validated with the real DeepGalaxy application. We evaluate access modes, file formats, and Lustre file system configurations, demonstrating that stripe counts optimized for the access pattern and application scale can reduce I/O and execution times, achieving up to 18 GiB/s of bandwidth and a 5X increase in IOPS. HDF5 provides balanced performance, while TFRecord stands out in bandwidth-intensive scenarios. Shared access minimizes contention and improves scalability in multi-node executions. The results are consolidated into configuration guidelines that offer practical recommendations for practitioners to tune DDL applications for efficient execution in HPC environments. Edixon Párraga, Betzabeth León, Sandra Méndez, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 5 |
| 2023 | A computational methodology applied to optimize the performance of a river model under uncertainty conditions
Adriana Gaudiani, Alvaro Wong, Emilio Luque, Dolores Rexachs |
J. Supercomput. | 3 |
| 2022 | A model of checkpoint behavior for applications that have I/OabstractAbstract Due to the increase and complexity of computer systems, reducing the overhead of fault tolerance techniques has become important in recent years. One technique in fault tolerance is checkpointing, which saves a snapshot with the information that has been computed up to a specific moment, suspending the execution of the application, consuming I/O resources and network bandwidth. Characterizing the files that are generated when performing the checkpoint of a parallel application is useful to determine the resources consumed and their impact on the I/O system. It is also important to characterize the application that performs checkpoints, and one of these characteristics is whether the application does I/O. In this paper, we present a model of checkpoint behavior for parallel applications that performs I/O; this depends on the application and on other factors such as the number of processes, the mapping of processes and the type of I/O used. These characteristics will also influence scalability, the resources consumed and their impact on the IO system. Our model describes the behavior of the checkpoint size based on the characteristics of the system and the type (or model) of I/O used, such as the number I/O aggregator processes, the buffering size utilized by the two-phase I/O optimization technique and components of collective file I/O operations. The BT benchmark and FLASH I/O are analyzed under different configurations of aggregator processes and buffer size to explain our approach. The model can be useful when selecting what type of checkpoint configuration is more appropriate according to the applications’ characteristics and resources available. Thus, the user will be able to know how much storage space the checkpoint consumes and how much the application consumes, in order to establish policies that help improve the distribution of resources. Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 5 |
| 2022 | Correction to: A model of checkpoint behavior for applications that have I/O
Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 5 |
| 2022 | Scalable performance analysis method for SPMD applicationsabstractAbstract The analysis of parallel scientific applications allows us to understand their computational and communication behavior. One way of obtaining performance information is through performance tools. One such tool is parallel application signatures for performance prediction (PAS2P), based on parallel application repeatability, focusing on performance analysis and prediction. The same resources that execute the parallel application are used to perform its analysis, creating a machine independent model of the application and identifying its common patterns. However, the analysis is costly in terms of execution time due to the high number of synchronization communications performed by PAS2P, degrading performance as the number of processes increases. To solve this problem, we propose a model that reduces data dependency between processes, reducing the number of communications performed by PAS2P in the analysis stage and taking advantage of the characteristics of single program, multiple sata applications. Our analysis proposal allows us to decrease the analysis time by 29 times when the application scales to 256 processes, while keeping error levels below 11% in the runtime prediction. It is important to mention that the analysis time is not considerably affected by increasing the number of application processes. Felipe Tirado, Alvaro Wong, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 4 |
| 2021 | Analysis of parallel application checkpoint storage for system configuration
Betzabeth León, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 4 |
| 2021 | Middleware to Manage Fault Tolerance Using Semi-Coordinated CheckpointsabstractCompute node failures are becoming a normal event for many long-running and scalable MPI applications. Keeping within the MPI standards and applying some of the methods developed so far in terms of fault tolerance, we developed a methodology that allows applications to tolerate failures through the creation of semi-coordinated checkpoints within the RADIC architecture. To do this, we developed the ULSC2-RADIC middleware that divides the application into independent MPI worlds where each MPI world would correspond to a compute node and make use of the DMTCP checkpoint library in a semi-coordinated environment. We performed experimental results using scientific applications and the NAS Parallel Benchmarks to assess the overhead and also the functionality in case of a node failure. We evaluated the computational cost of the semi-coordinated checkpoints compared with the coordinated checkpoints. Alvaro Wong, Elisa Heymann, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Soft errors detection and automatic recovery based on replication combined with different levels of checkpointing
Diego Montezanti, Enzo Rucci, Armando De Giusti, Marcelo R. Naiouf, Dolores Rexachs, Emilio Luque |
Future Gener. Comput. Syst. | 6 |
| 2020 | A Method for Projections of the Emergency Department Behaviour by Non-Communicable Diseases From 2019 to 2039abstractIn this paper, a new method for prediction of future performance and demand on emergency department (ED) in Spain is presented. Increased life expediency and population aging in Spain, along with their corresponding health conditions such as non-communicable diseases (NCDs), have been suggested to contribute to higher demands on ED. These lead to inferior performance of the department and cause longer ED length of stay (LoS). Prediction and quantification of behavior of ED is, however, challenging as ED is one of the most complex parts of hospitals. Using detailed computational approaches integrated with clinical data behavior of Spain's ED in future years was predicted. First, statistical models were developed to predict how the population and age distribution of patients with non-communicable diseases change in Spain in future years. Then, an agent-based modeling approach was used for simulation of the emergency department to predict impacts of the changes in population and age distribution of patients with NCDs on the performance of ED, reflected in ED LoS, between years 2019 and 2039. Results from different projection scenarios indicated that Spain would experience a continuous increase in total ED LoS from 5.7 million hours in 2019 to 6.2 million hours in 2039 if same human and physical resources, as well as same ED configuration, are used. The results from this study can provide health care provider with quantitative information on required staff and physical resources in the future and allow health care policymakers to improve modifiable factors contributing to the demand and performance of ED. Elham Shojaei, Alvaro Wong, Dolores Rexachs, Francisco Epelde, Emilio Luque |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | RaaS: Resilience as a ServiceabstractCloud computing is continuously increasing its popularity as key features such as scalability, pay-per-use and availability continue to evolve. It is also becoming a competitive platform for running high performance computing (HPC) and parallel applications due to the increasing performance of virtualized, highly-available instances. However, migrating HPC applications to cloud still requires native fault-tolerant solutions to fully leverage cloud features and maximize the resource utilization at the best cost - particularly for long-running parallel applications where faults can cause invalid states or data loss. This requires re-executing applications which increases completion time and cost. We propose Resilience as a Service (RaaS), a fault tolerant framework for HPC applications running in cloud. In this paper RADIC architecture (Redundant Array of Distributed Independent Fault Tolerance Controllers) is used to provide clouds with a highly available, distributed and scalable fault-tolerant service. The paper explores how traditional HPC protection and recovery mechanisms must be redesigned to natively leverage cloud properties and its multiple alternatives for implementing rollback recovery protocols using virtual machines, containers, object and block storage or database services. Results show that RaaS restores and completes the application execution using available resources while reducing overhead up to 8% for different fault-tolerant configuration alternatives. Jorge Villamayor, Dolores Rexachs, Emilio Luque, Diego Lugones |
CCGrid | 3 |
| 2018 | P3S: A Methodology to Analyze and Predict Application ScalabilityabstractExecuting message-passing parallel applications on a large number of resources in an efficient way is not a trivial task. Due to the complex interaction between the parallel applications and the HPC system, many applications may suffer performance inefficiencies when they scale. To achieve an efficient use of these large-scale systems using thousands of cores, a point to consider before executing an application is to know its behavior in the system. In this work, we propose a novel methodology called P3S (Prediction of Parallel Program Scalability), which allows us to analyze and predict the scalability of message-passing applications on a given system. The methodology strives to use a bounded analysis time, and a reduced set of resources to predict the application behavior for large-scale. The experimental validation proves that the P3S is able to predict the application scalability with an average accuracy greater than 95 percent using a reduced set of resources. Javier Panadero, Alvaro Wong, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Analyzing the Parallel I/O Severity of MPI ApplicationsabstractPerformance evaluation of parallel applications plays an important role in High Performance Computing (HPC). This is also applied to parallel I/O performance evaluation, which requires understanding the I/O pattern of the application and having knowledge about the performance capacity of the HPC I/O system. In this paper, we present a methodology to evaluate the I/O performance of parallel applications based on the I/O severity degree. We define the I/O severity concept taking into account the I/O requirements of a parallel application, the mapping of I/O processes and the configuration of the I/O subsystem. Requirements are expressed in units denominated I/O phases, which are defined using the temporal and spatial pattern of different files of the application. Our approach is applied to the I/O kernels of scientific applications such as S3DIO, FLASH-IO and BT-IO on the SuperMUC supercomputer. Experimental results show that our methodology allows us to identify if a parallel application is limited by the I/O subsystem and identifying possible root causes of the I/O problems. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CCGrid | 3 |
| 2017 | Improving the Network of Search Engine Services Through Application-Driven Routing
Joe Carrión, Daniel Franco 0002, Veronica Gil-Costa, Mauricio Marín, Emilio Luque |
Euro-Par | 5 |
| 2017 | Care HPS: A high performance simulation tool for parallel and distributed agent-based modeling
Francisco Borges, Albert Gutierrez-Milla, Emilio Luque, Remo Suppi |
Future Gener. Comput. Syst. | 3 |
| 2017 | An approach for an efficient execution of SPMD applications on Multi-core environments
Ronal Muresano, Hugo Meyer, Dolores Rexachs, Emilio Luque |
Future Gener. Comput. Syst. | 4 |
| 2017 | Hybrid Message Pessimistic Logging. Improving current pessimistic message logging protocols
Hugo Meyer, Ronal Muresano, Marcela Castro-León, Dolores Rexachs, Emilio Luque |
J. Parallel Distributed Comput. | 5 |
| 2015 | Fault tolerance at system level based on RADIC architectureabstractThe increasing failure rate in High Performance Computing encourages the investigation of fault tolerance mechanisms to guarantee the execution of an application in spite of node faults. This paper presents an automatic and scalable fault tolerant model designed to be transparent for applications and for message passing libraries. The model consists of detecting failures in the communication socket caused by a faulty node. In those cases, the affected processes are recovered in a healthy node and the connections are reestablished without losing data. The Redundant Array of Distributed Independent Controllers architecture proposes a decentralized model for all the tasks required in a fault tolerance system: protection, detection, recovery and masking. Decentralized algorithms allow the application to scale, which is a key property for current HPC system. Three different rollback recovery protocols are defined and discussed with the aim of offering alternatives to reduce overhead when multicore systems are used. A prototype has been implemented to carry out an exhaustive experimental evaluation through Master/Worker and Single Program Multiple Data execution models. Multiple workloads and an increasing number of processes have been taken into account to compare the above mentioned protocols. The executions take place in two multicore Linux clusters with different socket communications libraries. Marcela Castro-León, Hugo Meyer, Dolores Rexachs, Emilio Luque |
J. Parallel Distributed Comput. | 4 |
| 2015 | Parallel Application Signature for Performance Analysis and PredictionabstractPredicting the performance of parallel scientific applications is becoming increasingly complex. Our goal was to characterize the behavior of message-passing applications on different target machines. To achieve this goal, we developed a method called parallel application signature for performance prediction (PAS2P), which strives to describe an application based on its behavior. Based on the application's message-passing activity, we identified and extracted representative phases, with which we created a parallel application signature that enabled us to predict the application's performance. We experimented with using different scientific applications on different clusters. We were able to predict execution times with an average accuracy greater than 97 percent. Alvaro Wong, Dolores Rexachs, Emilio Luque |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | "Analysis of scalability: A parallel application model approach"abstractIn this paper we propose a methodology that allows us to predict the application scalability behavior in a specific system, providing information to select the most appropriate resources to run the application. We explain the general methodology, focusing on the presentation of a novel method to model the logical application trace for a large number of processes. This method is based on the projection of a set of executions of the application signature for a small number of processes. The generated traces are validated by comparing them with the real traces obtained with PAS2P tool. We present the experimental validation for the BT Nas Parallel Benchmark. The signatures for 16, 36, 64, 81 and 100 processes were executed and used to model and project the logical trace for 1024 processes. The results obtained show the accuracy of the method. The communication pattern was predicted without error, while the predicted error is less than 10% for the communication volume and less than 5% for the number of instructions. Javier Panadero, Alvaro Wong, Dolores Rexachs, Emilio Luque |
CLUSTER | 4 |
| 2012 | A Fault-Tolerant Cache Service for Web Search Engines: RADIC Evaluation
Carlos Gómez-Pantoja, Dolores Rexachs, Mauricio Marín, Emilio Luque |
Euro-Par | 4 |
| 2012 | Transparent Fault Tolerance Solution at Socket Level Based on RADICabstractWe present a transparent middleware for fault tolerance based on RADIC, Redundant Array of Distributed Independent Controllers, a transparent and scalable fault tolerant architecture for parallel applications. It is designed at socket level and makes a secure tunnel connection able to keep the tcp sessions established by the application in spite of node failures. It is located at user level and is independent of the message-passing communication library being used. The protection gets through uncoordinated checkpoints and log message and the recovery are done in a automatic way so in case of node failures there is no need of intervention of the administrator. We have tested our fault tolerance system by executing a master-worker (M/W) and SPMD applications that follow different communication patterns. Marcela Castro-León, Dolores Rexachs, Emilio Luque |
ISPA | 3 |
| 2012 | A Fault-Tolerant Cache Service for Web Search EnginesabstractLarge Web search engines are constructed as a collection of services that are deployed on dedicated clusters of distributed-memory processors. In particular, efficient user query throughput heavily relies on using result cache services devoted to maintaining the answers to most frequent queries. Load balancing and fault tolerance are critical to this service. This paper proposes the design of a result cache service based on consistent hashing and a strategy for enabling fault tolerance. Performance evaluation is performed by using actual queries from a commercial search engine. The results show that the proposed cache service outperforms baseline approaches, decreases the average query response time, increases query throughput and efficiently recovers performance after processor failures. Carlos Gómez-Pantoja, Veronica Gil-Costa, Dolores Rexachs, Mauricio Marín, Emilio Luque |
ISPA | 5 |
| 2012 | A methodology for transparent knowledge specification in a dynamic tuning environmentabstractSUMMARY The increasing use of parallel/distributed applications demands a continuous support to take significant advantages from parallel power. This includes the evolution of performance analysis and tuning tools which automatically allows for obtaining a better behavior of the applications. Different approaches and tools have been proposed and they are continuously evolving to cover the requirements and expectations of users. One such tool is MATE (Monitoring Analysis and Tuning Environment), which provides automatic and dynamic tuning for parallel/distributed applications. The knowledge used by MATE to analyze and take decisions is based on performance models which include a set of performance parameters and a set of mathematical expressions modeling the solution of the performance problem. These elements are used by the tuning environment to conduct the monitoring and analysis steps, respectively. The tuning phase depends on the results of the performance analysis. This paper presents a methodology to specify performance models. Each performance model specification can be automatically and transparently translated into a piece of software code encapsulating the knowledge to be straightforwardly included in MATE. Applying this methodology, the user does not have to be involved in the implementation details of MATE, which makes the usage of the tool more transparent. Copyright © 2011 John Wiley & Sons, Ltd. Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque |
Softw. Pract. Exp. | 4 |
| 2011 | Impact of parallel programming models and CPUs clock frequency on energy consumption of HPC systemsabstractEnergy consumption has become one of the greatest challenges in the field of high performance computing (HPC). The energy cost produced by supercomputers during the lifetime of the installation is similar to acquisition. Thus, besides its impact on the environment, energy is a limiting factor for the HPC. Our research aims to reduce the energy consumption of computer systems to run parallel HPC applications. In this article we analyse the possible influence on the energy consumption of parallel programming paradigms of shared memory (OpenMP) and message passing (MPI), and the behaviour of systems at different clock frequencies of CPUs. The results show that the programming model has a major impact on the energy consumption of computer systems. It was found that the impact of reduced clock frequencies on the execution time, energy efficiency, and maximum power consumption depends not only on the type of application but also on its implementation in a specific programming model. We believe that another criteria to consider when choosing a parallel programming model is the impact on energy consumption. Javier Aldo Balladini, Remo Suppi, Dolores Rexachs, Emilio Luque |
AICCSA | 4 |
| 2011 | Predicting parallel applications performance using signatures: The workload effectabstractBeing able to accurately estimate how an application will perform in a specific computational system provides many useful benefits and can result in smarter decisions. In this work we present a novel approach to model the behavior of message passing parallel applications. Based in the concept of signatures, which are the most relevant parts of an application (phases), we are able to build a model that allows us to predict the application execution time in different systems with variable input data size. Executing these signatures with different input data sizes defines a program's behavior partial function. Using regression we can generalize this behavior function to predict an application performance in a target system with other input data size within a predefined range. We explain our methodology and in order to validate the proposal present results using a synthetic program and well known applications. J. Martinez Canillas, Alvaro Wong, Dolores Rexachs, Emilio Luque |
AICCSA | 4 |
| 2011 | Predictive and Distributed Routing Balancing for High Speed Interconnection NetworksabstractCurrent parallel applications in parallel computing systems require an interconnection network to provide low and bounded communication delays. Communication characteristics such as traffic pattern and communication load change over time and, eventually, they may exceed available network capacity causing congestion and performance degradation. Congestion control based on adaptive routing should be applied in order to adapt quickly to changing traffic conditions. Studies on a vast range of parallel applications show repetitive behavior and can be characterized by a set of representative phases. This work presents a Predictive and Distributed Routing Balancing technique (PR-DRB) to control network congestion based on adaptive traffic distribution. PR-DRB uses speculative routing based on application repetitiveness. PR-DRB monitors messages latencies on routers and logs solutions to congestion, to quickly respond in future similar situations. Experimental results show that the predictive approach could be used to improve performance. Carlos Nunez Castillo, Diego Lugones, Daniel Franco 0002, Emilio Luque |
CLUSTER | 4 |
| 2011 | Performance Behavior Prediction Scheme for Shared-Memory Parallel ApplicationsabstractA current challenge in computing centers with different clusters to run applications is which multicore systems must we choose to run a given shared-memory parallel application. Our proposal is to generate a node performance profile database (NPPDB), composed by performance profiles given by distinct micro benchmark-target node combination. Then, applications are executed on a base node to identify different execution phases and their weights, and to collect performance and functional data for each phase. For similarity, the information to compare behavior is always obtained on the same node. When we want to project performance behavior, we look for similarity using the information from the performance profiles database with the phase characterization, in order to select the appropriate node for running the application. John Corredor, Juan C. Moure, Dolores Rexachs, Daniel Franco 0002, Emilio Luque |
CLUSTER | 5 |
| 2011 | Methodology for Performance Evaluation of the Input/Output System on Computer ClustersabstractThe increase of processing units, speed and computational power, and the complexity of scientific applications that use high performance computing require more efficient Input/Output (I/O) systems. In order to efficiently use the I/O it is necessary to know its performance capacity to determine if it fulfills applications I/O requirements. This paper proposes a methodology to evaluate I/O performance on computer clusters under different I/O configurations. This evaluation is useful to study how different I/O subsystem configurations will affect the application performance. This approach encompasses the characterization of the I/O system at three different levels: application, I/O system and I/O devices. We select different system configuration and/or I/O operation parameters and we evaluate the impact on performance by considering both the application and the I/O architecture. During I/O configuration analysis we identify configurable factors that have an impact on the performance of the I/O system. In addition, we extract information in order to select the most suitable configuration for the application. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2011 | Including the Workload Effect in the Parallel Program SignatureabstractPerformance prediction and application behavior modeling have been the subject of extensive research that aims to estimate applications performance with acceptable precision. In this paper we present a novel approach to model the behavior of message passing parallel applications. There are many dimensions to consider while predicting a deterministic application behavior. Two dimensions that affect an application performance are the computational resources available and the size of its input data used in the computation. Based on the concept of signatures, we are able to build a model that allows us to predict applications execution time in different systems with variable input data size within a predefined range. Our approach generates signatures, which consist of the most relevant parts of an application (phases). Executing these phases for different workloads partially defines a program's behavior function. By using regression analysis we are able to generalize this behavior function to predict an application performance in a target system with any input data size within a predefined range. We explain our methodology and in order to validate the proposal, we present results using a synthetic program and well-known applications. We were able to estimate the total execution time for a input data size range with an average error of 4 % executing, at most, three signatures that represent less than the 10 % of the total application execution time. J. Martinez Canillas, Alvaro Wong, Dolores Rexachs, Emilio Luque |
HPCC | 4 |
| 2011 | What is Missing in Current Checkpoint Interval Models?abstractThe growth in the number of components that compose parallel computers increases their fault frequency. Currently, in such systems faults are no longer a rare event but a common problem, thus some sort of fault tolerance should be provided. In general, fault tolerance protocols rely on checkpoints. A common question surrounding check pointing is the definition of the checkpoint interval. In this paper we propose the modelling of the relationship established between the parallel applications processes due to the messages exchange in order to incorporate this relationship into current checkpoint interval models. The experimental evaluation shows that the use of our checkpoint interval model based on the definition of the parallel application inter-process dependency factor is effective to calculate the checkpoint interval for parallel applications. Our results demonstrate that the overhead prediction error is smaller than 4% in comparison with the application execution. Leonardo Fialho, Dolores Rexachs, Emilio Luque |
ICDCS | 3 |
| 2011 | Predictive and Distributed Routing Balancing on High-Speed Cluster NetworksabstractIn high performance clusters current parallel application communication needs such as traffic pattern, communication volume, etc., change along time and are difficult to know in advance. Such needs often exceed or do not match available resources causing resource use imbalance, network congestion, throughput reduction and message latency increase, thus degrading the overall system performance. Studies on parallel applications show repetitive behavior that can be characterized by a set of representative phases. This work presents a Predictive and Distributed Routing Balancing (PRDRB) technique, a new method developed to gradually control network congestion, based on paths expansion, traffic distribution, applications pattern repetitiveness and speculative adaptive routing, in order to maintain low latency values. PRDRB monitors messages latencies on routers and logs solutions to congestion, to quickly respond in future similar situations. Traffic congestion experiments were conducted in order to evaluate the performance of the method, and improvements were observed. Carlos Nunez Castillo, Diego Lugones, Daniel Franco 0002, Emilio Luque |
SBAC-PAD | 4 |
| 2010 | Methodology for Efficient Execution of SPMD Applications on Multicore EnvironmentsabstractThe need to efficiently execute applications in heterogeneous environments is a current challenge for parallel computing programmers. The communication heterogeneities found in multicore clusters need to be addressed to improve efficiency and speedup. This work presents a methodology developed for SPMD applications, which is centered on managing communication heterogeneities and improving system efficiency on multicore clusters. The methodology is composed of three phases: characterization, mapping strategy, and scheduling policy. We focus on SPMD applications which are designed through a message-passing library for communication, and selected according to their synchronicity and communications volume. The novel contribution of this methodology is it determines the approximate number of cores necessary to achieve a suitable solution with a good execution time, while the efficiency level is maintained over a threshold defined by users. Applying this methodology gave results showing a maximum improvement in efficiency of around 43% in the SPMD applications tested. Ronal Muresano, Dolores Rexachs, Emilio Luque |
CCGRID | 3 |
| 2010 | A reconfigurable cache memory with heterogeneous banksabstractThe optimal size of a large on-chip cache can be different for different programs: at some point, the reduction of cache misses achieved when increasing cache size hits diminishing returns, while the higher cache latency hurts performance. This paper presents the Amorphous Cache (AC), a reconfigurable L2 on-chip cache aimed at improving performance as well as reducing energy consumption. AC is composed of heterogeneous sub-caches as opposed to common caches using homogenous sub-caches. The sub-caches are turned off depending on the application workload to conserve power and minimize latencies. A novel reconfiguration algorithm based on Basic Block Vectors is proposed to recognize program phases, and a learning mechanism is used to select the appropriate cache configuration for each program phase. We compare our reconfigurable cache with existing proposals of adaptive and non-adaptive caches. Our results show that the combination of AC and the novel reconfiguration algorithm provides the best power consumption and performance. For example, on average, it reduces the cache access latency by 55.8%, the cache dynamic energy by 46.5%, and the cache leakage power by 49.3% with respect to a non-adaptive cache. Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
DATE | 4 |
| 2010 | Extraction of Parallel Application Signatures for Performance PredictionabstractPredicting performance of parallel applications is becoming increasingly complex and the best performance predictor is the application itself, but the time required to run it thoroughly is a onerous requirement. We seek to characterize the behavior of message-passing applications on different systems by extracting a signature which will allow us to predict what system will allow the application to perform best. To achieve this goal, we have developed a method we called Parallel Application Signatures for Performance Prediction (PAS2P) that strives to describe an application based on its behavior. Based on the application's message-passing activity, we have been able to identify and extract representative phases, with which we created a Parallel Application Signature that has allowed us to predict the application's performance. We have experimented with different signature-extraction algorithms and found a reduction in the prediction error using different scientific applications on different clusters. We were able to predict execution times with an average accuracy of over 98%. Alvaro Wong, Dolores Rexachs, Emilio Luque |
HPCC | 3 |
| 2010 | A Performance Tuning Strategy for Complex Parallel ApplicationabstractDefining performance models associated with the application structure has been proven a useful strategy for implementing dynamic tuning tools. However, for extending this strategy to more complex applications (those composed by different structures) it must integrate a policy for the distribution of the resources among the different application components. Consequently, we propose to take advantage of the knowledge of these models and combine them with a resource management policy for obtaining a global model. In this sense, this work constitutes the ongoing effort in the development of performance models for dynamic tuning. Jose Alexander Guevara, Eduardo César, Joan Sorribes, Andreu Moreno, Tomàs Margalef, Emilio Luque |
PDP | 6 |
| 2010 | FT-DRB: A Method for Tolerating Dynamic Faults in High-Speed Interconnection NetworksabstractThe intensive and continuous use of high-performance computing systems for executing computationally intensive applications, coupled with the large number of elements that make them up, dramatically increase the likelihood of failures during their operation. The interconnection network is a critical part of such systems, therefore, network faults have an extremely high impact because most routing algorithms are not designed to tolerate faults. In such algorithms, just a single fault may stall messages in the network, preventing the finalization of applications, or may lead to deadlocked configurations. This paper introduces a novel fault-tolerant routing method provided with a new deadlock avoidance technique designed to solve an unbounded number of faults appearing at random during system operation. Our method provides escape paths for the stalled messages. In addition, the routing algorithm configures alternative paths to avoid the faulty areas taking advantage of communication path redundancy by means of multipath routing approaches. Deadlock avoidance is achieved by adding a small-sized queue and applying a simple set of actions when accessing output buffers with limited free space. Experiments show that our method allows applications to successfully finalize their execution in the presence of several number of faults, with an average performance value of 96% compared to the fault-free scenarios. Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
PDP | 4 |
| 2010 | Deadlock Avoidance for Interconnection Networks with Multiple Dynamic FaultsabstractThe intensive and continuous use of high-performance computing systems for executing computationally intensive applications, coupled with the large number of elements that make them up, dramatically increase the likelihood of failures during their operation. Clearly, network faults have an extremely high impact because most routing algorithms are not designed to tolerate faults. In such algorithms, just a single fault may lead to deadlocked configurations thus preventing the correct finalization of applications. This paper introduces a new deadlock avoidance mechanism for routing algorithms designed to deal with multiple dynamic faults. The mechanism is based on adding a small-sized buffer and applying a simple set of actions when accessing output buffers with limited free space. Unlike typical static solutions, this proposal allows the design of routing algorithms capable of treating an unbounded number of dynamic faults. Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
PDP | 4 |
| 2010 | Scalable dynamic Monitoring, Analysis and Tuning Environment for parallel applications
Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque |
J. Parallel Distributed Comput. | 4 |
| 2010 | Designing an effective P2P system for a VoD system to exploit the multicast communication
Xiaoyuan Yang 0001, Fernando Cores, Porfidio Hernández, Ana Ripoll, Emilio Luque |
J. Parallel Distributed Comput. | 5 |
| 2009 | Dynamic and Distributed Multipath Routing Policy for High-Speed Cluster NetworksabstractThe increasing demand of parallel applications in cluster computing requires the use of interconnection networks to provide low and bounded communication delays. However, message congestion appears when communication load between nodes is not fairly distributed over the network. Congestion spreading increases latency and reduces network throughput causing important performance degradation. In this paper we present dynamic routing balancing with multipath distribution (DRB-MD), a new method developed to control network congestion based on a uniform balancing of communication load. DRB-MD distributes the traffic load according to a gradual and load-controlled path expansion. It monitors message latency in network switches, makes decisions about how many alternative paths should be used, and finally decides which path (or paths) to use between each source-destination pair. Experiments with permutation patterns and hotspot traffic were conducted to evaluate DRB-MD performance under conditions commonly created by parallel scientific applications. Diego Lugones, Daniel Franco 0002, Emilio Luque |
CCGRID | 3 |
| 2009 | An assessment of multi-core for a performance prediction model of tomographic reconstructionabstractThree-dimensional (3D) reconstruction of structures from projection data is essential for helping people in a wide range of areas. Algebraic reconstruction techniques (ART) are iterative procedures for recovering the structure of the 3D objects from projection images. During the seventies, the ART were dismissed due to high-demanding computing requirements. Interesting recent research aims at acquiring experience with parallelization strategies and at demonstrating the effectiveness of the massively parallel processing approach in 3D reconstructions. Multi-core (MC) technology provides new levels of performance and, therefore, it is of paramount importance to make performance predictions. The objectives of this work are both adapting an analytical performance prediction model for the iterative reconstruction techniques (IRT) to a MC environment and finding a process's CPU affinity that produces the best overall performance. BPTomo+is a parallel distributed application for tomographic reconstruction that uses IRT. Besides, it includes a process's CPU affinity mask. The analytical performance prediction model is validated by comparison of the estimated times for representative datasets against BPTomo+computation times measured on a MC server. The analytical model is shown to be quite accurate. The percentage of deviation between estimated and measured times is less than 6%. Paula Cecilia Fritzsche, Ronal Muresano, Dolores Rexachs, Emilio Luque |
CLUSTER | 4 |
| 2009 | Fast-Response Dynamic Routing Balancing for high-speed interconnection networksabstractCommunication requirements in High Performance Computing systems demand the use of high-speed Interconnection networks to connect processing nodes. However, when communication load is unfairly distributed across the network resources, message congestion appears. Congestion spreading increases latency and reduces network throughput causing important performance degradation. The Fast-Response Dynamic Routing Balancing (FR-DRB) is a method developed to perform a uniform balancing of communication load over the interconnection network. FR-DRB distributes the message traffic based on a gradual and load-controlled path expansion. The method monitors network message latency and makes decisions about the number of alternative paths to be used between each source-destination pair for message delivery. FR-DRB performance has been compared with other routing policies under a representative set of traffic patterns which are commonly created by parallel scientific applications. Experiments results show an important improvement in latency and throughput. Diego Lugones, Daniel Franco 0002, Emilio Luque |
CLUSTER | 3 |
| 2009 | How SPMD applications could be efficiently executed on multicore environments?abstractA challenge for programmers of parallel programming environments is to execute applications efficiently. For this reason, applications with high levels of synchronism and communications such as SPMD (single program multiple data) create a challenge regarding how to distribute tasks between PE (processing element) in a multicore cluster; this kind of environment presents high heterogeneity in communication parameters due to different communication paths present. For this reason, this work is centered around developing a methodology to distribute SPMD tasks between PEs in a multicore cluster. The task assignment process is realized through mapping and scheduling strategies based on controlling the communications heterogeneities. Finally, the objective is to obtain a good execution time while maintaining the efficiency level over a threshold. The results obtained show an improvement around 40% of efficiency in a heat transfer application, when our methodology is applied. Ronal Muresano, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2009 | Increasing the availability provided by RADIC with low overheadabstractFor machines composed of a large number of processing units, fault probability tends to increase linearly with this number. This makes the use of a fault tolerant solution a major issue. A fault tolerant solution provides certain level of availability, which is usually influenced by time overhead, performance degradation, resources or cost. In the rollback-recovery protocol, the availability increase is usually achieved by increasing the checkpoint frequency or by making several replicas of checkpoints and/or logs. Such a replication allows the solution to tolerate concurrent correlated faults, i.e., a fault in a computing node and in the stable storage. These faults are theoretically less probable, however recent studies have shown that faults are temporally and spatially correlated, consequently increasing the concurrent fault probability. The major concern replicating the checkpoints and logs is the overhead caused by storing these replicas over various repositories, which may disallow its use. In this paper we present how we increased the availability provided by RADIC, without significantly increase of its overhead. Our approach consists of parallelizing the storing of these replicas using the pipeline technique. Such a technique allows us to make low-overhead copies of checkpoints and logs over N protectors. Furthermore, as secondary benefit, the pipelining between observer and protector reduces more than four times (in the best case) the pessimistic message logging overhead. Guna Santos, Leonardo Fialho, Dolores Rexachs, Emilio Luque |
CLUSTER | 4 |
| 2009 | Parallel application signatureabstractWe seek to achieve characterization or application signature from a parallel application that will allow us, through the execution of this signature, to evaluate its performance in different computers. Sequential applications behavior can be understood by means of tools such as SimPoint. This tool can identify and select significant phases describing the applications behavior. Our proposal is to extend those concepts towards parallel applications, with the goal of modeling and predicting the parallel application. To achieve this, we developed a methodology, enabling us to identify and extract repetitive behavior to create the application signature. We have validated our proposal using scientific applications such as the NAS Parallel Benchmarks, Sweep3D. We could predict the execution time of the entire application. Alvaro Wong, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2009 | A Multipath Fault-Tolerant Routing Method for High-Speed Interconnection Networks
Gonzalo Zarza, Diego Lugones, Daniel Franco 0002, Emilio Luque |
Euro-Par | 4 |
| 2009 | Models for high-speed interconnection networks performance analysisabstractModeling Interconnection networks is an important research topic enabling the study of the interconnection behavior and its significance in telecommunication applications and distributed systems. However, complexity of large-scale networks makes development of models and simulation tools a prohibitively difficult task. In this paper we have explored the network modeling space design to provide models following two different approaches: accurate simulation models based on finite state machines (FSM), and also, analytical models to provide profitable speedup with a minimal accuracy loss. Experiments results show that the proposed analytical model provides a faithful abstraction for the scale of systems that are of interest in the foreseeable future, it reaches an 8% error and speedup of around 30× vs. a FSM model. Diego Lugones, Daniel Franco 0002, Eduardo Argollo, Emilio Luque |
MASCOTS | 4 |
| 2009 | Task distribution using factoring load balancing in Master-Worker applications
Andreu Moreno, Eduardo César, Joan Sorribes, Tomàs Margalef, Emilio Luque |
Inf. Process. Lett. | 5 |
| 2008 | Increasing the Performability of Computer Clusters Using RADIC IIabstractPerformance and availability form an undissociable binomial for some kind of applications. Therefore, the fault tolerant solutions must take into consideration these two constraints when it has been designed. Our previous work, called RADIC, implemented a basic level protection allowing to recover from faults just using the active cluster resources, changing the system configuration. However, Such approach may genenerate some performance degradation in some cases. In this paper, we present RADIC II, which incorporates a new protection level using dynamic redundancy, allowing to mitigate or avoid the recovery side-effects. Such functionality allows restoring a changed system configuration and it can avoid the configuration changes. The results has shown that RADIC-II operates correctly and becomes itself as a good approach to provide high availability to the parallel applications without suffer a system degradation in post-recovery execution. Guna Santos, Angelo Duarte, Dolores Rexachs, Emilio Luque |
ARES | 4 |
| 2008 | Performance models for dynamic tuning of parallel applications on Computational GridsabstractPerformance is a main issue in parallel application development. Dynamic tuning is a technique that acts over application parameters to raise execution performance indexes. To perform that, it is necessary to collect measurements, analyze application behavior using a performance model and carry out tuning actions. Computational Grids present proclivity for dynamic changes on their features during application execution. Thus, dynamic tuning tools are indispensable to reach the expected performance indexes on those environments. A particular problem which provokes performance bottlenecks is the load distribution in master/worker applications. This paper addresses the performance modeling of such applications on Computational Grids for the perspective of dynamic tuning. It is inferred that grain size and number of workers are critical parameters to reduce execution time while raising the efficiency of resources usage. A heuristic to dynamically tune granularity and number of workers is proposed. The experimental simulated results of a matrix multiplication application in a heterogeneous Grid environment are shown. Genaro Costa, Josep Jorba 0001, Anna Sikora, Tomàs Margalef, Emilio Luque |
CLUSTER | 5 |
| 2008 | On-Line Performance Modeling for MPI Applications
Oleg Morajko, Anna Sikora, Tomàs Margalef, Emilio Luque |
Euro-Par | 4 |
| 2008 | Dynamic Pipeline Mapping (DPM)
Andreu Moreno, Eduardo César, Andreu Guevara, Joan Sorribes, Tomàs Margalef, Emilio Luque |
Euro-Par | 6 |
| 2008 | Performance Model for Parallel Mathematical Libraries Based on Historical Knowledgebase
Ihab Salawdeh, Eduardo César, Anna Sikora, Tomàs Margalef, Emilio Luque |
Euro-Par | 5 |
| 2008 | Providing Non-stop Service for Message-Passing Based Parallel Applications with RADIC
Guna Santos, Angelo Duarte, Dolores Rexachs, Emilio Luque |
Euro-Par | 4 |
| 2008 | A General Approach to Predict the Performance Order of TSP Family Problems
Paula Cecilia Fritzsche, Dolores Rexachs, Emilio Luque |
ICA3PP | 3 |
| 2008 | Increasing the Scalability and the Speedup of a Fish School Simulator
Christianne Dalforno, Diego Mostaccio, Remo Suppi, Emilio Luque |
ICCSA (2) | 4 |
| 2008 | Software Probes: Towards a Quick Method for Machine Characterization and Application Performance PredictionabstractComputers perform different applications in different ways. To characterize an application performance into a machine, the usual method is a throughout execution of it. This work is a step into a synthetic probe able to characterize a master-worker application's performance in a fraction of the time required to run it entirely. This is specially important for CPU-intensive scientific applications, who runs for very long, as it makes sense that it runs as efficiently (and fast) as possible. To know how, and for how long a master-worker application is going to run can guide the decision to use this machine or not. Our software probe takes into account only the performance-relevant parts of the application, discovering a program's relevant phases. Running solely these significant phases is a powerful way to quickly characterize the application's performance on a machine. It can help to select the best computing nodes in a grid or in a multi-cluster to run this application, and even quickly predict the total execution time for this application/data set in the machine analyzed. We also present ongoing work on a fully synthetic probe generated from programs' phases. Alexandre Otto Strube, Dolores Rexachs, Emilio Luque |
ISPDC | 3 |
| 2007 | Automatic Generation of Dynamic Tuning Techniques
Paola Caymes-Scutari, Anna Sikora, Tomàs Margalef, Emilio Luque |
Euro-Par | 4 |
| 2007 | Applying Data Mining to Define TSP Asymptotic Time ComplexityabstractComputational science is often referred to as the third science,complementing both theoretical and laboratory science. In this field, new challenges are continuously arising. The asymptotic time complexity definition of both deterministic and non-deterministic algorithms to solve all kinds of problems is one of the key points in computer science. Knowing the limit of the execution time of an algorithm when the size of the problem goes to infinity is essential. In particular, data-dependent applications is an extremely challenging problem because for a specific issue the input data sets may cause variability in execution times. The development of an entire approach to define the asymptotic time complexity of a hard data-dependent parallel application that solves the traveling salesman problem (TSP) is the focus of this study. Two different parallel TSP algorithms are presented. One of these is used to show the usefulness and the profits of the proposed approach, and the other one is used as witness. The experimental results are quite promising. Paula Cecilia Fritzsche, Dolores Rexachs, Emilio Luque |
ICTAI (2) | 3 |
| 2007 | Improving Web Services Interoperability with Binding ExtensionsabstractCurrent Web services are able to interoperate successfully with most basic data types. However, due to the limited functionality of existing data binding tools, they still experiment difficulties manipulating more complex XML data types, forcing programmers to work at the XML level. In this paper we propose a business model for web services where data binding tools not only generate the WSDL, but also provide portable binding extensions for manipulating theXSD types. These binding extensions can be integrated into any other binding tool, overcoming their limitations. Because these extensions are written in XML, this model conforms to the principle of platform independence. In addition, this model does not suppose any extra programming effort to neither service providers or clients. The approach has been validated with the creation of several extensions, which has been ported into Java and PHP clients. Our preliminary results show no performance penalty. Enric Jaén Villoldo, Joan Serrat 0001, Emilio Luque |
ICWS | 3 |
| 2007 | Functional Tests of the RADIC Fault Tolerance ArchitectureabstractClusters with thousand of nodes are a reality and the current trend indicates that they are becoming larger. Such large clusters are subject to a relatively high fault frequency so a fault-tolerance scheme is mandatory to assure the correct application completion. Message passing is the programming model often used in large clusters and the current implementations used to achieve fault tolerance in message passing systems do not focus in an architecture that simultaneously attends to scalability, transparency and independence of stable/central elements. The RADIC architecture was proposed and design as a fully distributed structure in order to achieve such requirements. Such architecture defines a fully distributed fault tolerance controller implemented by a set of system processes, which collaborate in order to perform all the basic functions of a fault tolerance protocol. This paper presents the test methodology used to verify the functionality of the RADIC architecture using RADICMPI, a prototype on the MPI semantic Angelo Duarte, Dolores Rexachs, Emilio Luque |
PDP | 3 |
| 2007 | MATE: Monitoring, Analysis and Tuning Environment for parallel/distributed applicationsabstractAbstract The main goal of parallel/distributed applications is to solve the considered problem as fast as possible using the available resources. In this context, the application performance becomes a crucial issue. Developers of these applications must optimize them if they are to fulfill the promise of high‐performance computation. To improve performance, developers search for bottlenecks by analyzing application behavior, try to identify performance problems, determine their causes and overcome them by changing the source code of the application. Current approaches require developers to do these tasks manually and imply a high degree of expertise. Therefore, another approach is needed to help developers during the optimization process. This paper presents the dynamic tuning approach that addresses these issues. In this approach, many tasks are automated and the user intervention and required experience may be significantly reduced. An application is monitored, its performance bottlenecks are detected and it is modified automatically during execution, without recompiling or re‐running it. The introduced modifications adapt the application behavior to changing conditions. We present an environment called MATE (Monitoring, Analysis and Tuning Environment) that has been developed to provide dynamic tuning of parallel/distributed applications. We also show practical experiments conducted with MATE to prove its effectiveness and profitability. Copyright © 2006 John Wiley & Sons, Ltd. Anna Sikora, Paola Caymes-Scutari, Tomàs Margalef, Emilio Luque |
Concurr. Comput. Pract. Exp. | 4 |
| 2007 | Cooperating CoScheduling: A Coscheduling Proposal Aimed at Non-Dedicated Heterogeneous NOWs
Francesc Giné, Francesc Solsona Tehàs, Mauricio Hanzich, Porfidio Hernández, Emilio Luque |
J. Comput. Sci. Technol. | 5 |
| 2007 | Design and implementation of a dynamic tuning environment
Anna Sikora, Tomàs Margalef, Emilio Luque |
J. Parallel Distributed Comput. | 3 |
| 2007 | The Convergence of Realistic Distributed Load-Balancing Algorithms
F. Cedo, Ana Cortés, Ana Ripoll, Miquel A. Senar, Emilio Luque |
Theory Comput. Syst. | 5 |
| 2006 | Increasing the cluster availability using RADICabstractThe redundant array of distributed independent checkpoints (RADIC) is a fault tolerant architecture based on a fully distributed array of dedicated process. These processes collaborate to create a fault tolerance controller which transparently manages all fault tolerance activities. The architecture is designed as a software layer between the application and the cluster structure and it was developed to attend to the requirements of scalability, user transparency and independency of dedicated/stable cluster resources. RADIC only requires the resources already available in the nodes used by the parallel application and it uses a pessimistic message-log rollback-recovery protocol in order to operate without any global synchronization. Such protocol, together with the independence of central elements, makes RADIC a scalable architecture that works transparently to the user. We tested the functionality and performance of the architecture in a real scenario using a prototype based on the MPI standard (RADICMPI) Angelo Duarte, Dolores Rexachs, Emilio Luque |
CLUSTER | 3 |
| 2006 | A Performance Prediction Methodology for Data-dependent Parallel ApplicationsabstractThe increase in the use of parallel distributed architectures in order to solve large-scale scientific problems has generated the need for performance prediction for both deterministic applications and non-deterministic applications. In particular, the performance prediction of data-dependent programs is an extremely challenging problem because for a specific issue the input datasets may cause different execution times. Generally, a parallel application is characterized as a collection of tasks and their interrelations. If the application is time-critical it is not enough to work with only one value per task, and consequently knowledge of the distribution of task execution times is crucial. The development of a new prediction methodology to estimate the performance of data-dependent parallel applications is the primary target of this study. This approach makes it possible to evaluate the parallel performance of an application without the need of implementation. A real data-dependent arterial structure detection application model is used to apply the methodology proposed. The predicted times obtained using the new methodology for genuine datasets are compared with predicted times that arise from using only one execution value per task. Finally, the experimental study shows that the new methodology generates more precise predictions Paula Cecilia Fritzsche, Concepció Roig, Ana Ripoll, Emilio Luque, Aura Hernández-Sabaté |
CLUSTER | 4 |
| 2006 | Using Simulation, Historical and Hybrid Estimation Systems for Enhacing Job Scheduling on NOWsabstractThe computation capacity of the workstations in an open laboratory is enough to execute not only the local workload but some distributed computation. Unfortunately, the local workload introduces much uncertainty into the predictability of the system, which hinders the applicability of the job scheduling strategies. In this work, we introduce an estimation engine into our job scheduling system, termed CISNE. This prediction capacity allows us guarantee some limits to the turnaround time of parallel jobs. With this aim, three different estimation methods have been proposed and implemented in the CISNE system: a simulation tool, a historical system and an integration of both (hybrid). In this framework, we have compared our proposals to representative estimation methods in the literature. Likewise, we have analyzed these estimation methods in relation to different scheduling policies. These results reveal that the hybrid method achieves the best performance due to the fact that it combines the flexibility of a simulator to represent such a dynamic system as a non-dedicated cluster together with the accuracy given by the historical information Mauricio Hanzich, Porfidio Hernández, Emilio Luque, Francesc Giné, Francesc Solsona Tehàs, Josep L. Lérida |
CLUSTER | 3 |
| 2006 | Tuning Application in a Multi-cluster Environment
Eduardo Argollo, Adriana Gaudiani, Dolores Rexachs, Emilio Luque |
Euro-Par | 4 |
| 2006 | Exploiting Throughput for Pipeline Execution in Streaming Image Processing Applications
Fernando Guirado, Ana Ripoll, Concepció Roig, Aura Hernández-Sabaté, Emilio Luque |
Euro-Par | 5 |
| 2006 | Using On-the-Fly Simulation for Estimating the Turnaround Time on Non-dedicated Clusters
Mauricio Hanzich, Josep L. Lérida, Matías Torchinsky, Francesc Giné, Porfidio Hernández, Emilio Luque |
Euro-Par | 6 |
| 2006 | Providing VCR in a Distributed Client Collaborative Multicast Video Delivery Scheme
Xiaoyuan Yang 0001, Porfidio Hernández, Fernando Cores, Ana Ripoll, Remo Suppi, Emilio Luque |
Euro-Par | 6 |
| 2006 | Multi-Collaboration Domain Multicast P2P Delivery Architecture for VoD SystemabstractAlthough a P2P scheme based unicast forward or application level multicast (ALM) is able to decentralize the video delivery process to reduce the VoD server load, the performance of such a scheme is limited by the high network resource requirement of the clients' communications. In this paper, we propose a new P2P architecture that uses multicast communications for peer collaborations. We introduce the concept of n-peers to m-peers cooperative collaboration to achieve high client delivery efficiency as well as unbounded scalable client collaboration capacity. Furthermore, the architecture selects peers by taking into account the underlying network infrastructure in order to avoid network saturation. We compared the new design with video delivery architecture based on Patching and Chaining schemes in a synthetic network model as well as a real Ethernet network topology. The experimental results showed that our design is better than Patching and Chaining in terms of admitted client requests and VoD server resource requirement. Furthermore, the results showed that our architecture is adaptable to any underlying physical networks. Xiaoyuan Yang 0001, Porfidio Hernández, Leandro Souza, Ana Ripoll, Remo Suppi, Emilio Luque, Fernando Cores |
ICC | 6 |
| 2006 | Wide and efficient trace prediction using the local trace predictorabstractHigh prediction bandwidth enables performance improvements and power reduction techniques. This paper explores a mechanism to increase prediction width (instructions per prediction) by predicting instruction traces. Our analysis shows that predicting traces including multiple branches is not significantly less accurate than predicting single branches. A novel Local Trace Predictor organization is proposed. It increases prediction width without reducing the ratio of prediction accuracy versus memory resources with respect to a Basic Block Predictor.Compared to the previously proposed Next-Trace Predictor, the Local Trace Predictor reduces memory requirements by codifying trace predictions, and by limiting the number of traces starting at the same instruction to 2 or 4. The limit lessens prediction width only slightly, and does not affect prediction accuracy. The overall result is that the Local Trace Predictor outperforms the Next-Trace Predictor for sizes higher than 12 KBytes. Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
ICS | 4 |
| 2006 | DVoDP2P: distributed P2P assisted multicast VoD architectureabstractFor a high scalable VoD system, the distributed server architecture (DVoD) with more than one server-node is a cost-effective design solution. However, such a design is highly vulnerable to workload variations because the service capacity is limited. In this paper, we propose a new and efficient VoD architecture that combines DVoD with a P2P system. The DVoD's server-nodes is able to offer a minimum required quality of service (QoS) and the P2P system is able to provide the mechanism to increase the system service capacity according to client demands. Our P2P system is able to synchronize a group of clients in order to replace server-nodes in the delivery process. We compared the new VoD architecture with DVoD architecture based on classic multicast and P2P delivery policies (patching and chaining). The experimental results showed that our design is better than previous solutions, providing higher scalability Xiaoyuan Yang 0001, Porfidio Hernández, Fernando Cores, Leandro Souza, Ana Ripoll, Remo Suppi, Emilio Luque |
IPDPS | 7 |
| 2006 | MetaLoRaS: A Predictable MetaScheduler for Non-dedicated Multiclusters
Josep L. Lérida, Francesc Solsona Tehàs, Francesc Giné, Mauricio Hanzich, Porfidio Hernández, Emilio Luque |
ISPA | 6 |
| 2006 | On the Relevance of Network Topologies in Distributed Video-on-Demand ServersabstractDistributed video-on-demand servers (DVS) are proposed as a solution to the limited streaming capacity and null scalability of large-scale centralized systems. Server interconnection topology plays an important role in video-on-demand systems' performance. This paper presents an analysis of different topologies and their influence over storage management and distribution, delivery policies performance, refusing requests occurrence, network consumption and scalability. To accomplish the proposal study, we have designed a complete simulation framework for DVS systems. Experimental results obtained under different workload conditions allow us to draw two important conclusions: First, a better connectivity implies a lower mean request service distance and lesser network requirements, improving multicast policies efficiency. Second, topology regularity is essential, as it allows a greater traffic balancing and provides more alternative routing paths. The analysis of global results shows that hypercube presents the best trade-off among all the evaluated metrics, providing a gradual and unlimited scalability for the DVS system. Leandro Souza, Ana Ripoll, Xiaoyuan Yang 0001, Emilio Luque, Fernando Cores |
PDP | 4 |
| 2006 | Modeling Master/Worker applications for automatic performance tuning
Eduardo César, Andreu Moreno, Joan Sorribes, Emilio Luque |
Parallel Comput. | 4 |
| 2005 | Optimizing Latency under Throughput Requirements for Streaming Applications on Cluster ExecutionabstractParallelism in applications that act on a stream of input data can be exploited with two different approaches, spatial and temporal. In this paper we propose a new task mapping algorithm, called EXPERT, to exploit temporal parallelism efficiently when the streaming application is running in a pipeline fashion. We compare the performance of spatial and temporal approaches, in terms of latency and throughput for a video compression application. The results show that the pipeline execution with the task assignment provided by EXPERT algorithm, significantly overcomes spatial parallelism. Additionally, this temporal parallelism presents better scalability results when the dimension of the problem is augmented Fernando Guirado, Ana Ripoll, Concepció Roig, Emilio Luque |
CLUSTER | 4 |
| 2005 | Modeling Pipeline Applications in POETRIES
Eduardo César, Joan Sorribes, Emilio Luque |
Euro-Par | 3 |
| 2005 | CISNE: A New Integral Approach for Scheduling Parallel Applications on Non-dedicated Clusters
Mauricio Hanzich, Francesc Giné, Porfidio Hernández, Francesc Solsona Tehàs, Emilio Luque |
Euro-Par | 5 |
| 2005 | Topic 13 Routing and Communication in Interconnection Networks
Emilio Luque, Cruz Izu, Olav Lysne, José Legatheaux Martins |
Euro-Par | 1 |
| 2005 | Automatic Tuning of Master/Worker Applications
Anna Sikora, Eduardo César, Paola Caymes-Scutari, Tomàs Margalef, Joan Sorribes, Emilio Luque |
Euro-Par | 6 |
| 2005 | Target Encoding for Efficient Indirect Jump Prediction
Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
Euro-Par | 4 |
| 2005 | Dynamic Distributed Collaborative Merging Policy to Optimize the Multicasting Delivery Scheme
Xiaoyuan Yang 0001, Porfidio Hernández, Fernando Cores, Ana Ripoll, Remo Suppi, Emilio Luque |
Euro-Par | 6 |
| 2005 | Performance and Power Evaluation of an Intelligently Adaptive Data Cache
Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
HiPC | 4 |
| 2005 | Is evolution or revolution the way for improving the teaching methodology in computer science?abstractNo abstract available. Emilio Luque |
ITiCSE | 1 |
| 2005 | Enhancing wildland fire prediction on cluster systems applying evolutionary optimization techniques
Baker Abdalhaq, Ana Cortés, Tomàs Margalef, Emilio Luque |
Future Gener. Comput. Syst. | 4 |
| 2004 | Topic 3: Scheduling and Load Balancing
Emilio Luque, José G. Castaños, Evangelos P. Markatos, Raffaele Perego 0001 |
Euro-Par | 1 |
| 2004 | MATE: Dynamic Performance Tuning Environment
Anna Sikora, Oleg Morajko, Tomàs Margalef, Emilio Luque |
Euro-Par | 4 |
| 2004 | Supporting Caching and Mirroring in Distributed Video-on-Demand Architectures
Xiaoyuan Yang 0001, Fernando Cores, Ana Ripoll, Porfidio Hernández, Bahjat Qazzaz, Remo Suppi, Emilio Luque |
Euro-Par | 7 |
| 2004 | Modeling Master-Worker Applications in POETRIESabstractParallel/distributed application development is a very difficult task for non-expert programmers, and therefore support tools are needed for all phases of this kind of application development cycle. This means that developing applications using predefined programming structures (frameworks) should be easier than doing it from scratch. We propose to take advantage of the knowledge about the structure of the application in order to develop a dynamic and automatic tuning tool. In this sense, we have designed POETRIES, which is a dynamic performance tuning tool based on the idea that a performance model could be associated to the high-level structure of the application. This way, the tool could efficiently make better tuning decisions. Specifically, we focus this work on the definition of the performance model associated to applications developed with the master-worker framework. Eduardo César, José G. Mesa, Joan Sorribes, Emilio Luque |
HIPS | 4 |
| 2004 | Graduate students learning strategies through research collaborationabstractIt is already known that the learning process can be accelerated with the mixture of theoretical classes and experimental work. This paper describes an interesting experiment with that combination in the teaching of computer architecture for Ph.D. students in collaboration with a researcher in a real design investigation. As the work progressed, a simple cyclical methodology arose as reference for future works. Eduardo Argollo, Mauricio Hanzich, Diego Mostaccio, Germán Bianchini, Paula Cecilia Fritzsche, Ferran Bonàs, Emilio Luque, Juan C. Moure, Dolores Rexachs |
ITiCSE | 7 |
| 2004 | Efficient resource management applied to master worker applications
Elisa Heymann, Miquel A. Senar, Emilio Luque, Miron Livny |
J. Parallel Distributed Comput. | 3 |
| 2003 | POETRIES: Performance Oriented Environment for Transparent Resource-Management, Implementing End-User Parallel/Distributed Applications
Eduardo César, José G. Mesa, Joan Sorribes, Emilio Luque |
Euro-Par | 4 |
| 2003 | Exploiting Traffic Balancing and Multicast Efficiency in Distributed Video-on-Demand Architectures
Fernando Cores, Ana Ripoll, Bahjat Qazzaz, Remo Suppi, Xiaoyuan Yang 0001, Porfidio Hernández, Emilio Luque |
Euro-Par | 7 |
| 2003 | Cooperating Coscheduling in a Non-dedicated Cluster
Francesc Giné, Francesc Solsona Tehàs, Porfidio Hernández, Emilio Luque |
Euro-Par | 4 |
| 2003 | Predicting the Best Mapping for Efficient Exploitation of Task and Data Parallelism
Fernando Guirado, Ana Ripoll, Concepció Roig, Xiao Yuan 0003, Emilio Luque |
Euro-Par | 5 |
| 2003 | Optimizing a Decoupled Front-End Architecture: The Indexed Fetch Target Buffer (iFTB)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 3 |
| 2003 | Admission control policies for video on demand brokersabstractA VOD server has to employ admission control algorithms to decide whether a new client (request) can be admitted without violating the requirements of clients who are already being serviced. This paper presents and analyzes four specific admission control algorithms pertaining to the deterministic and predictive category. These algorithms are: maximum, adaptive maximum, average and adaptive average. They are studied in detail for CPU and network resources. The adaptive algorithms are our own contribution and we show that they use an exponential weighted average to guess the future resource needs of a stream, rather than just using the maximum or mean over the whole period. Thus, they provide smoothing and so a closer matching of resources used to those reserved and so higher capacity and less waiting time. Bahjat Qazzaz, Xiaoyuan Yang 0001, Porfidio Hernández, Remo Suppi, Emilio Luque |
ICME | 6 |
| 2003 | Clustering and reassignment-based mapping strategy for message-passing architectures
Miquel A. Senar, Ana Ripoll, Ana Cortés, Emilio Luque |
J. Syst. Archit. | 4 |
| 2002 | Optimization of Fire Propagation Model Inputs: A Grand Challenge Application on Metacomputers (Research Note)
Baker Abdalhaq, Ana Cortés, Tomàs Margalef, Emilio Luque |
Euro-Par | 4 |
| 2002 | Double P-Tree: A Distributed Architecture for Large-Scale Video-on-Demand
Fernando Cores, Ana Ripoll, Emilio Luque |
Euro-Par | 3 |
| 2002 | Adjusting Time Slices to Apply Coscheduling Techniques in a Non-dedicated NOW (Research Note)
Francesc Giné, Francesc Solsona Tehàs, Porfidio Hernández, Emilio Luque |
Euro-Par | 4 |
| 2002 | Speeding Up Target Address Generation Using a Self-indexed FTB (Research Note)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 3 |
| 2002 | Parasite: Distributing Processing Using Java Applets (Research Note)
Remo Suppi, Marc Solsona, Emilio Luque |
Euro-Par | 3 |
| 2002 | AMEEDA: A General-Purpose Mapping Tool for Parallel Applications on Dedicated Clusters (Research Note)
Xiao Yuan 0003, Concepció Roig, Ana Ripoll, Miquel A. Senar, Fernando Guirado, Emilio Luque |
Euro-Par | 6 |
| 2002 | Cost-effective distributed architecture for large-scale video-on-demandabstractIn spite of the attractiveness of video-on-demand (VoD) services, their implementation up to the present has not been as widespread as could have been desired, due to centralized VoD systems having a limited streaming capacity and poor scaling. One-level proxy-based systems have been proposed to increase system capacity, but their scalability is still limited by the main net bandwidth. To achieve a scalable large-scale system (LVoD), we propose a hierarchical architecture, called double P-tree, based on a tree topology of independent local nets with proxies. To decentralize the architecture and facilitate the system's growth, the functionality of the proxy has been modified in such a way that it works at the same time as caching for the most-watched videos, and as a distributed mirror for the remaining videos. Moreover, the double P-tree topology interconnects local networks in order to improve connectivity and efficiency. The evaluation of this new architecture, through an analytical model, shows that the double P-tree architecture is a good approach to the design of flexible LVoD systems, without the need for using a complex server or high-bandwidth networks. Fernando Cores, Ana Ripoll, Emilio Luque |
ICME (1) | 3 |
| 2002 | The KScalar simulatorabstractModern processors increase their performance with complex microarchitectural mechanisms, which makes them more and more difficult to understand and evaluate. KScalar is a graphical simulation tool that facilitates the study of such processors. It allows students to analyze the performance behavior of a wide range of processor microarchitectures: from a very simple in-order, scalar pipeline, to a detailed out-of-order, superscalar pipeline with non-blocking caches, speculative execution, and complex branch prediction. The simulator interprets executables for the Alpha AXP instruction set: from very short program fragments to large applications. The object's program execution may be simulated in varying levels of detail: either cycle-by-cycle, observing all the pipeline events that determine processor performance, or million cycles at once, taking statistics of the main performance issues.Instructors may use KScalar in several ways. First, it may be used to provide demonstrations in lectures or online learning environments. Second, it allows students to investigate the characteristics of specific processor microarchitectures as practical short assignments associated to a lecture course. Third, students may undertake major projects involving the optimization of real programs at the software-hardware interface, or involving the optimization of a processor microarchitecture for a given application workload.A preliminary version of KScalar has been successfully used in several lecture courses during the last two years in the University Autónoma of Barcelona. It runs on a x86/Linux/KDE system. The graphical interface has been developed using the KDE and QT libraries. The simulator engine running behind the graphical interface is a heavily-modified version of SimpleScalar. KScalar code is available under the terms of the GNU and SimpleScalar General Public License Juan C. Moure, Dolores Rexachs, Emilio Luque |
ACM J. Educ. Resour. Comput. | 3 |
| 2002 | An asynchronous and iterative load balancing algorithm for discrete load model
Ana Cortés, Ana Ripoll, F. Cedo, Miquel A. Senar, Emilio Luque |
J. Parallel Distributed Comput. | 5 |
| 2001 | Evaluation of Strategies to Reduce the Impact of Machine Reclaim in Cycle-Stealing EnvironmentsabstractWe investigate the scheduling problem that arises in parallel applications executing on a network of machines by using a mode of cycle-stealing. In this mode of execution a parallel application executes its tasks in several machines whenever they are idle. When the user reclaims the machine, tasks must relinquish control immediately. In this case, the parallel application has the risk of losing work in progress on reclaimed machines and, therefore, the total execution time of the parallel application will be affected by the need for rescheduling the pre-empted task. We first evaluate the impact on the performance of an application when it runs on two different scenarios: a set of N dedicated machines, and a set of N non-dedicated machines (in which pre-emption may occur). This study shows that losing machines may have a considerable impact on the execution time of the application and therefore, we propose and evaluate three simple strategies to alleviate this problem. All strategies are based on the use of additional machines, but they differ in the way that these extra machines are used. In the first strategy additional machines are added to the common pool of machines used by the application. The other two are based on task replication, in which the additional machines are used to execute certain tasks that are already running in other machines. Elisa Heymann, Miquel A. Senar, Emilio Luque, Miron Livny |
CCGRID | 3 |
| 2001 | Improving Single-Thread Fetch Performance on a Multithreaded ProcessorabstractMultithreaded processors, by simultaneously using both the thread-level parallelism and the instruction-level parallelism of applications, achieve larger instruction per cycle rate than single-thread processors. On a multi-thread workload, a clustered organization maximizes performances. On a single-thread workload, however, all but one of the clusters are idle, degrading single-thread performance significantly. Using a clustered multi-thread performance as a baseline, we propose and analyze several mechanisms and policies to improve single-thread execution exploiting the existing hardware without a significant multi-thread performance loss. We focus on the fetch unit, which is maybe the most performance-critical stage. Essentially, we analyze three ways of exploiting the idle fetch clusters: allowing a single thread accessing its neighbor clusters, use the idle fetch clusters to provide multiple-path execution, or use them to widen the effective single-three fetch block. Juan C. Moure, R. B. García, Dolores Rexachs, Emilio Luque |
DSD | 4 |
| 2001 | Self-Adjusting Scheduling of Master-Worker Applications on Distributed Clusters
Elisa Heymann, Miquel A. Senar, Emilio Luque, Miron Livny |
Euro-Par | 3 |
| 2001 | Dynamic Performance Tuning Environment
Anna Sikora, Eduardo César, Tomàs Margalef, Joan Sorribes, Emilio Luque |
Euro-Par | 5 |
| 2001 | Predictive Coscheduling Implementation in a Non-dedicated Linux Cluster
Francesc Solsona Tehàs, Francesc Giné, Porfidio Hernández, Emilio Luque |
Euro-Par | 4 |
| 2001 | CMC: A Coscheduling Model for non-Dedicated Cluster ComputingabstractCoscheduling remains an open question when parallel jobs are executed in a non-dedicated Cluster system jointly with the local workload. Our efforts are directed towards maximizing parallel efficiency and minimizing impact on the local workload performance in such a systems, by using predictive coscheduling techniques. A coscheduling model for non-dedicated Cluster computing is presented in this paper. Some new coscheduling performance metrics based on this model are defined. A new predictive coscheduling algorithm for this model is also proposed. The performance of our algorithm is analyzed and its performance compared with other coscheduling algorithms in the literature by simulation. Francesc Solsona Tehàs, Francesc Giné, Porfidio Hernández, Emilio Luque |
IPDPS | 4 |
| 2001 | Coscheduling under Memory Constraints in a NOW Environment
Francesc Giné, Francesc Solsona Tehàs, Porfidio Hernández, Emilio Luque |
JSSPP | 4 |
| 2000 | Integrating Automatic Techniques in a Performance Analysis Session (Research Note)
Antonio Espinosa 0001, Tomàs Margalef, Emilio Luque |
Euro-Par | 3 |
| 2000 | Exploiting Knowledge of Temporal Behaviour in Parallel Programs for Improving Distributed Mapping
Concepció Roig, Ana Ripoll, Miquel A. Senar, Fernando Guirado, Emilio Luque |
Euro-Par | 5 |
| 2000 | Implementing Explicit and Implicit Coscheduling in a PVM Environment (Research Note)
Francesc Solsona Tehàs, Francesc Giné, Porfidio Hernández, Emilio Luque |
Euro-Par | 4 |
| 2000 | Evaluation of an Adaptive Scheduling Strategy for Master-Worker Applications on Clusters of Workstations
Elisa Heymann, Miquel A. Senar, Emilio Luque, Miron Livny |
HiPC | 3 |
| 2000 | An Efficient Method for Improving Large Optimistic PDESabstractSTW (Switch Time Warp) is a mechanism for limiting the optimism of the TW (Time Warp) method. The proposed method achieves significant time/performance improvements for rollback reduction in optimistic parallel discrete event simulation (PDES). The STW uses a cost function to decide if a process is running in an overoptimistic state. The STW mechanism will decide the LP's execution order within the processor and the amount of CPU time to assign to a logical process to reduce the number of rollbacks. Remo Suppi, Fernando Cores, Emilio Luque |
MASCOTS | 3 |
| 1999 | A new method to make communication latency uniform: distributed routing balancingabstractArticle A new method to make communication latency uniform: distributed routing balancing Share on Authors: D. Franco Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile , I. Garcés Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile , E. Luque Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, Spain Unitat d'Arquitectura d'ordinadors i Sistemes Operatius, Departament d'Informàtica, Universitat Autònoma de Barcelona, 08193-Bellaterra, Barcelona, SpainView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 210–219https://doi.org/10.1145/305138.305195Online:01 May 1999Publication History 19citation351DownloadsMetricsTotal Citations19Total Downloads351Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Daniel Franco 0002, Indhira Garcés, Emilio Luque |
International Conference on Supercomputing | 3 |
| 1999 | Analytical Modeling of the Network Traffic PerformanceabstractInterconnection network modeling is an important field in order to study and understand interconnection network behaviour and its significance in telecommunication applications and distributed systems. In this paper, we show an analytical model that represents interconnection networks. The model accepts as inputs the network load consisting of the network topology, routing, and the communication pattern of the application. Any topology of any size and different parameters for the router are supported. The model outputs the latency behaviour of the interconnection network for each channel on each link. The accuracy of the model is shown by comparison with network simulation. The model is useful to study the latency/load curve of communication patterns, to calculate average network delays to identify hot-spots in the network and to perform network analysis and design. Indhira Garcés, Daniel Franco 0002, Emilio Luque |
MASCOTS | 3 |
| 1998 | Distributed routing balancing for interconnection network communicationabstractAn efficient design of the interconnection network is crucial because of its impact on the parallel computer performance. A high speed routing scheme that minimises contention and avoids the formation of hot-spots should be included in the design. We have developed a new method to uniformly balance communication traffic over the interconnection network called distributed routing balancing (DRB) that is based on limited and load-controlled path expansion in order to maintain a low message latency. The method uniformly distributes the communication load between all links of the interconnection network and maintains latency control provided that total bandwidth requirements do not exceed total available link bandwidth in the interconnection network. DRB defines how to create alternative paths to expand single paths (expanded path definition) and when to use them depending on traffic load (expanded path selection carried out by DRB routing). Some conclusions of the experimentation and comparisons with existing methods are given. It is demonstrated that DRB is a method to effectively balance network traffic. Indhira Garcés, Daniel Franco 0002, Emilio Luque |
HiPC | 3 |
| 1998 | On the Stability of a Distributed Dynamic Load Balancing AlgorithmabstractWe present a new fully distributed dynamic load balancing algorithm called DASUD (Diffusion Algorithm Searching Unbalanced Domains). Since DASUD is iterative and runs in an asynchronous way, a mathematical model that describes DASUD behaviour has been proposed and has been used to prove DASUD's convergence. DASUD has been evaluated by comparison with another well known strategy from the literature, namely, the SID (Sender Initiated Diffusion) algorithm. The comparison was carried out by considering a large set of load distributions which were applied to ring, torus and hypercube topologies, and the number of processors ranged from 8 to 128. From these experiments we have observed that DASUD outperforms the SID strategy as it provides the best trade-off between the global balance degree obtained at the final state and the number of iterations required to reach such a state. Ana Cortés, Ana Ripoll, Miquel A. Senar, F. Cedo, Emilio Luque |
ICPADS | 5 |
| 1996 | Teaching parallel processing: development of curriculum and software toolsabstractarticle Free Access Share on Teaching parallel processing: development of curriculum and software tools Authors: Jan Kwiatkowski Technical University of Wroclaw, Computer Science Department, Poland Technical University of Wroclaw, Computer Science Department, PolandView Profile , Marek Andruszkiewicz Technical University of Wroclaw, Computer Science Department, Poland Technical University of Wroclaw, Computer Science Department, PolandView Profile , Emilio Luque Universitat Autonoma de Barcelona, Computer Science Department, Spain Universitat Autonoma de Barcelona, Computer Science Department, SpainView Profile , Tomas Margalef Universitat Autonoma de Barcelona, Computer Science Department, Spain Universitat Autonoma de Barcelona, Computer Science Department, SpainView Profile , Jose Cunha Universidade Nova de Lisboa, Departamento de Informatica, Portugal Universidade Nova de Lisboa, Departamento de Informatica, PortugalView Profile , Joao Lourenco Universidade Nova de Lisboa, Departamento de Informatica, Portugal Universidade Nova de Lisboa, Departamento de Informatica, PortugalView Profile , Henryk Krawczyk Technical University of Gdansk Electronics Faculty, Poland Technical University of Gdansk Electronics Faculty, PolandView Profile , Stanislaw Szejko Technical University of Gdansk Electronics Faculty, Poland Technical University of Gdansk Electronics Faculty, PolandView Profile Authors Info & Claims ACM SIGCSE BulletinVolume 28Issue SI1996 pp 159–161https://doi.org/10.1145/237477.237633Published:01 January 1996Publication History 4citation199DownloadsMetricsTotal Citations4Total Downloads199Last 12 Months15Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Jan Kwiatkowski, Marek Andruszkiewicz, Emilio Luque, Tomàs Margalef, José C. Cunha, João Lourenço, Henryk Krawczyk, Stanislaw Szejko |
ITiCSE | 3 |
| 1996 | Parallel systems development in education: a guided methodabstractArticle Free Access Share on Parallel systems development in education: a guided method Authors: E. Luque Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , J. Sorribes Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , R. Suppi Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , E. Cesar Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , J. L. Falguera Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile , M. Serrano Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, Spain Department d'Informàtica, University Autonoma of Barcelona, 08193 Bellaterra, Barcelona, SpainView Profile Authors Info & Claims ITiCSE '96: Proceedings of the 1st conference on Integrating technology into computer science educationJune 1996 Pages 156–158https://doi.org/10.1145/237466.237629Online:01 January 1996Publication History 2citation169DownloadsMetricsTotal Citations2Total Downloads169Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Emilio Luque, Joan Sorribes, Remo Suppi, Eduardo César, J. Falguera, Massimo Serranó |
ITiCSE | 1 |
| 1994 | Scheduling of parallel programs including dynamic loops
Emilio Luque, Ana Ripoll, Tomàs Margalef, Ana Cortés |
Future Gener. Comput. Syst. | 1 |
| 1994 | Programming environment for a transputer based computer
Emilio Luque, Miquel A. Senar, Daniel Franco 0002, Porfidio Hernández, Elisa Heymann, Juan C. Moure |
Future Gener. Comput. Syst. | 1 |
| 1994 | Simulation of parallel systems: PSEE (Parallel System Evaluation Environment)
Emilio Luque, Remo Suppi, Joan Sorribes |
Future Gener. Comput. Syst. | 1 |
| 1992 | A quantitative approach for teaching parallel computingabstractParallel computing teaching has an important difficulty, there are few tools to directly learn the behavior of the parallel algorithms and the parallel architectures. Normally the student is formed to think in sequential algorithms running in sequential machines. We present PSEE, a tool to reduce the gap between the basic concepts and its utilization. PSEE is an integrated and interactive graphic environment which allows to simulate and evaluate the performance of parallel algorithms in parallel architectures. PSEE permits to manage the main characteristic parameters involved in the system in order to show the tuning grade of the algorithm/architecture couple. PSEE includes a graphic editor for algorithms and architectures in modelled form, an interactive simulator to run (simulate) the algorithm on the architecture and a performance evaluation instrument. Emilio Luque, Remo Suppi, Joan Sorribes |
SIGCSE | 1 |
| 1992 | Designing parallel systems: a performance prediction problem
Emilio Luque, Remo Suppi, Joan Sorribes |
Inf. Softw. Technol. | 1 |
| 1991 | Simulation and visualization tools for link-based parallel architectures
Emilio Luque, Remo Suppi, Joan Sorribes, M. A. Mayosky, Miquel A. Senar |
Microprocessing and Microprogramming | 1 |
| 1990 | Task Duplication Static-Scheduling for Multiprocessor Systems with Non-Fixed Execution Time Tasks
Emilio Luque, Ana Ripoll, Porfidio Hernández, Tomàs Margalef |
ICPP (1) | 1 |
| 1990 | Impact of task duplication on static-scheduling performance in multiprocessor systems with variable execution-time tasks
Emilio Luque, Ana Ripoll, Porfidio Hernández, Tomàs Margalef |
ICS | 1 |
| 1987 | Coprocessor for real-time dynamic vertical migration
Emilio Luque, Joan Sorribes, Ana Ripoll |
Microprocessing and Microprogramming | 1 |
| 1985 | Self-tuning machines
Mario De Blasi, Anna Gentile, Emilio Luque, Ana Ripoll |
Microprocessing and Microprogramming | 3 |
| 1984 | Integer Linear Programming for Microprograms Register Allocation
Emilio Luque, Ana Ripoll |
Inf. Process. Lett. | 1 |
| 1984 | A development system for self tuning machines
Mario De Blasi, Anna Gentile, Emilio Luque, Joan Sorribes |
Microprocessing and Microprogramming | 3 |
| 1981 | Microprogramming: A tool for vertical migration
Emilio Luque, Ana Ripoll |
Microprocessing and Microprogramming | 1 |
| 1980 | Tuning Architecture via Microprogramming
Emilio Luque, Ana Ripoll |
Inf. Process. Lett. | 1 |
| 1980 | Dynamic microprogramming in computer architecture redefinition
Emilio Luque, Ana Ripoll, José J. Ruz |
Euromicro Newsletter | 1 |
| 1979 | Database Concurrent Processor
Emilio Luque, José J. Ruz, Ana Ripoll, Alfredo Bautista |
VLDB | 1 |
| 1978 | A general purpose computer emulator
Emilio Luque, Lorenzo Moreno Ruiz, Francisco Tirado |
Euromicro Newsletter | 1 |