EDBT 2026 Demo / reviewers in the wild / expert
David E. Singh
dblp:73/4490 · also David Exposito Singh
· DBLP profile ↗
39ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0002-8125-0049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | I/O Behind the Scenes: Bandwidth Requirements of HPC Applications with Asynchronous I/OabstractI/O bandwidth is a critical resource in an HPC cluster. As with all shared resources, its availability is impacted significantly by the users and the applications they execute. Without proper restrictions, jobs consuming more prominent portions of the I/O bandwidth can severely affect other jobs by notably prolonging their runtime. In such a context, applications that perform asynchronous I/O bring unique properties that allow for the reduction of such effects. That is, by limiting the bandwidth to the required one to perform the I/O in the background of the compute phases, I/O bursts can be flattened without significantly prolonging the application time, if at all. Hence, the bandwidth consumption of such applications is limited to what they need, sparing much of the system bandwidth to other applications. At the same time, these applications achieve higher parallel efficiency due to the overlapping of different resources (e.g., compute and I/O). This paper shows these aspects and demonstrates our approach to finding the required bandwidth for applications that use asynchronous I/O. Moreover, we apply it automatically using an MPI implementation of a bandwidth limitation approach at the application level. We validate our approach with several experiments on a large production cluster and show the impact of our approach on the application behavior and its importance for the system throughput. Ahmad Tarraf, Javier Fernández 0001, David E. Singh, Taylan Özden, Jesús Carretero 0001, Felix Wolf 0001 |
CLUSTER | 3 |
| 2024 | Performance-driven scheduling for malleable workloadsabstractAbstract The development of adaptive scheduling algorithms that take advantage of malleability has become a crucial area of research in many large-scale projects. Malleable workloads can improve the system’s performance but, at the same time, provide an extra dimension to the scheduling problem. This paper proposes an adaptive, performance-based job scheduling method that emphasizes the backfilling concept with malleability. The proposed method performs the malleability operations only when the estimated execution time of the involved applications is better than or equal to the execution time on the allocated resources without reconfiguration. The reconfiguration feasibility is determined by performance models considering the application scalability and reconfiguration overheads. Different policies for implementing malleability are presented, each targeting a specific workload in terms of job size and scalability. The comprehensive evaluation shows an improvement in the slowdown up to 49% compared to the non-adaptive baseline scheduling algorithm. Njoud O. Almaaitah, David E. Singh, Taylan Özden, Jesús Carretero 0001 |
J. Supercomput. | 2 |
| 2024 | Detailed parallel social modeling for the analysis of COVID-19 spreadabstractAbstract Agent-based epidemiological simulators have been proven to be one of the most successful tools for the analysis of COVID-19 propagation. The ability of these tools to reproduce the behavior and interactions of each single individual leads to accurate and detailed results, which can be used to model fine-grained health-related policies like selective vaccination campaigns or immunity waning. One characteristic of these tools is the large amount of input data and computational resources that they require. This relies on the development of parallel algorithms and methodologies for generating, accessing, and processing large volumes of data from multiple data sources. This work presents a parallel workflow for extending the social modeling of EpiGraph, an agent-based simulator. We have included two novel parallel social generation stages that generate a detailed and realistic social model and one new visualization stage. This work also presents a description of the algorithms used in each stage, different optimization techniques that permit to reduce the application convergence time, and a practical evaluation of large workloads on HPC systems. Results show that this contribution can be efficiently executed in parallel architectures and the results allow to increase the simulation detail level, representing a significant advance in the simulator scenario modeling. As a summary of results, the first contribution of this paper is the development of two models (a spatial and a social one) that assign geographical and socioeconomic indicators to each simulated individual (i.e., agents), reproducing the real social distribution of the city of Madrid. The second contribution presents an improved parallel and distributed algorithm that executes the two aforementioned models using different parallelization strategies and preserving the load balance. Aymar Cublier Martínez, Jesús Carretero 0001, David E. Singh |
J. Supercomput. | 3 |
| 2024 | Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future OpportunitiesabstractWith the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research. Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2023 | COVID-19 incidence estimates and forecast by metaprediction for the Comunidad de Madrid *abstractEpidemiological mathematical models have been proved crucial in supporting the decision-making of the health authorities during the COVID-19 pandemic. In this context, this work presents two contributions. The first one is a methodology to integrate different data sources into a single time series that provides realistic COVID-19 incidence rates considering both the reported and unreported cases in Spain and Comunidad de Madrid. The second contribution is a novel ensemble forecast model that uses as input the predictions of three different COVID-19 forecasts models. These approaches have been used to provide forecast predictions in the scope of PredCov project, supporting both the Spanish and the European Union -via the European Centre for Disease Prevention and Control-health authorities. The output generated by the ensemble model provides a combined -and more accurate-prediction of the COVID-19 incidence. This work includes a description of both contributions and discusses the results provided by them. Aymar Cublier Martínez, Mario Muñoz Organero, David Moriña, Diana Gomez-Barroso, David E. Singh |
BIBM | 5 |
| 2023 | Fine-grained parallel social modelling for analyzing COVID-19 propagationabstractAgent-based epidemiological simulators have been proven to be one of the most successful tools for the analysis of the COVID-19 propagation. The ability of these tools to reproduce the behavior and interactions of each single individual leads to accurate and detailed results, that can be used to model fine-grained health-related policies like selective vaccination campaigns or immunity waning. One characteristic of these tools is the large amount of input data and computational resources and that they require. This relies on the development of parallel algorithms and methodologies for generating, accessing and processing large volumes of data from multiple data sources. This work presents a parallel workflow for extending the social modelling of EpiGraph, an agent-based simulator. We have included two novel parallel social generation stages -that provide detailed and realistic social model- and one new visualization stage. The work presents a description of the algorithms used in each stage and a practical evaluation on a real platform. Results show that this contribution can be efficiently executed in parallel architectures and increases the simulation detail level, representing a significant advance in the simulator scenario modelling. Aymar Cublier Martínez, Álejandro Alvarez Isabel, Jesús Carretero 0001, David E. Singh |
PDP | 4 |
| 2023 | Evaluating the spread of Omicron COVID-19 variant in SpainabstractThis work analyzes the propagation the highly transmissible COVID-19 variant Omicron across Spain via simulation by using EpiGraph. EpiGraph is an agent-based parallel simulator that reproduces the COVID-19 propagation over wide areas. In this work we consider a population of 19,574,086 individuals of the 63 most populated cities of Spain, for the time interval between May 15th 2021 and March 6th 2022. The main variants existing at the start of the simulation were the Alpha and Delta, with prevalence of 4% and 96%. Then, during the second half of November 2021, the Omicron variant appears in Spain. Due to the higher transmission of this new variant - about 2 times larger than Delta, it quickly spreads through all the cities and becomes the dominant strain in the country. In this work we analyze the propagation of this variant under different mobility restrictions and patient zero scenarios. We first define a baseline scenario which reproduces the existing conditions of the COVID-19 propagation in Spain for our period of study. We then consider alternative scenarios for different starting locations of the propagation. Finally, for each one of these scenarios, we evaluate different transportation intensities - i.e. movement of individuals between the cities. The main conclusion is that, independently of the initial location of the Omicron variant and the existing transportation conditions, the Omicron variant spreads through all the country in a short time interval. The work presented in this paper also implements and evaluates a power monitoring and optimization system aimed at reducing the energy consumption of such massive simulations as the ones performed in EpiGraph. Miguel Guzmán-Merino, Maria-Cristina V. Marinescu, Alberto Cascajo, Jesús Carretero 0001, David E. Singh |
Future Gener. Comput. Syst. | 5 |
| 2022 | Evaluating the spread of Omicron COVID-19 variant in SpainabstractThis work analyzes the propagation of Omicron, a high transmissible COVID-19 variant, across Spain by means of EpiGraph. EpiGraph is an agent-based parallel simulator that reproduces the COVID-19 propagation over wide areas. In this work we consider a population of 19,574,086 individuals related to the 63 most populated cities of Spain for the time interval comprised between May 15th of 2021 and March 6th of 2022. The main variants existing at the start of the simulation were the Alpha and Delta, with a a 4% and 96% prevalence of the existing infections, respectively. Then, during the second mid of November of 2021 the Delta variant appears in Spain. Given to the higher transmission of this new variant -about 2 times larger than Delta-, it quickly spreads through all the cities and becomes the dominant strain of the country. In this work we analyze the propagation of this variant under multiple conditions. First, we define a baseline scenario, that reproduces the existing conditions of the COVID-19 propagation in Spain for this period. Then, we consider alternative scenarios, in which different locations of the initial spread of Omicron variant are considered. Finally, for each one of these scenarios, we evaluate different transportation intensities -i.e. movement of individuals between the cities-. The main conclusion of this work is that, independently of the initial location of the Omicron variant, and the existing transportation conditions, the Omicron variant spreads though all the country in a short time interval. Miguel Guzmán-Merino, Maria-Cristina V. Marinescu, David E. Singh |
CCGRID | 3 |
| 2022 | Improving Congestion Control through Fine-Grain Monitoring of InfiniBand NetworksabstractCongestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms. Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001 |
HOTI | 5 |
| 2021 | LIMITLESS - LIght-weight MonItoring Tool for LargE Scale SystemsabstractThis work presents LIMITLESS, a HPC framework that provides new strategies for monitoring clusters. LIMITLESS is a scalable light-weight monitor that is integrated with other HPC runtimes in order to obtain an holistic view of the system that combines both platform and application monitoring. This paper presents a description of the novel components of the architecture, including new approaches for reaching a higher scalability based on a combination of in-transit processing and performance prediction. This work also includes a practical evaluation on simulated and real platforms, that shows significant monitoring scalability, retrieving data capacity and reduced overheads. Alberto Cascajo, David E. Singh, Jesús Carretero 0001 |
PDP | 2 |
| 2020 | Mapping and scheduling HPC applications for optimizing I/OabstractIn HPC platforms, concurrent applications are sharing the same file system. This can lead to conflicts, especially as applications are more and more data intensive. I/O contention can represent a performance bottleneck. The access to bandwidth can be split in two complementary yet distinct problems. The mapping problem and the scheduling problem. The mapping problem consists in selecting the set of applications that are in competition for the I/O resource. The scheduling problem consists then, given I/O requests on the same resource, in determining the order to these accesses to minimize the I/O time. In this work we propose to couple a novel bandwidth-aware mapping algorithm to I/O list-scheduling policies to develop a cross-layer optimization solution. Jesús Carretero 0001, Emmanuel Jeannot, Guillaume Pallez, David E. Singh, Nicolas Vidal 0001 |
ICS | 4 |
| 2019 | Combining malleability and I/O control mechanisms to enhance the execution of multiple applications
David E. Singh, Jesús Carretero 0001 |
J. Syst. Softw. | 1 |
| 2016 | QuizMonitor: A learning platform that leverages student monitoringabstractThis work presents the design, implementation, and evaluation of a learning platform that addresses two main objectives: first it provides and on-line quiz tool for students which can be used as a complementary learning approach to the classroom courses. Secondly, this tool performs a detailed analysis of learners use, considering not only the number of mistakes students have made but also the student temporal use distribution and opinion (obtained by a survey) about the difficulty of the learning contents. All this information is processed and used to provide feedback to the teachers identifying the most difficult contents of the area of study and the students with a low learning performance. We have developed and evaluated this tool in two university degree subjects. The obtained results are promising, showing that students can improve their final grades through its use and that teachers can identify the student learning problems, in order to assess where students may require more support or challenge. In addition, we analyze the impact of different metrics obtained by this tool (academic performance, intensity of study and student proactivity) in the final subject grades. Carlos Gomez, David E. Singh, Jesús Carretero 0001 |
EDUCON | 2 |
| 2016 | Improving the Energy Efficiency of MPI Applications by Means of MalleabilityabstractThis work presents two novel techniques for increasing the energy efficiency of parallel applications by means of malleability. These techniques are implemented as an extension of Flex-MPI, a library implemented on top of MPI, which provides performance-aware dynamic reconfiguration for MPI-based applications. During the application execution, Flex-MPI performs energy and performance monitoring by means of energy and performance counters. It leverages this information in order to adapt the program performance using two energy policies: energy minimization and performance-per-watt maximization. The evaluation results show that these new energy-aware capabilities permit MPI applications to be executed in an optimized way both in terms of performance and energy efficiency. Manuel Rodriguez-Gonzalo, David E. Singh, Francisco Javier García Blas, Jesús Carretero 0001 |
PDP | 2 |
| 2015 | Towards efficient large scale epidemiological simulations in EpiGraphabstractThe work we present in this paper focuses on understanding the propagation of flu-like infectious outbreaks between geographically distant regions due to the movement of people outside their base location. Our approach incorporates geographic location and a transportation model into our existing region-based, closed-world EpiGraph simulator to model a more realistic movement of the virus between different geographic areas. This paper describes the MPI-based implementation of this simulator, including several optimization techniques such as a novel approach for mapping processes onto available processing elements based on the temporal distribution of process loads. We present an extensive evaluation of EpiGraph in terms of its ability to simulate large-scale scenarios, as well as from a performance perspective. Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
Parallel Comput. | 2 |
| 2015 | Enhancing the performance of malleable MPI applications by using performance-aware dynamic reconfiguration
Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
Parallel Comput. | 2 |
| 2013 | FLEX-MPI: An MPI Extension for Supporting Dynamic Load Balancing on Heterogeneous Non-dedicated Systems
Gonzalo Martín 0001, Maria-Cristina V. Marinescu, David E. Singh, Jesús Carretero 0001 |
Euro-Par | 3 |
| 2013 | Parallel algorithm for simulating the spatial transmission of influenza in EpiGraphabstractThis paper introduces an approach to modeling and simulating the propagation of flu-like infectious diseases over large, widely spread urban areas connected by transportation networks. We incorporate geographic location and a transportation model into our region-based, closed-world EpiGraph simulator to realistically model the movement of the virus between different geographic regions. The resulting simulator can assist in understanding how outbreaks propagate between far apart regions due to the movement of people outside their base location. This paper describes the MPI-based implementation of EpiGraph and its performance evaluation when simulating large-scale scenarios. We evaluate the simulator both on a distributed memory system and on a shared memory system. Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
EuroMPI | 2 |
| 2012 | Runtime Support for Adaptive Resource Provisioning in MPI Applications
Gonzalo Martín 0001, David E. Singh, Maria-Cristina V. Marinescu, Jesús Carretero 0001 |
EuroMPI | 2 |
| 2012 | Dynamic-CoMPI: dynamic optimization techniques for MPI parallel applications
Rosa Filgueira, Jesús Carretero 0001, David E. Singh, Alejandro Calderón 0001, Alberto Nuñez |
J. Supercomput. | 3 |
| 2010 | Lessons Learnt Porting Parallelisation Techniques for Irregular Codes to NUMA SystemsabstractThis work presents a study undertaken to characterise the behaviour of some parallelisation techniques for irregular codes, previously developed for SMP architectures, on a several-node SMP NUMA system. The main objective is to determine the performance effect of bus contention and cache coherency in such a complex architecture. Results show that: (1) cores which share a socket can be considered as independent processors in this context; (2) for big data sizes, the effect of sharing a bus degrades the performance but masks the cache coherency effects and (3) the NUMA-ratio is a critical factor on irregular codes. These results allow us to study the effect in performance of the thread-to-core mappings and memory allocation policies. Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, David LaFrance-Linden, Francisco F. Rivera, David E. Singh |
PDP | 5 |
| 2010 | SENFIS: a Sensor Node File System for increasing the scalability and reliability of Wireless Sensor Networks applications
Soledad Escolar, Florin Isaila, Alejandro Calderón 0001, Luis Miguel Sánchez, David E. Singh |
J. Supercomput. | 5 |
| 2009 | A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001 |
J. Supercomput. | 1 |
| 2009 | A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001 |
J. Supercomput. | 1 |
| 2008 | View-Based Collective I/O for MPI-IOabstractThis paper presents the design and implementation of a new file system independent collective I/O optimization based on file views: view-based collective I/O. View-based collective I/O has been implemented and evaluated inside ROMIO implementation of MPI-IO standard. The evaluation section shows that view-based I/O outperforms the original two-phase collective I/O from ROMIO in most of the cases for three well-known parallel I/O benchmarks. This is especially due to a smaller cost of scatter/gather operations, a reduction of the metadata overhead, and a smaller number of collective communication and synchronization primitives used in the implementation. Francisco Javier García Blas, Florin Isaila, David E. Singh, Jesús Carretero 0001 |
CCGRID | 3 |
| 2008 | Exploiting data compression in collective I/O techniquesabstractThis paper presents Two-Phase Compressed I/O (TPC I/O,) an optimization of the Two-Phase collective I/O technique from ROMIO, the most popular MPI-IO implementation. In order to reduce network traffic, TPC I/O employs LZO algorithm to compress and decompress exchanged data in the inter-node communication operations. The compression algorithm has been fully implemented in the MPI collective technique, allowing to dynamically use (or not) compression. Compared with Two-Phase I/O, Two-Phase Compressed I/O obtains important improvements in the overall execution time for many of the considered scenarios. Rosa Filgueira, David E. Singh, Juan Carlos Pichel, Jesús Carretero 0001 |
CLUSTER | 2 |
| 2008 | Reordering Algorithms for Increasing Locality on Multicore ProcessorsabstractIn order to efficiently exploit available parallelism, multicore processors must address contention for shared resources as cache hierarchy. This fact becomes even more important when irregular codes are executed on them, which is the case for sparse matrix ones. In this paper a technique for increasing locality of sparse matrix codes on multicore platforms is presented. The technique consists on reorganizing the data guided by a locality model which introduces the concept of windows of locality. The evaluation of the reordering technique has been performed on two different leading multicore platforms: Intel Core2Duo and Intel Xeon. Experimental results show important performance improvements when using our reordered matrices with respect to original ones. In particular, an average execution time reduction of about 30% is achieved considering different number of running threads. These results are due to an improved overall cache behavior. Likewise, a comparison of our proposal with some standard reordering techniques is included in the paper. Results point out that the reordering technique always outperforms standard algorithms and is effective for matrices with any structure. Juan Carlos Pichel, David E. Singh, Jesús Carretero 0001 |
HPCC | 2 |
| 2007 | Optimization and evaluation of parallel I/O in BIPS3D parallel irregular applicationabstractThis paper presents the optimization and evaluation of parallel I/O for the BIPS3D parallel irregular application, a 3-dimensional simulation of BJT and HBT bipolar devices. The parallel version of BIPS3D employs Metis, a library for partitioning graphs, finite element meshes, or sparse matrices. First, we show how the partitioning information provided by Metis can be used in order to improve the performance of parallel I/O. Second, we propose a novel technique, called Interval Data Grouping (IDG), which exploits the data replication of mesh nodes for optimizing the scheduling of the parallel file operations. Finally, we evaluate the parallel I/O version of BIPS3D for various existing parallel I/O techniques and present an in-depth analysis of the IDG performance. Rosa Filgueira, David E. Singh, Florin Isaila, Jesús Carretero 0001, Antonio J. García-Loureiro |
IPDPS | 2 |
| 2007 | An Inspector/Executor Based Strategy to Efficiently Parallelize N-Body Simulation Programs on Shared Memory SystemsabstractReordering of data is becoming more and more significant in order to achieve a higher performance in memory data access and, particularly, in program runtime. This fact becomes specially important in parallel applications that are executed in shared memory systems. This work presents a new parallelizing, run time strategy for irregular structures associated to N-Body problem simulation algorithms. Such strategy, so-called STPCLS (Step Classification), is based on the inspector-executor paradigm. It has been tested in a shared memory system using a significant set of irregular loops. The outcomes show that the efficiency of our solution is high, and the benefits overcome the overheads imposed by our algorithm. Juan Ángel Lorenzo del Castillo, Julio L. Albín, Tomás F. Pena, Francisco F. Rivera, David E. Singh |
ISPDC | 5 |
| 2007 | Multiple-Phase Collective I/O Technique for Improving Data Access LocalityabstractThis paper presents multiple-phase collective I/O, a novel collective I/O technique for distributed memory multiprocessors. Multiple-phase collective I/O is a refinement of two-phase collective I/O technique. The communication phase is structured into several steps, which progressively increase the locality of the data to be written to a file system. Besides the description of multiple-phase collective I/O, our paper addresses two additional objectives. First, the authors target to improve the efficiency of the sulphur transport Eurelian model 2 (STEM-II) application. STEM-II is an air quality model that simulates transport, chemical transformations, emission and deposition processes in a unified framework. Due to the large amount of processed data, I/O becomes a critical factor for the application performance. Multiple-phase collective I/O, considerably enhances the performance of the I/O stage in particular and, consequently, of the whole application in general. Second objective consists of evaluating and comparing the performance of multiple-phase collective I/O with that of other well known parallel I/O techniques David E. Singh, Florin Isaila, Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001 |
PDP | 1 |
| 2006 | On the Reliability of Web Clusters with Partial Replication of ContentsabstractTraditionally, distributed Web servers have used two strategies for allocating files on server nodes: full replication and full distribution. While full replication provides a highly reliable solution, it limits storage capacity to the capacity of the smallest node. On the other hand, full distribution provides higher storage capacity at the cost of lower reliability. A hybrid solution is partial replication where every file is allocated to a small number of nodes. The most promising architecture for a partial replication strategy is the Web cluster architecture. However, Web clusters present a big flaw from reliability perspective as they contain a single point of failure. To correct this flaw, in this paper we present a modified architecture: the Web cluster with distributed Web switch. Reliability of Web clusters is evaluated for different replication strategies. System evaluations show that our proposal leads to a highly reliable solution with high scalability. José Daniel García, Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, David E. Singh, Alejandro Calderón 0001 |
ARES | 5 |
| 2006 | Integrating Logical and Physical File Models in the MPI-IO Implementation for "Clusterfile"abstractThis paper presents the design and implementation of the MPI-IO interface for the Clusterfile parallel file system. The approach offers the opportunity of achieving a high correlation between the file access patterns of parallel applications and the physical file distribution. First, any physical file distribution can be expressed by means of MPI data types. Second, mechanisms such as views and collective I/O operations are portably implemented inside the file system, unifying the I/O scheduling strategies of the MPI-IO library and the file system. The experimental section demonstrates performance benefits of more than one order of magnitude. Florin Isaila, David E. Singh, Jesús Carretero 0001, Félix García Carballeira, Gabor Szeder, Thomas Moschny |
CCGRID | 2 |
| 2006 | Improving the Performance of Cluster Applications through I/O Proxy ArchitectureabstractClusters are the most common solutions for high performance computing at the present time. In this kind of systems, an important challenge is the I/O subsystem design. Typically, these environments are not flexible enough and the only way to solve performance bottlenecks is adding new hardware. In this paper, we show how an I/O proxy-based architecture can improve the I/O performance of cluster applications in three ways: adapting to the application requirements, reducing the load on the I/O nodes, and finally, increasing the global performance of the storage system Luis Miguel Sánchez, Florin Isaila, Alejandro Calderón 0001, David E. Singh, José Daniel García |
CLUSTER | 4 |
| 2006 | A Quantitative Justification to Partial Replication of Web Contents
José Daniel García, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Alejandro Calderón 0001, David E. Singh |
ICCSA (4) | 6 |
| 2006 | Image segmentation based on merging of sub-optimal segmentations
Juan Carlos Pichel, David E. Singh, Francisco F. Rivera |
Pattern Recognit. Lett. | 2 |
| 2003 | Increasing the Parallelism of Irregular Loops with Dependences
David E. Singh, María J. Martín, Francisco F. Rivera |
Euro-Par | 1 |
| 2003 | High performance air pollution modeling for a power plant environment
María J. Martín, David E. Singh, José Carlos Mouriño, Francisco F. Rivera, Ramón Doallo, Javier D. Bruguera |
Parallel Comput. | 2 |
| 2002 | Improving Locality in the Parallelization of Doacross Loops (Research Note)
María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
Euro-Par | 2 |
| 2002 | Exploiting Locality in the Run-Time Parallelization of Irregular LoopsabstractThe goal of this work is the efficient parallel execution of loops with indirect array accesses, in order to be embedded in a parallelizing compiler framework. In this kind of loop pattern, dependences can not always be determined at compile-time as, in many cases, they involve input data that are only known at run-time and/or the access pattern is too complex to be analyzed In this paper we propose runtime strategies for the parallelization of these loops. Our approaches focus not only on extracting parallelism among iterations of the loop, but also on exploiting data access locality to improve memory hierarchy behavior and, thus, the overall program speedup. Two strategies are proposed one based on graph partitioning techniques and other based on a block-cyclic distribution. Experimental results show that both strategies are complementary and the choice of the best alternative depends on some features of the loop pattern. María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
ICPP | 2 |