EDBT 2026 Demo / reviewers in the wild / expert
Javier Fernández 0001
dblp:91/4494-1 · also Javier Fernández Muñoz
· DBLP profile ↗
33ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0001-8539-5491ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | I/O Behind the Scenes: Bandwidth Requirements of HPC Applications with Asynchronous I/OabstractI/O bandwidth is a critical resource in an HPC cluster. As with all shared resources, its availability is impacted significantly by the users and the applications they execute. Without proper restrictions, jobs consuming more prominent portions of the I/O bandwidth can severely affect other jobs by notably prolonging their runtime. In such a context, applications that perform asynchronous I/O bring unique properties that allow for the reduction of such effects. That is, by limiting the bandwidth to the required one to perform the I/O in the background of the compute phases, I/O bursts can be flattened without significantly prolonging the application time, if at all. Hence, the bandwidth consumption of such applications is limited to what they need, sparing much of the system bandwidth to other applications. At the same time, these applications achieve higher parallel efficiency due to the overlapping of different resources (e.g., compute and I/O). This paper shows these aspects and demonstrates our approach to finding the required bandwidth for applications that use asynchronous I/O. Moreover, we apply it automatically using an MPI implementation of a bandwidth limitation approach at the application level. We validate our approach with several experiments on a large production cluster and show the impact of our approach on the application behavior and its importance for the system throughput. Ahmad Tarraf, Javier Fernández 0001, David E. Singh, Taylan Özden, Jesús Carretero 0001, Felix Wolf 0001 |
CLUSTER | 2 |
| 2024 | Performance and programmability of GrPPI for parallel stream processing on multi-coresabstractAbstract GrPPI library aims to simplify the burdening task of parallel programming. It provides a unified, abstract, and generic layer while promising minimal overhead on performance. Although it supports stream parallelism, GrPPI lacks an evaluation regarding representative performance metrics for this domain, such as throughput and latency. This work evaluates GrPPI focused on parallel stream processing. We compare the throughput and latency performance, memory usage, and programmability of GrPPI against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks and benchmarks with handwritten parallel code using the same backends supported by GrPPI. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is often competitive with handwritten parallel code, the infeasibility of fine-tuning GrPPI is a crucial drawback for emerging applications. Despite this, programmability experiments estimate that GrPPI can potentially reduce the development time of parallel applications by about three times. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
J. Supercomput. | 5 |
| 2023 | Dynamic management of processes and communicators in malleable MPI applicationsabstractA malleable application is defined as one that can increase or decrease its resources dynamically based on workload variations. These applications often leverage a job manager to handle the resources. The MPI standard incorporates strategies to increase the number of processes connected to a running application. However, it does not clarify how to use these same features to reduce said processes. This paper presents a strategy compatible with the latest version of the MPI standard to develop malleable MPI applications. It allows the addition and removal of resources at the computing node level in a discretionary manner. The results obtained from the evaluation show that the proposed strategy has acceptable performance and scalability for the functionality it provides. Javier Fernández 0001, Alberto Cascajo, Jesús Carretero 0001 |
ICPADS | 1 |
| 2023 | A Latency, Throughput, and Programmability Perspective of GrPPI for Streaming on Multi-coresabstractSeveral solutions aim to simplify the burdening task of parallel programming. The GrPPI library is one of them. It allows users to implement parallel code for multiple backends through a unified, abstract, and generic layer while promising minimal overhead on performance. An outspread evaluation of GrPPI regarding stream parallelism with representative metrics for this domain, such as throughput and latency, was not yet done. In this work, we evaluate GrPPI focused on stream processing. We evaluate performance, memory usage, and programming effort and compare them against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is competitive with handwritten code in some cases, in other cases, the infeasibility of fine-tuning GrPPI is a crucial drawback. Despite this, programmability experiments estimate that GrPPI has the potential to reduce by about three times the development time of parallel applications. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, André Sacilotto Santos, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
PDP | 6 |
| 2022 | Convergence of HPC and Big Data in extreme-scale data analysis through the DCEx programming modelabstractHigh-level programming models can help application developers to access and use resources without the need to manage low-level architectural entities, as a parallel programming model defines a set of programming abstractions that simplify the way by which a programmer structures and expresses her/his algorithm. Early proposals of Exascale programming tools are based on the adaptation of traditional parallel programming languages and hybrid solutions. This incremental approach is too conservative, often resulting in very complex code. This paper describes the design features, the programming constructs, and the runtime mechanisms of the Data Centric programming model for Exascale systems (DCEx). DCEx is based on structuring applications into data-parallel blocks. Blocks are units of shared-and distributed-memory parallel computation, communication, and migration in the memory/storage hierarchy. Blocks and their message queues are mapped onto processes and placed in memory/storage by the DCEx runtime. Those data-parallel blocks are orchestrated by using distributed parallel patterns that simplify the development cost. DCEx aims to reach the convergence of traditional HPC programming models, mainly based on MPI, with the emerging technologies based on the data intensive paradigms. To demonstrate the potential of DCEx, we carried out an experimental evaluation developing a real-world diffusion-weighted magnetic resonance imaging data processing application in a neuroimaging research context. Francisco Javier García Blas, Javier Fernández 0001, Jesús Carretero 0001, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Alberto Fernández-Pena, Daniel Martín de Blas |
SBAC-PAD | 2 |
| 2020 | Relaxing the one definition rule in interpreted C++abstractMost implementations of the C++ programming language generate binary executable code. However, interpreted execution of C++ sources has its own use cases as the Cling interpreter from CERN's ROOT project has shown. Some limitations are derived from the ODR (One Definition Rule) that rules out multiple definitions of entities within a single translation unit (TU). ODR is there to ensure uniform view of a given C++ entity across translation units. Ensuring uniform view of C++ entities helps when producing ABI compatible binaries. Interpreting C++ presumes a single ever-growing translation unit that define away some of the ODR use-cases. Therefore, it may well be desirable to relax the ODR and, consequently, to support the ability of developers to override any existing definition for a given declaration. This approach is especially well-suited for iterative prototyping. In this paper, we extend Cling, a Clang/LLVM-based C++ interpreter, to enable redefinitions of C++ entities at the prompt. To achieve this, top-level declarations are nested into inline namespaces and the translation unit lookup table is adjusted to invalidate previous definitions that would otherwise result in ambiguities. Formally, this technique refactors the code to an equivalent that does not violate the ODR, as each definition is nested in a different namespace. Furthermore, any previous definition that has been shadowed is still accessible by means of its fully-qualified name. A prototype implementation of the presented technique has been integrated into the Cling C++ interpreter, showing that our technique is feasible and usable. Javier López-Gómez, Javier Fernández 0001, David del Rio Astorga, Vassil Vassilev, Axel Naumann, José Daniel García |
CC | 2 |
| 2020 | Detecting semantic violations of lock-free data structures through C++ contracts
Javier López-Gómez, David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
J. Supercomput. | 4 |
| 2019 | Exploring stream parallel patterns in distributed MPI environments
Javier López-Gómez, Javier Fernández 0001, David del Rio Astorga, Manuel F. Dolz, José Daniel García |
Parallel Comput. | 2 |
| 2019 | Hybrid static-dynamic selection of implementation alternatives in heterogeneous environments
David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, Francisco Javier García Blas |
J. Supercomput. | 3 |
| 2018 | Parallelizing and Optimizing LHCb-Kalman for Intel Xeon Phi KNL ProcessorsabstractReal time data processing is an important component of particle physics experiments with large computing resource requirements. As the Large Hadron Collider (LHC) at CERN is preparing for its next upgrade the LHCb experiment is upgrading its detector for a 30x increase in data throughput. In preparation for this upgrade the experiment is considering a number of architectural improvements encompassing both its software and hardware infrastructure. One of the hardware platforms under consideration is the Intel Xeon-Phi Knights Landing processor. Thanks to its on-package high-bandwidth memory and many-core architecture it offers an interesting alternative to more traditional server systems. We present a scalable, multi-threaded and NUMA-aware Kalman filter proto-application for particle track fitting expressed in terms of generic parallel patterns using the GrPPI interface. We show how code maintainability and readability improves, while maintaining comparable levels of performance to the baseline implementation. This is achieved by keeping the parallel algorithms in the underlying framework generic, but topology aware through the use of the Portable Hardware Locality (hwloc) library, which allows us to target different architectures with the same program. We measure the performance of our topology-aware GrPPI Kalman filter implementation on the Intel Xeon-Phi Knights Landing platform and conclude on the feasibility of integrating such high-level parallelization libraries in complex software frameworks such as LHCb's Gaudi framework. Placido Fernández, David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, Omar Awile, José Daniel García |
PDP | 4 |
| 2018 | Supporting MPI-distributed stream parallel patterns in GrPPIabstractIn the recent years, the large volumes of stream data and the near real-time requirements of data streaming applications have exacerbated the need for new scalable algorithms and programming interfaces for distributed and shared-memory platforms. To contribute in this direction, this paper presents a new distributed MPI back end for GrPPI, a C++ high-level generic interface of data-intensive and stream processing parallel patterns. This back end, as a new execution policy, supports the distributed and hybrid (distributed and shared-memory) parallel execution of the Pipeline and Farm patterns, where the hybrid mode combines the MPI policy with a GrPPI shared-memory one. A detailed analysis of the GrPPI MPI execution policy reports considerable benefits from the programmability, flexibility and readability points of view. The experimental evaluation on a streaming application with different distributed and shared-memory scenarios reports considerable performance gains with respect to the sequential versions at the expense of negligible GrPPI overheads. Javier Fernández 0001, Manuel F. Dolz, David del Rio Astorga, Javier Prieto Cepeda, José Daniel García |
EuroMPI | 1 |
| 2018 | Paving the way towards high-level parallel pattern interfaces for data stream processing
David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
Future Gener. Comput. Syst. | 3 |
| 2018 | Assessing and discovering parallelism in C++ code for heterogeneous platforms
David del Rio Astorga, Rafael Sotomayor, Luis Miguel Sánchez, Francisco Javier García Blas, Alejandro Calderón 0001, Javier Fernández 0001 |
J. Supercomput. | 6 |
| 2017 | Probabilistic-Based Selection of Alternate Implementations for Heterogeneous Platforms
Javier Fernández 0001, Andrés Sánchez Cuadrado, David del Rio Astorga, Manuel F. Dolz, José Daniel García |
ICA3PP | 1 |
| 2017 | A generic parallel pattern interface for stream and data processingabstractSummary Current parallel programming frameworks aid developers to a great extent in implementing applications that exploit parallel hardware resources. Nevertheless, developers require additional expertise to properly use and tune them to operate efficiently on specific parallel platforms. On the other hand, porting applications between different parallel programming models and platforms is not straightforward and demands considerable efforts and specific knowledge. Apart from that, the lack of high‐level parallel pattern abstractions, in those frameworks, further increases the complexity in developing parallel applications. To pave the way in this direction, this paper proposesGRPPI, a generic and reusable parallel pattern interface for both stream processing and data‐intensive C++ applications.GRPPIaccommodates a layer between developers and existing parallel programming frameworks targeting multi‐core processors, such as C++ threads, OpenMP and Intel TBB, and accelerators, as CUDA Thrust. Furthermore, thanks to its high‐level C++ application programming interface and pattern composability features,GRPPIallows users to easily expose parallelism via standalone patterns or patterns compositions matching in sequential applications. We evaluate this interface using an image processing use case and demonstrate its benefits from the usability, flexibility, and performance points of view. Furthermore, we analyze the impact of using stream and data pattern compositions on CPUs, GPUs and heterogeneous configurations. David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Enabling semantics to improve detection of data races and misuses of lock-free data structuresabstractSummary The rapid progress of multi/many‐core architectures has caused data‐intensive parallel applications not yet fully optimized to deliver the best performance. In the advent of concurrent programming, frameworks offering structured patterns have alleviated developers' burden adapting such applications to multithreaded architectures. While some of these patterns are implemented using synchronization primitives, others avoid them by means of lock‐free data mechanisms. However, lock‐free programming is not straightforward, ensuring an appropriate use of their interfaces can be challenging, since different memory models plus instruction reordering at compiler/processor levels can interfere in the occurrence of data races. The benefits of race detectors are formidable in this sense; however, they may emit false positives if are unaware of the underlying lock‐free structure semantics. To mitigate this issue, this paper extends ThreadSanitizer, a race detection tool, with the semantics of 2 lock‐free data structures: the single‐producer/single‐consumer and the multiple‐producer/multiple‐consumer queues. With it, we are able to drop false positives and detect potential semantic violations. The experimental evaluation, using different queue implementations on a set ofμbenchmarks and real applications, demonstrates that it is possible to reduce, on average, 60% the number of data race warnings and detect wrong uses of these structures. Manuel F. Dolz, David del Rio Astorga, Javier Fernández 0001, Massimo Torquati, José Daniel García, Félix García Carballeira, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Improving MPI applications with a new MPI_Info and the use of the memoizationabstractThe MPI forum is actively working for a better MPI standard. The results are the new version 3 of the MPI standard, and the efforts for the incoming MPI 3.1/4.0. The technological changes provide many opportunities for improvements and new ideas. This paper introduces two main contributions in this direction: (1) how to improve the MPI_Info object implementation, and (2) a new way of using the former improved MPI_Info object as a storage solution. Alejandro Calderón 0001, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Daniel Higuero, Borja Bergua |
EuroMPI | 4 |
| 2012 | An Adaptive, Scalable, and Portable Technique for Speeding Up MPI-Based Applications
Rosa Filgueira, Malcolm P. Atkinson 0001, Alberto Nuñez, Javier Fernández 0001 |
Euro-Par | 4 |
| 2012 | A Comparative Evaluation of Parallel Programming Models for Shared-Memory ArchitecturesabstractNowadays, most computers that are commercially available off-the-shelf (COTS) include hardware features that increase the performance of parallel general-purpose threads (hyper threading, multicore, ccNUMA architectures) or SIMD kernels (CPU vector instructions, GPUs). The purpose of this paper is to perform a compared evaluation of several parallel programming models where each one is fitted to exploit some of these features but also each one requires a different level of programming skills. Four parallel programming models (OpenMP, Intel TBB, Intel ArBB, and CUDA) have been selected. The idea is to cover a wide spectrum of programming models and most of the parallel hardware features included in modern computers. On one hand, OpenMP and TBB platforms, that exploits parallel threads running on multicore systems. On the other hand, ArBB, that combines multicore parallel threads and multicore SIMD features with a simpler programming model, and CUDA that exploits SIMD features of the GPU hardware. Our results obtained with the benchmarks used on this paper suggest that OpenMP and TBB have a lower performance compared to ArBB and CUDA. But also that ArBB performance tends to be comparable with CUDA performance in most cases (although it is normally lower). Thus, there are evidences that a careful designed top range multicore and multisocket architecture, can be comparable in terms of performance with top range GPU cards for many applications, with the advantage of a simpler programming model. Luis Miguel Sánchez, Javier Fernández 0001, Rafael Sotomayor, José Daniel García |
ISPA | 2 |
| 2011 | Optimizing Distributed Architectures to Improve Performance on Checkpointing ApplicationsabstractNowadays, satisfying the global throughput targets of each application in High Performance Computing systems is a difficult task because of the high number of architectural configurations having a considerable impact on the overall system performance, such as the number of storage servers, features of the communication links, number of CPU cores per node, etc. In this paper we have performed a thorough study of the compared performance of scaling up HPC cluster architectures using a checkpointing application model. This study is specifically focused on multi-core HPC clusters and the scaling process is oriented towards the three main resources: computing power, communications and storage. The main goal of this work is to evaluate and analyze how evolves both scalability and bottlenecks existent on different HPC multi-core architectures using different architectural configurations. In order to achieve this goal, a set of simulation experiments has been achieved using a simulation framework, called SIMCAN, specifically designed for modeling and simulating HPC architectures. The results obtained show that the computing power is well suited thanks to the multi-core processors, while the problems are found on the storage and on the communications channels, being the storage network the main bottleneck. Alberto Nuñez, Javier Fernández 0001, Jesús Carretero 0001, Laura Prada, Mario Blaum |
HPCC | 2 |
| 2010 | New Contributions for Simulating Large Distributed SystemsabstractNowadays, simulation of large distributed environments is a very important research field. Due to the large number of components to be simulated, the execution of those simulations requires a high level of resources such as CPU and memory. Thus, in order to increase the performance of those simulations, a feasible solution consists on splitting the complete model in sub-domains, where each sub-domain is executed in a single machine. In this paper we propose a strategy to automatically accomplish the parallelization of those environments, which has been implemented and tested in the SIMCAN simulation platform. Alberto Nuñez, Javier Fernández 0001, Jesús Carretero 0001 |
DS-RT | 2 |
| 2010 | Branch replication scheme: A new model for data replication in large scale data grids
José María Pérez, Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, Javier Fernández 0001 |
Future Gener. Comput. Syst. | 5 |
| 2010 | New techniques for simulating high performance MPI applications on large storage networks
Alberto Nuñez, Javier Fernández 0001, José Daniel García, Félix García Carballeira, Jesús Carretero 0001 |
J. Supercomput. | 2 |
| 2009 | Fault tolerant file models for parallel file systems: introducing distribution patterns for every file
Alejandro Calderón 0001, Félix García Carballeira, Luis Miguel Sánchez, José Daniel García, Javier Fernández 0001 |
J. Supercomput. | 5 |
| 2008 | New techniques for simulating high performance MPI applications on large storage networksabstractIn this paper we present new techniques for simulating high performance MPI applications on large storage networks. Performance analysis of high performance application on large storage networks is a very complex and time-consuming task. However, modelling and studying the behaviour of any application on complex network architectures is crucial to obtain good performance. The goal of this work is to predict both scalability degree and performance of any high computing applications on any network architecture. A very interesting feature of this work is that our approach does not require to modify the application in order to simulate its behaviour. Also, there is no need to modify the simulator code to test different architectures. It can be done just creating a new configuration file. In order to perform those analyses we have used SIMCAN, a simulation tool to analyzing high-performance I/O architectures, developed at University Carlos III de Madrid. To validate this work we have used the BIPS3D application on several hardware-based architectures and on our simulator. The comparative results of those environments are presented to show the accuracy and efficiency of our approach. Alberto Nuñez, Javier Fernández 0001, José Daniel García, Jesús Carretero 0001 |
CLUSTER | 2 |
| 2008 | M-PLAT: Multi-Programming Language Adaptive TutorabstractIn this paper we introduce M-PLAT, an intelligent tutoring system for helping students to learn the basics of programming languages. In fact, the M-PLAT system represents a full collection of intelligent tutoring systems, and due to its modular and hierarchical architecture it can be upgraded to deal with a new programming language that is not yet included in the system. Thus, this tutoring system is not limited to a unique programming language, making M-PLAT a very scalable system. The best important feature of our system is that M-PLAT dynamically adapts itself to the learning style of each student, optimizing the learning time to each student. Alberto Nuñez, Javier Fernández 0001, José Daniel García, Laura Prada, Jesús Carretero 0001 |
ICALT | 2 |
| 2008 | Model for on-demand virtual computing architectures - OVCAabstractHigh performance computers are becoming popular for non scientific areas. Growth of computer capacity, virtualization techniques, and network capacity, give an opportunity to exploit the potential of those systems and to integrate different concepts for improving resource utilization. In this paper, we propose an architecture which manages groups of distributed resources, deploying infrastructures on-demand. Resources, no matter if they are local or remote, physical or virtual, are managed homogeneously. Highly customized environments are dynamically created to fulfill clientpsilas requirements, by means of virtual machines. The proposed architecture is based on distributed systems not linked to any specific implementation, which offers high flexibility and applicability. A functional prototype has been implemented and tested at our lab. Evaluation results demonstrate that our initial assumptions are correct. Alejandra Rodríguez, Javier Fernández 0001, Jesús Carretero 0001 |
ISCC | 2 |
| 2007 | Dispatching Requests in Partially Replicated Web Clusters - An Adaptation of the LARD Algorithm
José Daniel García, Laura Prada, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Luis Miguel Sánchez |
WEBIST (1) | 5 |
| 2006 | On the Reliability of Web Clusters with Partial Replication of ContentsabstractTraditionally, distributed Web servers have used two strategies for allocating files on server nodes: full replication and full distribution. While full replication provides a highly reliable solution, it limits storage capacity to the capacity of the smallest node. On the other hand, full distribution provides higher storage capacity at the cost of lower reliability. A hybrid solution is partial replication where every file is allocated to a small number of nodes. The most promising architecture for a partial replication strategy is the Web cluster architecture. However, Web clusters present a big flaw from reliability perspective as they contain a single point of failure. To correct this flaw, in this paper we present a modified architecture: the Web cluster with distributed Web switch. Reliability of Web clusters is evaluated for different replication strategies. System evaluations show that our proposal leads to a highly reliable solution with high scalability. José Daniel García, Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, David E. Singh, Alejandro Calderón 0001 |
ARES | 3 |
| 2006 | A Quantitative Justification to Partial Replication of Web Contents
José Daniel García, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Alejandro Calderón 0001, David E. Singh |
ICCSA (4) | 4 |
| 2003 | Video Forwarding Techniques for Mixed Wired and Wireless NetworksabstractDuring the last years, Internet video streaming has experiences a phenomenal growth. This is happening despite the notorious difficulties of transmitting data packets with a deadline over the Internet, due to variability in throughput, delays and losses. These problems arise significantly when using wireless networks where the available bandwidth is low and the losses are important due to its error prone transmission nature. In this paper we propose a fast-forwarding technique that is based on segmenting the movie on different files. Normal movie reproduction requires all the files, but fast-forwarding reproduction only requires one file. Those files can me merged by the client or by the server. The segmentation is frame based, grouping all the frames that can be independently decoded together. The resulting file can be showed with any existing player. This group of frames would be the ones to use in a fast-forward reproduction. Our techniques can also be useful in adaptive environments, like wireless networks, because there is no problem for the fast-forward file to use the same optimizations that exist for full movie files. This method also reduces the storage bandwidth and the storage size needed (there is no extra data for fast-forwarding). We also propose a video server architecture that takes advantage of this technique to achieve full interactive video reproduction. The evaluation results shown in this paper demonstrates that our technique enhances video fast-forwarding operations. Javier Fernández 0001, Jesús Carretero 0001, Félix García Carballeira, José María Pérez, Alejandro Calderón 0001, José J. Muñoz |
ISCC | 1 |
| 2003 | A hierarchical disk scheduler for multimedia systems
Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, Alok N. Choudhary |
Future Gener. Comput. Syst. | 2 |
| 2001 | New Techniques for Collective Communications in Clusters: A Case Study with MPIabstractThe paper describes new techniques to increase the performance of collective communication operations in clusters. These techniqnes are based in multithreading operations and on-line data compression. The techniques proposed have been implemented in MiMPI, a thread-safe implementation of MPI. We have evaluated, and compared, the performance of MiMPI with other implementations of MPI available for clusters with Linux and Windows 2000. The benchmark used has been MPBench, a flexible and portable framework to allow benchmarking of MPI implementations. Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001, Javier Fernández 0001, Oscar Pérez |
ICPP | 4 |