VLDB 2026 Research / reviewers in the wild / expert
Juan Fernández Peinador
dblp:f/JuanFernandezPeinador · also Juan Fernández 0001
· DBLP profile ↗
29ranked-venue papers
5as first author
0since 2021 · last 2014
0000-0001-9462-7338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Parallel and multicore computing · 36% Cloud and datacenter computing · 18% Interconnection networks and networks-on-chip · 17% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
many-core systems |
0.1 | 1 | 2012 | Efficient Hardware Barrier Synchronization in Many-Core CMPs · IEEE Trans. Parallel Distributed Syst. 2012 |
Embedded and real-time systems › real-time scheduling › multiprocessor scheduling
gang scheduling |
0.1 | 2 | 2006 | STORM: Scalable Resource Management for Large-Scale Parallel Computers · IEEE Trans. Computers 2006 Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 |
Parallel and multicore computing
parallel scheduling |
0.1 | 2 | 2006 | STORM: Scalable Resource Management for Large-Scale Parallel Computers · IEEE Trans. Computers 2006 Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.1 | 2 | 2006 | STORM: Scalable Resource Management for Large-Scale Parallel Computers · IEEE Trans. Computers 2006 STORM: lightning-fast resource management · SC 2002 |
Cloud and datacenter computing
job scheduling |
0.1 | 2 | 2005 | Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 STORM: lightning-fast resource management · SC 2002 |
Parallel and multicore computing › parallel scheduling
coscheduling |
0.1 | 1 | 2005 | Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2005 | Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 |
High-performance computing
collective communication |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Storage systems
data reduction |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Interconnection networks and networks-on-chip › network interface
network interface card |
0.0 | 1 | 2003 | Scalable NIC-based Reduction on Large-scale Clusters · SC 2003 |
Parallel and multicore computing
parallel programming models and runtimes |
0.0 | 1 | 2003 | BCS-MPI: A New Approach in the System Software Design for Large-Scale Parallel Computers · SC 2003 |
Performance modeling and evaluation › network performance analysis
protocol performance analysis |
0.0 | 1 | 2003 | BCS-MPI: A New Approach in the System Software Design for Large-Scale Parallel Computers · SC 2003 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2005 | Adaptive Parallel Job Scheduling with Flexible Coscheduling · IEEE Trans. Parallel Distributed Syst. 2005 |
Cloud and datacenter computing
resource management |
0.0 | 1 | 2002 | STORM: lightning-fast resource management · SC 2002 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2performance modeling · 0.1implicit coscheduling · 0.1gang scheduling · 0.1batch scheduling · 0.1collective communication primitives · 0.0collective algorithm design · 0.0buffered coscheduling · 0.0NIC-based reduction · 0.0low-level network features · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Selective dynamic serialization for reducing energy consumption in hardware transactional memory systems
Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2013 | On the design of energy-efficient hardware transactional memory systemsabstractSUMMARY Transactional memory is currently being advocated as a promising alternative to lock‐based synchronization because it simplifies multithreaded programming. In this way, future many‐core chip multiprocessor architectures may need to provide hardware support for transactional memory. On the other hand, energy consumption constitutes nowadays a first class consideration in multicore processor designs. In this work, we characterize the performance and energy consumption of two well‐known hardware transactional memory systems that employ opposite policies for data versioning and conflict management. More specifically, we compare a LogTM‐SEeager‐eagersystem and a version of the Scalable Transactional Coherence and Consistencylazy‐lazysystem that enable parallel commits. To do so, we extended the Multifacet GEMS simulator to estimate the energy consumed in the on‐chip caches according to CACTI and used the interconnection network energy model given by Orion 2. Results show that the energy consumption of the eager‐eager system is 38% higher in average than in the lazy‐lazy case, whereas performance differences between the two systems are 26% in average. We found that even though lazy‐lazy beats eager‐eager on average, there are considerable deviations in performance depending on the particular characteristics of each application and the settings of both systems. Finally, from this characterization, we observe that a significant part of the energy consumed in some applications in eager‐eager is spent on the back‐off delay phase and explore more energy‐efficient hardware back‐off mechanisms. For lazy‐lazy systems, the way in which memory lines are assigned to the L2 cache banks affects the number of parallel commits in some applications, and we study an alternative fine‐grained assignment. Copyright © 2012 John Wiley & Sons, Ltd. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Design of an efficient communication infrastructure for highly contended locks in many-core CMPs
José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
J. Parallel Distributed Comput. | 2 |
| 2012 | Design of a collective communication infrastructure for barrier synchronization in cluster-based nanoscale MPSoCsabstractBarrier synchronization is a key programming primitive for shared memory embedded MPSoCs. As the core count increases, software implementations cannot provide the needed performance and scalability, thus making hardware acceleration critical. In this paper we describe an interconnect extension implemented with standard cells and with a mainstream industrial toolflow. We show that the area overhead is marginal with respect to the performance improvements of the resulting hardware-accelerated barriers. We integrate our HW barrier into the OpenMP programming model and discuss synchronization efficiency compared with traditional software implementations. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, Davide Bertozzi, Daniele Bortolotti, Andrea Marongiu, Luca Benini |
DATE | 2 |
| 2012 | Dynamic Serialization: Improving Energy Consumption in Eager-Eager Hardware Transactional Memory SystemsabstractIn the search for new paradigms to simplify multithreaded programming, Transactional Memory (TM) is currently being advocated as a promising alternative to deadlock-prone lock-based synchronization. In this way, future many-core CMP architectures may need to provide hardware support for TM. On the other hand, power dissipation constitutes a first class consideration in multicore processor designs. In this work, we propose Dynamic Serialization (DS) as a new technique to improve energy consumption without degrading performance in applications with conflicting transactions. Our proposal, which is implemented on top of a hardware transactional memory system with an eager conflict management policy, detects and serializes conflicting transactions dynamically. Particularly, in case of conflict one transaction is allowed to continue whilst the rest are completely stalled. Once the executing transaction has finished it wakes up several of the stalling transactions. This brings important benefits in terms of energy consumption due to the reduction in the amount of wasted work that DS implies. Results for a 16-core CMP show that Dynamic Serialization obtains reductions of 10% on average in energy consumption (more than 20% in high contention scenarios) without affecting, on average, execution time. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Manuel E. Acacio, Juan Fernández Peinador |
PDP | 4 |
| 2012 | Stencil computations on heterogeneous platforms for the Jacobi method: GPUs versus Cell BE
José M. Cecilia, José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, José M. García 0001, Manuel Ujaldon |
J. Supercomput. | 3 |
| 2012 | Efficient Hardware Barrier Synchronization in Many-Core CMPsabstractTraditional software-based barrier implementations for shared memory parallel machines tend to produce hotspots in terms of memory and network contention as the number of processors increases. This could limit their applicability to future many-core CMPs in which possibly several dozens of cores would need to be synchronized efficiently. In this work, we develop GBarrier, a hardware-based barrier mechanism especially aimed at providing efficient barriers in future many-core CMPs. Our proposal deploys a dedicated G-line-based network to allow for fast and efficient signaling of barrier arrival and departure. Since GBarrier does not have any influence on the memory system, we avoid all coherence activity and barrier-related network traffic that traditional approaches introduce and that restrict scalability. Through detailed simulations of a 32-core CMP, we compare GBarrier against one of the most efficient software-based barrier implementations for a set of kernels and scientific applications. Evaluation results show average reductions of 54 and 21 percent in execution time, 53 and 18 percent in network traffic, and also 76 and 31 percent in the energy-delay2product metric for the full CMP when the kernels and scientific applications, respectively, are considered. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | GLocks: Efficient Support for Highly-Contended Locks in Many-Core CMPsabstractSynchronization is of paramount importance to exploit thread-level parallelism on many-core CMPs. In these architectures, synchronization mechanisms usually rely on shared variables to coordinate multithreaded access to shared data structures thus avoiding data dependency conflicts. Lock synchronization is known to be a key limitation to performance and scalability. On the one hand, lock acquisition through busy waiting on shared variables generates additional coherence activity which interferes with applications. On the other hand, lock contention causes serialization which results in performance degradation. This paper proposes and evaluates \textit{GLocks}, a hardware-supported implementation for highly-contended locks in the context of many-core CMPs. \textit{GLocks} use a token-based message-passing protocol over a dedicated network built on state-of-the-art technology. This approach skips the memory hierarchy to provide a non-intrusive, extremely efficient and fair lock implementation with negligible impact on energy consumption or die area. A comprehensive comparison against the most efficient shared-memory-based lock implementation for a set of micro benchmarks and real applications quantifies the goodness of \textit{GLocks}. Performance results show an average reduction of 42% and 14% in execution time, an average reduction of 76% and 23% in network traffic, and also an average reduction of 78% and 28% in energy-delay$^2$ product (ED$^2$P) metric for the full CMP for the micro benchmarks and the real applications, respectively. In light of our performance results, we can conclude that \textit{GLocks} satisfy our initial working hypothesis. \textit{GLocks} minimize cache-coherence network traffic due to lock synchronization which translates into reduced power consumption and execution time. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
IPDPS | 2 |
| 2010 | A G-Line-Based Network for Fast and Efficient Barrier Synchronization in Many-Core CMPsabstractBarrier synchronization in shared memory parallel machines has been widely implemented through busy-waiting on shared variables. However, typical implementations of barrier synchronization tend to produce hot-spots in terms of memory and network contention, thus creating performance bottlenecks that become markedly more pronounced as the number of cores or processors increases. To overcome such limitations, we present a novel hardware-based barrier mechanism in the context of many-core CMPs. Our proposal is based on global interconnection lines (G-lines) and the S-CSMA technique, which have been recently used to enhance a flow control mechanism (EVC) in the context of networks-on-chip. Based on this technology, we have designed a simple and scalable G-line-based network that operates independently of the main data network, and that is aimed at carrying out barrier synchronizations efficiently. In the ideal case, our design takes only 4 cycles to perform a barrier synchronization once all cores or threads have arrived at the barrier. As a proof of concept, we examine the benefits of our proposal by comparing it with one of the best software approaches (a binary combining-tree barrier). To do so, we run several kernels and scientific applications on top of the Sim-PowerCMP performance simulator that models a 32-core CMP with a 2D-mesh network configuration. Our proposal entails average reductions in terms of execution time of 68% and 21% for kernels and scientific applications, respectively. Additionally, network traffic is also lowered by 74% and 18%, respectively. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
ICPP | 2 |
| 2010 | Characterizing Energy Consumption in Hardware Transactional Memory SystemsabstractTransactional Memory is currently being advocated as a promising alternative to lock-based synchronization because it simplifies multithreaded programming. In this way, future many-core CMP architectures may need to provide hardware support for transactional memory. On the other hand, power dissipation constitutes a first class consideration in multicore processor design. In this work, we characterize the performance and energy consumption of two well-known Hardware Transactional Memory systems that employ opposite policies for data versioning and conflict management. More specifically, we compare the Log TM-SE Eager-Eager system and a version of the Scalable TCC Lazy-Lazy system that enables parallel commits. To the best of our knowledge, this is the first characterization in terms of energy consumption of hardware transactional memory systems. To do that, we extended the GEMS simulator to estimate the energy consumed in the on-chip caches according to CACTI, and used the interconnection network energy model given by Orion 2. Results show that the energy consumption of the Eager-Eager system is 60% higher on average than in the Lazy-Lazy case, whereas performance differences between the two systems are 42% on average. Finally, we found that although on average Lazy-Lazy beats Eager-Eager there are considerable deviations in performance depending on the particular characteristics of each application. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
SBAC-PAD | 3 |
| 2010 | Characterizing the basic synchronization and communication operations in Dual Cell-based Blades through CellStats
José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
J. Supercomput. | 2 |
| 2009 | Fast and Efficient Synchronization and Communication Collective Primitives for Dual Cell-Based Blades
Epifanio Gaona-Ramírez, Juan Fernández Peinador, Manuel E. Acacio |
Euro-Par | 2 |
| 2009 | A Parallel Implementation of the 2D Wavelet Transform Using CUDAabstractThere is a multicore platform that is currently concentrating an enormous attention due to its tremendous potential in terms of sustained performance: the NVIDIA Tesla boards. These cards intended for general-purpose computing on graphic processing units (GPGPUs) are used as data-parallel computing devices. They are based on the Computed Unified Device Architecture (CUDA) which is common to the latest NVIDIA GPUs. The bottom line is a multicore platform which provides an enormous potential performance benefit driven by a non-traditional programming model. In this paper we try to provide some insight into the peculiarities of CUDA in order to target scientific computing by means of a specific example. In particular, we show that the parallelization of the two-dimensional fast wavelet transform for the NVIDIA Tesla C870 achieves a speedup of 20.8 for an image size of 8192x8192, when compared with the fastest host-only version implementation using OpenMP and including the data transfers between main memory and device memory. Joaquín Franco, Gregorio Bernabé, Juan Fernández Peinador, Manuel E. Acacio |
PDP | 3 |
| 2008 | CellStats: A Tool to Evaluate the Basic Synchronization and Communication Operations of the Cell BEabstractThe Cell Broadband Engine (Cell BE) is a recent heterogeneous chip-multiprocessor (CMP) architecture jointly developed by IBM, Sony and Toshiba to offer very high performance, especially on game and multimedia applications. The significant number of processor cores that it contains (nine in its first generation), along with their heterogeneity (they are of two different types) and the variety of synchronization and communication primitives offered to programmers, make the task of developing efficient applications for the Cell BE very challenging. In this work, we present CellStats, a tool aimed at characterizing the performance of the main synchronization and communication primitives provided by the Cell BE under varying workloads. In particular, the current implementation of CellStats allows to evaluate the DMA transfer mechanism, the read-modify-write atomic operations, the mailboxes, the signals and the time taken by thread creation. As an example of application of CellStats, we present a characterization of the Cell BE incorporated into the PlayStation 3. From this characterization, we extract some recommendations that can help programmers to identify the most appropriate primitive under different assumptions. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
PDP | 2 |
| 2007 | Multicore Surprises: Lessons Learned from Optimizing Sweep3D on the Cell Broadband EngineabstractThe Cell Broadband Engine (BE) processor provides the potential to achieve an impressive level of performance for scientific applications. This level of performance can be reached by exploiting several dimensions of parallelism, such as thread-level parallelism using several synergistic processing elements, data streaming parallelism, vector parallelism in the form of 128-bit SIMD operations, and pipeline parallelism by issuing multiple instructions in the same clock cycle. In our exploration to achieve the optimum level of performance for Sweep3D, we have enjoyed many pleasant surprises, such as a very high floating point performance, reaching 64% of the theoretical peak in double precision, and an over all performance speedup ranging from 4.5 times when compared with "heavy iron" processors, up to over 20 times with conventional processors. Fabrizio Petrini, Gordon C. Fossum, Juan Fernández Peinador, Ana Lucia Varbanescu, Michael Kistler, Michael Perrone |
IPDPS | 3 |
| 2007 | Challenges in Mapping Graph Exploration Algorithms on Advanced Multi-core ProcessorsabstractMulti-core processors are a shift of paradigm in computer architecture that promises a dramatic increase in performance. But multi-core processors also bring an unprecedented level of complexity in algorithmic design and software development. In this paper we describe the challenges and design choices involved in parallelizing a breadth-first search (BFS) algorithm on a state-of-the-art multi-core processor, the Cell Broadband Engine (Cell BE). Our experiments obtained on a pre-production Cell BE board running at 3.2 GHz show almost linear speedups when using multiple synergistic processing units, and an impressive level of performance when compared to other processors. The Cell BE is typically an order of magnitude faster than conventional processors, such as the AMD Opteron and the Intel Pentium 4 and Woodcrest, an order of magnitude faster than the MTA-2 multi-threaded processor, and two orders of magnitude faster than a BlueGene/L processor. Oreste Villa, Daniele Paolo Scarpazza, Fabrizio Petrini, Juan Fernández Peinador |
IPDPS | 4 |
| 2006 | An Abstract Interface for System Software on Large-Scale ClustersabstractScalable management of distributed resources is one of the major challenges when building large-scale clusters for high-performance computing. This task includes transparent fault tolerance, efficient deployment of resources and support for all the needs of parallel applications: parallel I/O, deterministic behavior and responsiveness. These challenges may seem daunting with commodity hardware and operating systems, since they were not designed to support a global, single management view of a large-scale system. In this paper we propose and demonstrate an abstract network interface in the cluster interconnect to facilitate the implementation of a simple yet powerful global operating system. This system, which can be thought of as a coarse-grain SIMD operating system, can allow commodity clusters to grow to thousands of nodes, while still retaining the usability and performance of the single-node workstation. Juan Fernández Peinador, Eitan Frachtenberg, Fabrizio Petrini, José Carlos Sancho |
Comput. J. | 1 |
| 2006 | STORM: Scalable Resource Management for Large-Scale Parallel ComputersabstractAlthough clusters are a popular form of high-performance computing, they remain more difficult to manage than sequential systems - or even symmetric multiprocessors. In this paper, we identify a small set of primitive mechanisms that are sufficiently general to be used as building blocks to solve a variety of resource-management problems. We then present STORM, a resource-management environment that embodies these mechanisms in a scalable, low-overhead, and efficient implementation. The key innovation behind STORM is a modular software architecture that reduces all resource management functionality to a small number of highly scalable mechanisms. These mechanisms simplify the integration of resource management with low-level network features. As a result of this design, STORM can launch large, parallel applications an order of magnitude faster than the best time reported in the literature and can gang-schedule a parallel application as fast as the node OS can schedule a sequential application. This paper describes the mechanisms and algorithms behind STORM and presents a detailed performance model that shows that STORM's performance can scale to thousands of nodes Eitan Frachtenberg, Fabrizio Petrini, Juan Fernández Peinador, Scott Pakin |
IEEE Trans. Computers | 3 |
| 2005 | Adaptive Parallel Job Scheduling with Flexible CoschedulingabstractMany scientific and high-performance computing applications consist of multiple processes running on different processors that communicate frequently. Because of their synchronization needs, these applications can suffer severe performance penalties if their processes are not all coscheduled to run together. Two common approaches to coscheduling jobs are batch scheduling, wherein nodes are dedicated for the duration of the run, and gang scheduling, wherein time slicing is coordinated across processors. Both work well when jobs are load-balanced and make use of the entire parallel machine. However, these conditions are rarely met and most realistic workloads consequently suffer from both internal and external fragmentation, in which resources and processors are left idle because jobs cannot be packed with perfect efficiency. This situation leads to reduced utilization and suboptimal performance. Flexible coscheduling (FCS) addresses this problem by monitoring each job's computation granularity and communication pattern and scheduling jobs based on their synchronization and load-balancing requirements. In particular, jobs that do not require stringent synchronization are identified, and are not coscheduled; instead, these processes are used to reduce fragmentation. FCS has been fully implemented on top of the STORM resource manager on a 256-processor alpha cluster and compared to batch, gang, and implicit coscheduling algorithms. This paper describes in detail the implementation of FCS and its performance evaluation with a variety of workloads, including large-scale benchmarks, scientific applications, and dynamic workloads. The experimental results show that FCS saturates at higher loads than other algorithms (up to 54 percent higher in some cases), and displays lower response times and slowdown than the other algorithms in nearly all scenarios. Eitan Frachtenberg, Dror G. Feitelson, Fabrizio Petrini, Juan Fernández Peinador |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2004 | Designing Parallel Operating Systems via Parallel Programming
Eitan Frachtenberg, Kei Davis, Fabrizio Petrini, Juan Fernández Peinador, José Carlos Sancho |
Euro-Par | 4 |
| 2004 | Architectural Support for System Software on Large-Scale ClustersabstractScalable management of distributed resources is one of the major challenges in deployment of large-scale clusters. Management includes transparent fault tolerance, efficient allocation of resources, and support for all the needs of parallel computing: parallel I/O, deterministic behavior, and responsiveness. Meeting these requirements with commodity hardware and operating systems is difficult because they were not designed to support global management of a large-scale system. We propose a small set of hardware mechanisms in the cluster interconnect to facilitate the implementation of a simple yet powerful global operating system. This system, inspired by concepts from the BSP and SIMD computational models, allows commodity clusters to grow to thousands of nodes while still retaining the usability and responsiveness of the single-node workstation. Our results on a software prototype show that it is possible to implement efficient and scalable system software using the proposed set of mechanisms. Juan Fernández Peinador, Eitan Frachtenberg, Fabrizio Petrini |
ICPP | 1 |
| 2004 | On the Feasibility of Incremental Checkpointing for Scientific ComputingabstractSummary form only given. In the near future large-scale parallel computers will feature hundreds of thousands of processing nodes. In such systems, fault tolerance is critical as failures will occur very often. Checkpointing and rollback recovery has been extensively studied as an attempt to provide fault tolerance. However, current implementations do not provide the total transparency and full flexibility that are necessary to support the new paradigm of autonomic computing - systems able to self-heal and self-repair. We provide an in-depth evaluation of incremental checkpointing for scientific computing. The experimental results, obtained on a state-of-the art cluster running several scientific applications, show that efficient, scalable, automatic and user-transparent incremental checkpointing is within reach with current technology. José Carlos Sancho, Fabrizio Petrini, Greg Johnson, Juan Fernández Peinador, Eitan Frachtenberg |
IPDPS | 4 |
| 2003 | Parallel Job Scheduling under Dynamic Workloads
Eitan Frachtenberg, Dror G. Feitelson, Juan Fernández Peinador, Fabrizio Petrini |
JSSPP | 3 |
| 2003 | BCS-MPI: A New Approach in the System Software Design for Large-Scale Parallel ComputersabstractBuffered CoScheduled MPI (BCS-MPI) introduces a new approach to design the communication layer for large-scale parallel machines. The emphasis of BCS-MPI is on the global coordination of a large number of communicating processes rather than on the traditional optimization of the point-to-point performance. BCS-MPI delays the inter-processor communication in order to schedule globally the communication pattern and it is designed on top of a minimal set of collective communication primitives. In this paper we describe a prototype implementation of BCS-MPI and its communication protocols. Several experimental results, executed on a set of scientific applications, show that BCS-MPI can compete with a production-level MPI implementation, but is much simpler to implement, debug and model. Juan Fernández Peinador, Eitan Frachtenberg, Fabrizio Petrini |
SC | 1 |
| 2003 | Scalable NIC-based Reduction on Large-scale ClustersabstractMany parallel algorithms require efficient reduction collectives. In response, researchers have designed algorithms considering a range of parameters including data size, system size, and communication characteristics. Throughout this past work, however, processing was limited to the host CPU. Today, modern Network Interface Cards (NICs) sport programmable processors with substantial memory, and thus introduce a fresh variable into the equation. In this paper, we investigate this new option in the context of large-scale clusters. Through experiments on the 960-node, 1920-processor ASCI Linux Cluster (ALC) at Lawrence Livermore National Laboratory, we show that NIC-based reductions outperform host-based algorithms in terms of reduced latency and increased consistency. In particular, in the largest configuration tested - 1812 processors - our NIC-based algorithm summed single-element vectors of 32-bit integers and 64-bit floating-point numbers in 73 µs and 118 µs, respectively. These results represent respective improvements of 121% and 39% over the production-level MPI library. Adam Moody, Juan Fernández Peinador, Fabrizio Petrini, Dhabaleswar K. Panda 0001 |
SC | 2 |
| 2002 | Scalable Resource Management in High Performance ComputersabstractClusters of workstations have emerged as an important platform for building cost-effective, scalable, and highly-available computers. Although many hardware solutions are available today, the largest challenge in making largescale clusters usable lies in the system software. In this paper we present STORM, a resource management tool designed to provide scalability, low overhead, and the flexibility necessary to efficiently support and analyze a wide range of job-scheduling algorithms. STORM achieves these feats by using a small set of primitive mechanisms that are common in modern high-performance interconnects. The architecture of STORM is based on three main technical innovations. First, a part of the scheduler runs in the thread processor located on the network interface. Second, we use hardware collectives that are highly scalable both for implementing control heartbeats and to distribute the binary of a parallel job in near-constant time. Third, we use an I/O bypass protocol that allows fast data movements front the file system to the communication buffers in the network interface and vice versa. The experimental results show that STORM can launch a job with a binary of 12 MB on a 64-processor, 32-node cluster in less than 250 ms. This paper provides expert. mental and analytical evidence that these results scale to a much larger number of nodes. To the best of our knowledge, STORM significantly outperforms existing production schedulers in launching jobs, performing resource management tasks, and gang-scheduling tasks. Eitan Frachtenberg, Fabrizio Petrini, Juan Fernández Peinador, Salvador Coll |
CLUSTER | 3 |
| 2002 | Improving the Performance of Real-Time Communication Services on High-Speed LANs under Topology ChangesabstractIn this paper, we propose and evaluate a new protocol that provides topology change- and fault-tolerant real-time communication services on NOW and clusters. This protocol overcomes the main drawback of our previously proposed protocol, called Dynamically Re-established Real-Time Channels (DRRTC), which is physically limited by the number of virtual channels per port. The new protocol allows different real-time channels to share the same virtual channel. In this way, the new protocol allows us to establish a greater number of real-time channels than the previous one. Moreover, its only limitation is the bandwidth devoted to real-time traffic. However, this introduces two new problems that are successfully managed by the new protocol: the existence of cyclic dependencies among different real-time channels and the increased complexity of deadline requirements. We present and analyze the performance evaluation results when a single switch or a single link is deactivated/activated for different topologies and workloads. The new protocol overwhelms the DRRTC protocol while guaranteeing deadline requirements and channel recovery. Juan Fernández Peinador, José M. García 0001, José Duato |
LCN | 1 |
| 2002 | STORM: lightning-fast resource managementabstractAlthough workstation clusters are a common platform for high-performance computing (HPC), they remain more difficult to manage than sequential systems or even symmetric multiprocessors. Furthermore, as cluster sizes increase, the quality of the resource-management subsystem — essentially, all of the code that runs on a cluster other than the applications — increasingly impacts application efficiency. In this paper, we present STORM, a resource-management framework designed for scalability and performance. The key innovation behind STORM is a software architecture that enables resource management to exploit low-level network features. As a result of this HPC-application-like design, STORM is orders of magnitude faster than the best reported results in the literature on two sample resource-management functions: job launching and process scheduling. Eitan Frachtenberg, Fabrizio Petrini, Juan Fernández Peinador, Scott Pakin, Salvador Coll |
SC | 3 |
| 2001 | Performance Evaluation of Real-Time Communication Services on High-Speed LANs under Topology Changes
Juan Fernández Peinador, José M. García 0001, José Duato |
HiPC | 1 |