VLDB 2026 Research / reviewers in the wild / expert
Jesús Labarta
dblp:87/4934 · also Jesús Labarta Mancho
· DBLP profile ↗
178ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-7489-4727ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 146 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 4Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Compiler-Assisted Workflow for Efficiency-Guided Selective Tracing
Sebastian Kreutzer, Valentin Seitz, Joan Vinyals-Ylla-Catala, Tim Heldmann, Christian Iwainsky, Marta Garcia-Gasulla, Jesús Labarta, Christian H. Bischof |
Euro-Par (1) | 7 |
| 2026 | Designing a QEMU plugin to profile multicore long vector RISC-V architectures: RAVE
Pablo Vizcaino, Roger Ferrer, Jesús Labarta, Filippo Mantovani |
Future Gener. Comput. Syst. | 3 |
| 2025 | NSYS2PRV: Detailed and Quantitative Analysis of Large-Scale GPU Execution Traces with ParaverabstractThis work presents a tool, a methodology, a set of metrics, and practical examples for evaluating the performance of large-scale AI and traditional HPC applications using GPUs. NSYS2PRV is a tool that converts NVIDIA Nsight Systems reports into traces compatible with Paraver, enabling significantly enhanced insight compared to current performance analysis practices. By leveraging the capabilities of a well-established HPC performance analysis tool, we enable the comparison of execution traces and the quantification of microscopic-level differences to explain behaviors across hundreds or more computing devices. We argue that large-scale GPU applications and AI workloads can greatly benefit from the type of large-scale performance analysis introduced here, an approach that is not yet widely adopted in this domain. Translating nsys-generated traces to Paraver allows analysts to combine the fine-grained, highly accurate execution data obtainable from proprietary tools with the flexibility and scalability of an open-source, parallel performance analysis environment. Paraver also enables easy, customizable computation of efficiency metrics. This work demonstrates a more effective and insightful analysis experience than that offered by the native visualization tools in Nsight Systems. Additionally, we introduce a set of Paravercompatible metrics that guide the analysis process, and we showcase examples where these metrics were successfully applied to real-world AI and HPC workloads. Marc Clascà, Jesús Labarta, Marta Garcia-Gasulla |
CLUSTER | 2 |
| 2025 | Leveraging iterative applications to improve the scalability of task-based programming models on distributed systemsabstractDistributed tasking models such as OmpSs-2@Cluster, StarPU-MPI, and PaRSEC express HPC applications as task graphs with explicit dependencies. The single task graph unifies the representation of parallelism across CPU cores, accelerators, and distributed-memory nodes, offering higher programmer productivity compared to traditional MPI + X. Most task-based models construct the task graph sequentially, which provides a clear and familiar programming model, simplifying code development, maintenance, and porting. However, this design introduces a bottleneck in task creation and dependency management, limiting performance and scalability. As a result, unless the tasks are very coarse-grained, current distributed sequential tasking models cannot match the performance of MPI + X. Many scientific applications, however, are iterative in nature, constructing the same directed acyclic task graph at each timestep. We exploit this structure to eliminate the sequential bottleneck and control message overhead in a sequentially-constructed distributed tasking model, while preserving its simplicity and productivity. Our approach builts on the recently proposed taskiter directive for OpenMP and OmpSs-2, allowing a single iteration to be expressed as a cyclic graph. The runtime partitions the cyclic graph across nodes, precomputes the MPI transfers, and then executes the loop body at low overhead. By integrating the MPI communications directly into the application’s task graph, our approach naturally overlaps computation and communication, in some cases exposing dramatically more parallelism than fork–join MPI + OpenMP. We define the programming model and describe the full runtime implementation, and integrate our proposal into OmpSs-2@Cluster. We evaluate it using five benchmarks on up to 128 nodes of the MareNostrum 5 supercomputer. For applications with fork–join parallelism, our approach has performance similar to fork–join MPI + OpenMP, making it a viable productive alternative, unlike the existing OmpSs-2@Cluster model, which is up to 7.7 times slower than MPI + OpenMP. For a 2D Gauss–Seidel stencil computation, our approach enables 3D wavefront computation, giving performance up to 22 times faster than fork–join MPI + OpenMP and on-a-par with state-of-the-art TAMPI + OmpSs-2. All software, comprising the compiler, runtime, and benchmarks, is released open source. 1 Omar Shaaban Ibrahim ali, Juliette Fournis d'Albiat, Isabel Piedrahita, Vicenç Beltran 0001, Xavier Martorell, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | A Mess of Memory System Benchmarking, Simulation and Application ProfilingabstractThe Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detailed memory system characterization. It is based on hundreds of measurements that are represented as a family of bandwidth-latency curves. The benchmark increases the coverage of all the previous tools and leads to new findings in the behavior of the actual and simulated memory systems. We deploy the Mess benchmark to characterize Intel, AMD, IBM, Fujitsu, Amazon and NVIDIA servers with DDR4, DDR5, HBM2 and HBM2E memory. The Mess memory simulator uses bandwidth-latency concept for the memory performance simulation. We integrate Mess with widely-used CPUs simulators enabling modeling of all high-end memory technologies. The Mess simulator is fast, easy to integrate and it closely matches the actual system performance. By design, it enables a quick adoption of new memory technologies in hardware simulators. Finally, the Mess application profiling positions the application in the bandwidth-latency space of the target memory system. This information can be correlated with other application runtime activities and the source code, leading to a better overall understanding of the application's behavior. The current Mess benchmark release covers all major CPU and GPU ISAs, x86, ARM, Power, RISC-V, and NVIDIA's PTX. We also release as open source the ZSim, gem5 and OpenPiton Metro-MPI integrated with the Mess simulator for DDR4, DDR5, Optane, HBM2, HBM2E and CXL memory expanders. The Mess application profiling is already integrated into a suite of production HPC performance analysis tools. Pouya Esmaili-Dokht, Francesco Sgherzi, Valéria Soldera Girelli, Isaac Boixaderas, Mariana Carmin, Alireza Monemi, Adrià Armejach, Estanislao Mercadal, Germán Llort, Petar Radojkovic, Miquel Moretó, Judit Giménez, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, Emanuele Confalonieri, Rishabh Dubey, Jason Adlard |
MICRO | 15 |
| 2024 | $\mathcal{O}(n)$O(n) Key-Value Sort With Active Compute MemoryabstractWe propose the Active Compute Memory (ACM), a near-memory-processing architecture capable of performing key–value sort directly in the DRAM. In the ACM architecture, sort is merely the writing of data into memory with one addressing protocol (perspective) and reading it back with different perspective. The first perspective is conventional, based on the data address; the second perspective is the sorted order. The ACM requires additional tables to store the meta-data and moderate control logic enhancements that can be implemented directly in the DRAM silicon. By these modest enhancements to DRAM, ACM exploits the parallelism inherently available in the row buffer to enable sort with$O(n)$complexity. This leads to an order of magnitude improvement in ACM performance and energy compared to conventional$O(n\log{}n)$CPU-centric sort algorithms. The ACM also shows superior performance compared to other near-memory sort accelerators. This is because the ACM processing is done near the row buffer and it exploits much lower memory access latency, higher bandwidth and wider parallel processing. The sort operation covered in this paper is just an example of an address management operation that can be efficiently implemented directly in the DRAM silicon. We release as an open source the simulation infrastructure for the ACM performance and energy modeling. We would encourage the community to use it, adapt it to other PIM proposals, and share their own evaluations. Pouya Esmaili-Dokht, Miquel Guiot, Petar Radojkovic, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, Jason Adlard, Paolo Amato, Marco Sforzin |
IEEE Trans. Computers | 6 |
| 2023 | Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extensionabstractSummary Novel architectures leveraging long and variable vector lengths like the NEC SX‐Aurora or the vector extension of RISCV are appearing as promising solutions on the supercomputing market. These architectures often require re‐coding of scientific kernels. For example, traditional implementations of algorithms for computing the fast Fourier transform (FFT) cannot take full advantage of vector architectures. In this article, we present the implementation of FFT algorithms able to leverage these novel architectures. We evaluate these codes on NEC SX‐Aurora , comparing them with the optimized NEC libraries; and in a prototype of a RISC‐V core with a vector processing unit. We present the benefits and limitations of two approaches of RADIX‐2 FFT vector implementations. We show that our approach makes better use of the vector unit of the NEC SX‐Aurora , reaching higher or equal performance than the optimized NEC library. More generally, we prove the importance of maximizing the vector length usage of the algorithm, taking advantage of the FFT properties to reduce long‐latency vector operations, and reordering the instructions according to the specific hardware features to boost the performance of FFT‐like computational kernels. Pablo Vizcaino, Filippo Mantovani, Roger Ferrer, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | ecoHMEM: Improving Object Placement Methodology for Hybrid Memory Systems in HPCabstractRecent byte-addressable persistent memory (PMEM) technology offers capacities comparable to storage devices and access times much closer to DRAMs than other non-volatile memory technology. To palliate the large gap with DRAM performance, DRAM and PMEM are usually combined. Users have the choice to either manage the placement to different memory spaces by software or leverage the DRAM as a cache for the virtual address space of the PMEM. We present novel methodology for automatic object-level placement, including efficient runtime object matching and bandwidth-aware placement. Our experiments leveraging Intel® Optane™ Persistent Memory show from matching to greatly improved performance with respect to state-of-the-art software and hardware solutions, attaining over 2x runtime improvement in miniapplications and over 6% in OpenFOAM, a complex production application. Marc Jordà, Siddharth Rai, Eduard Ayguadé, Jesús Labarta, Antonio J. Peña |
CLUSTER | 4 |
| 2022 | Towards Reconfigurable Accelerators in HPC: Designing a Multipurpose eFPGA Tile for Heterogeneous SoCsabstractThe goal of modern high performance computing platforms is to combine low power consumption and high throughput. Within the European Processor Initiative (EPI), such an SoC platform to meet the novel exascale requirements is built and investigated. As part of this project, we introduce an embedded Field Programmable Gate Array (eFPGA), adding flexibility to accelerate various workloads. In this article, we show our approach to design the eFPGA tile that supports the EPI SoC. While eFPGAs are inherently reconfigurable, their initial design has to be determined for tape-out. The design space of the eFPGA is explored and evaluated with different configurations of two HPC workloads, covering control and dataflow heavy applications. As a result, we present a well-balanced eFPGA design that can host several use cases and potential future ones by allocating 1% of the total EPI SoC area. Finally, our simulation results of the architectures on the eFPGA show great performance improvements over their software counterparts. Tim Hotfilter, Fabian Kreß, Fabian Kempf, Jürgen Becker 0001, Juan Miguel De Haro Ruiz, Daniel Jiménez-González, Miquel Moretó, Carlos Álvarez 0001, Jesús Labarta, Imen Baili |
DATE | 9 |
| 2022 | OmpSs-2@Cluster: Distributed Memory Execution of Nested OpenMP-style Tasks
Jimmy Aguilar Mena, Omar Shaaban, Vicenç Beltran 0001, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
Euro-Par | 6 |
| 2022 | Transparent load balancing of MPI programs using [email protected] and DLBabstractAbstract Load imbalance is a long-standing source of inefficiency in high performance computing. The situation has only got worse as applications and systems increase in complexity, e.g., adaptive mesh refinement, DVFS, memory hierarchies, power and thermal management, and manufacturing processes. Load balancing is often implemented in the application, but it obscures application logic and may need extensive code refactoring. This paper presents an automated and transparent dynamic load balancing approach for MPI applications with OmpSs-2 tasks, which relieves applications from this burden. Only local and trivial changes are required to the application. Our approach exploits the ability of [email protected] to offload tasks for execution on other nodes, and it reallocates compute resources among ranks using the Dynamic Load Balancing (DLB) library. It employs LeWI to react to fine-grained load imbalances and DROM to address coarse-grained load imbalances by reserving cores on other nodes that can be reclaimed on demand. We use an expander graph to limit the amount of point-to-point communication and state. The results show 46% reduction in time-to-solution for micro-scale solid mechanics on 32 nodes and a 20% reduction beyond DLB for n-body on 16 nodes, when one node is running slow. A synthetic benchmark shows that performance is within 10% of optimal for an imbalance of up to 2.0 on 8 nodes. All software is released open source. Jimmy Aguilar Mena, Omar Shaaban, Víctor López 0003, Marta Garcia-Gasulla, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
ICPP | 7 |
| 2022 | Feature Space Curvature Map: A Method To Homogenize Cluster DensitiesabstractThe majority of density-based clustering algorithms can not perform properly when data expose very different density through the feature space. These algorithms implicitly presume that all clusters almost have the same density, therefore, they normally use global parameters. Consequently, they are often biased towards finding dense clusters in front of sparse ones. In this paper, we propose a parametric multilinear transformation method to homogenize cluster densities while preserving the topological structure of the dataset. The transformed clusters have approximately the same density while all inter-cluster regions become globally low-density. In our method, the feature space is locally bent by dense data point concentrations the same way as stars bend the space-time dimensions in Theory of Relativity. We present a new Gravitational Self-organization Map to model the feature space curvature by plugging the concepts of gravity and fabric of space into the Self-organization Map algorithm to mathematically describe the density structure of the data. To homogenize the cluster density, we introduce a novel mapping mechanism to project the data from a non-Euclidean curved space to a new Euclidean flat space. Specifically, this mechanism transfers the basis vectors instead of the feature vectors to guarantee the continuity of the mapping function and optimize the computation cost of the algorithm. As a result, our method can efficiently and explicitly homogenize the density of any dataset globally to then apply existing clustering algorithms without modification. Our experimental results over both real-world and synthetic datasets show that our approach outperforms the current statistical-based methods. Kaveh Mahdavi, Jesús Labarta, Judit Giménez, Atefeh Mousavinia, Atiyeh Mousavinia |
IJCNN | 2 |
| 2022 | OmpSs@cloudFPGA: An FPGA Task-Based Programming Model with Message PassingabstractNowadays, a new parallel paradigm for energy-efficient heterogeneous hardware infrastructures is required to achieve better performance at a reasonable cost on high-performance computing applications. Under this new paradigm, some application parts are offloaded to specialized accelerators that run faster or are more energy-efficient than CPUs. Field-Programmable Gate Arrays (FPGA) are one of those types of accelerators that are becoming widely available in data centers. This paper proposes OmpSs@cloudFPGA, which includes novel extensions to parallel task-based programming models that enable easy and efficient programming of heterogeneous clusters with FPGAs. The programmer only needs to annotate, with OpenMP-like pragmas, the tasks of the application that should be accelerated in the cluster of FPGAs. Next, the proposed programming model framework automatically extracts parts annotated with High-Level Synthesis (HLS) pragmas and synthesizes them into hardware accelerator cores for FPGAs. Additionally, our extensions include and support two novel features: 1) FPGA-to-FPGA direct communication using a Message Passing Interface (MPI) similar Application Programming Interface (API) with one-to-one and collective communications to alleviate host communication channel bottleneck, and 2) creating and spawning work from inside the FPGAs to their own accelerator cores based on an MPI rank-like identification. These features break the classical host-accelerator model, where the host (typically the CPU) generates all the work and distributes it to each accelerator. We also present an evaluation of OmpSs@cloudFPGA for different parallel strategies of the N-Body application on the IBM cloudFPGA research platform. Results show that for cluster sizes up to 56 FPGAs, the performance scales linearly. To the best of our knowledge, this is the best performance obtained for N-body over FPGA platforms, reaching 344 Gpairs/s with 56 FPGAs. Finally, we compare the performance and power consumption of the proposed approach with the ones obtained by a classical execution on the MareNostrum 4 supercomputer, demonstrating that our FPGA approach reduces power consumption by an order of magnitude. Juan Miguel De Haro Ruiz, Rubén Cano, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, François Abel, Burkhard Ringlein, Beat Weiss |
IPDPS | 7 |
| 2022 | Automatic aggregation of subtask accesses for nested OpenMP-style tasksabstractTask-based programming is a high performance and productive model to express parallelism. Tasks encapsulate work to be executed across multiple cores or offloaded to GPUs, FPGAs, other accelerators or other nodes. In order to maintain parallelism and afford maximum freedom to the scheduler, the task dependency graph should be created in parallel and well in advance of task execution. A key limitation with OpenMP and OmpSs-2 tasking is that a task cannot be created until all its accesses and its descendents' accesses are known. Current approaches to work around this limitation either stop task creation and execution using a taskwait or they substitute “fake” accesses known as sentinels. This paper proposes the auto clause, which indicates that the task may create subtasks that access unspecified memory regions or it may allocate and return memory at addresses that are of course not yet known. Unlike approaches using taskwaits, there is no interruption to the concurrent creation and execution of tasks, maintaining parallelism and the scheduler's ability to optimize load balance and data locality. Unlike existing approaches using sentinels, all tasks can be given a precise specification of their own data accesses, so that a single mechanism is used to control task ordering, program data transfers on distributed memory and optimize data locality, e.g. on NUMA systems. The auto clause also provides an incremental path to develop programs with nested tasks, by removing the need for every parent task to have a complete specification of the accesses of its descendent tasks. This is redundant information that can be time consuming and error-prone to describe. We present a straightforward runtime implementation that achieves a 1.4 times speedup for n-body with OmpSs-2@Cluster task offloading to 32 nodes and <4% slowdown for three benchmarks with task offloading to 8 nodes. All code is open source. Omar Shaaban, Jimmy Aguilar Mena, Vicenç Beltran 0001, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta |
SBAC-PAD | 6 |
| 2022 | The MAMe dataset: on the relevance of high resolution and variable shape image properties
Ferran Parés, Anna Arias-Duart, Dario Garcia-Gasulla, Gema Campo-Francés, Nina Viladrich, Eduard Ayguadé, Jesús Labarta |
Appl. Intell. | 7 |
| 2022 | Asymmetric HMMs for Online Ball-Bearing Health AssessmentsabstractThe degradation of critical components inside large industrial assets, such as ball-bearings, has a negative impact on production facilities, reducing the availability of assets due to an unexpectedly high failure rate. Machine learning-based monitoring systems can estimate the remaining useful life (RUL) of ball bearings, reducing the downtime by early failure detection. However, traditional approaches for predictive systems require run-to-failure (RTF) data as training data, which in real scenarios can be scarce and expensive to obtain as the expected useful life could be measured in years. Therefore, to overcome the need of RTF, we propose a new methodology based on online novelty detection and asymmetrical hidden Markov models (As-HMMs) to work out the health assessment. This new methodology does not require previous RTF data and can adapt to natural degradation of mechanical components over time in data-stream and online environments. As the system is designed to work online within the electrical cabinet of machines, it has to be deployed using embedded electronics. Therefore, a performance analysis of As-HMM is presented to detect the strengths and critical points of the algorithm. To validate our approach, we use real life ball-bearing data sets and compare our methodology with other methodologies where no RTF data are needed and check the advantages in RUL prediction and health monitoring. As a result, we showcase a complete end-to-end solution from the sensor to actionable insights regarding RUL estimation toward maintenance application in real industrial environments. Carlos Puerto-Santana, Concha Bielza, Javier Diaz-Rozo, Guillem Ramirez-Gargallo, Filippo Mantovani, Gaizka Virumbrales, Jesús Labarta, Pedro Larrañaga |
IEEE Internet Things J. | 7 |
| 2021 | Organization Component Analysis: The method for extracting insights from the shape of clusterabstractClustering analysis is widely used to stratify data in the same cluster when they are similar according to specific metrics. The process of understanding and interpreting clusters is mostly intuitive. However, we observe each cluster has unique shape that comes out of metrics on data, which can represent the organization of categorized data mathematically. In this paper, we apply novel topological based method to study potentially complex high-dimensional categorized data by quantifying their shapes and extracting fine-grain insights about them to interpret the clustering result. We introduce our Organization Component Analysis method for the purpose of the automatic arbitrary cluster-shape study without assumption about the data distribution. Our method explores a topology-preserving map of a data cluster manifold to extract the main organization structure of a cluster by the leveraging of the self-organization map technique. To do this, we represent self-organization map as graph. We introduce organization components to geometrically describe the shape of cluster and their endogenous phenomena. Specifically, we propose an innovative way to measure the alignment between two sequences of momentum changes on geodesic path over the embedded graph to quantify the extent to which the feature is related to a given component. As a result, we can describe variability among stratified data, correlated features in terms of lower number of organization components. We illustrate the utilization of our method by applying it to two quite different types of data, in each case mathematically detecting the organization structure of categorized data which are much profounder and finer than those produced by standard methods. Kaveh Mahdavi, Jesús Labarta, Judit Giménez |
IJCNN | 2 |
| 2021 | OmpSs@FPGA Framework for High Performance FPGA ComputingabstractThis article presents the new features of the OmpSs@FPGA framework. OmpSs is a data-flow programming model that supports task nesting and dependencies to target asynchronous parallelism and heterogeneity. OmpSs@FPGA is the extension of the programming model addressed specifically to FPGAs. OmpSs environment is built on top of Mercurium source to source compiler and Nanos++ runtime system. To address FPGA specifics Mercurium compiler implements several FPGA related features as local variable caching, wide memory accesses or accelerator replication. In addition, part of the Nanos++ runtime has been ported to hardware. Driven by the compiler this new hardware runtime adds new features to FPGA codes, such as task creation and dependence management, providing both performance increases and ease of programming. To demonstrate these new capabilities, different high performance benchmarks have been evaluated over different FPGA platforms using the OmpSs programming model. The results demonstrate that programs that use the OmpSs programming model achieve very competitive performance with low to moderate porting effort compared to other FPGA implementations. Juan Miguel De Haro Ruiz, Jaume Bosch, Antonio Filgueras, Miquel Vidal, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
IEEE Trans. Computers | 9 |
| 2020 | sLASs: A fully automatic auto-tuned linear algebra library based on OpenMP extensions implemented in OmpSs (LASs Library)
Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Tetsuzo Usui, Jesús Labarta |
J. Parallel Distributed Comput. | 5 |
| 2019 | Unsupervised Feature Selection for Noisy Data
Kaveh Mahdavi, Jesús Labarta, Judit Giménez |
ADMA | 2 |
| 2019 | The OTree: Multidimensional Indexing with efficient data Sampling for HPCabstractSpatial big data is considered an essential trend in future scientific and business applications. Indeed, research instruments, medical devices, and social networks generate hundreds of petabytes of spatial data per year. However, many authors have pointed out that the lack of specialized frameworks for multidimensional Big Data is limiting possible applications and precluding many scientific breakthroughs. Paramount in achieving High-Performance Data Analytics is to optimize and reduce the I/O operations required to analyze large data sets. To do so, we need to organize and index the data according to its multidimensional attributes. At the same time, to enable fast and interactive exploratory analysis, it is vital to generate approximate representations of large datasets efficiently. In this paper, we propose the Outlook Tree (or OTree), a novel Multidimensional Indexing with efficient data Sampling (MIS) algorithm. The OTree enables exploratory analysis of large multidimensional datasets with arbitrary precision, a vital missing feature in current distributed data management solutions. Our algorithm reduces the indexing overhead and achieves high performance even for write-intensive HPC applications. Indeed, we use the OTree to store the scientific results of a study on the efficiency of drug inhalers. Then we compare the OTree implementation on Apache Cassandra, named Qbeast, with PostgreSQL and plain storage. Lastly, we demonstrate that our proposal delivers better performance and scalability. Cesare Cugnasco, Hadrien Calmet, Pol Santamaria, Raül Sirvent, Beatriz Eguzkitza, Guillaume Houzeaux, Yolanda Becerra 0001, Jordi Torres, Jesús Labarta |
IEEE BigData | 9 |
| 2019 | Accelerating Conjugate Gradient using OmpSsabstractIn this paper, we present the benefits of using the clause concurrent of OmpSs when performing reductions, more specifically, when applied to the dot product (DOT) operations. We analyze its benefits through the implementation of different versions of the Conjugate Gradient (CG) method. We start from a parallel version of the code based on tasks and dependencies; later, we introduce the use of the concurrent clause, which allows to overlap the execution of tasks that have data dependencies among them. In this way, we want to show the benefits of the concurrent clause, which might be included in OpenMP standard as previously done with other OmpSs features. Our tests, performed on a single node of the (Intel-based) Marenostrum 4 Supercomputer and a single socket of the (ARM-based) Dibona cluster, show that the use of the concurrent clause may improve performance with respect to the version where only tasks and dependencies are used around 37% and 23% respectively. Sandra Catalán, Xavier Martorell, Jesús Labarta, Tetsuzo Usui, Leonel Toledo, Pedro Valero-Lara |
PDCAT | 3 |
| 2019 | BLAS-3 Optimized by OmpSs Regions (LASs Library)abstractIn this paper we propose a set of optimizations for the BLAS-3 routines of LASs library (Linear Algebra routines on OmpSs) and perform a detailed analysis of the impact of the proposed changes in terms of performance and execution time. OmpSs allows to use regions in the dependences of the tasks. This helps not only in the programming of the algorithmic optimizations, but also in the reduction of the execution time achieved by such optimizations. Different strategies are implemented in order to reduce the amount of tasks created (when there is enough parallelism) during the execution of BLAS-3 operations in the original LASs. Also a better IPC is obtained thanks to a better memory hierarchy exploitation. More specifically, we increase the performance, in particular on big matrices, about 12% for TRSM, and 17% for GEMM with respect to the original version of LASs, even using less cores in the case of GEMM/SYMM. Moreover, when LASs is compared to the OpenMP reference dense linear algebra library PLASMA, performance is increased up to 12.5% for GEMM/SYMM, while for TRSM/TRMM this value raises to 15%. Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Jesús Labarta |
PDP | 4 |
| 2019 | Integrating blocking and non-blocking MPI primitives with task-based programming models
Kevin Sala, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Vicenç Beltran 0001, Jesús Labarta |
Parallel Comput. | 6 |
| 2019 | MPI+OpenMP tasking scalability for multi-morphology simulations of the human brain
Pedro Valero-Lara, Raül Sirvent, Antonio J. Peña, Jesús Labarta |
Parallel Comput. | 4 |
| 2018 | Application Acceleration on FPGAs with OmpSs@FPGAabstractOmpSs@FPGA is the flavor of OmpSs that allows offloading application functionality to FPGAs. Similarly to OpenMP, it is based on compiler directives. While the OpenMP specification also includes support for heterogeneous execution, we use OmpSs and OmpSs@FPGA as prototype implementation to develop new ideas for OpenMP. OmpSs@FPGA implements the tasking model with runtime support to automatically exploit all SMP and FPGA resources available in the execution platform. In this paper, we present the OmpSs@FPGA ecosystem, based on the Mercurium compiler and the Nanos++ runtime system. We show how the applications are transformed to run on the SMP cores and the FPGA. The application kernels defined as tasks to be accelerated, using the OmpSs directives are: 1) transformed by the compiler into kernels connected with the proper synchronization and communication ports, 2) extracted to intermediate files, 3) compiled through the FPGA vendor HLS tool, and 4) used to configure the FPGA. Our Nanos++ runtime system schedules the application tasks on the platform, being able to use the SMP cores and the FPGA accelerators at the same time. We present the evaluation of the OmpSs@FPGA environment with the Matrix Multiplication, Cholesky and N-Body benchmarks, showing the internal details of the execution, and the performance obtained on a Zynq Ultrascale+ MPSoC (up to 128x). The source code uses OmpSs@FPGA annotations and different Vivado HLS optimization directives are applied for acceleration. Jaume Bosch, Xubin Tan, Antonio Filgueras, Miquel Vidal, Marc Mateu, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
FPT | 10 |
| 2018 | Runtime-Guided Management of Stacked DRAM Memories in Task Parallel ProgramsabstractStacked DRAM memories have become a reality in High-Performance Computing (HPC) architectures. These memories provide much higher bandwidth while consuming less power than traditional off-chip memories, but their limited memory capacity is insufficient for modern HPC systems. For this reason, both stacked DRAM and off-chip memories are expected to co-exist in HPC architectures, giving raise to different approaches for architecting the stacked DRAM in the system. Lluc Alvarez, Marc Casas, Jesús Labarta, Eduard Ayguadé, Mateo Valero, Miquel Moretó |
ICS | 3 |
| 2018 | Reducing Data Movement on Large Shared Memory Systems by Exploiting Computation DependenciesabstractShared memory systems are becoming increasingly complex as they typically integrate several storage devices. That brings different access latencies or bandwidth rates depending on the proximity between the cores where memory accesses are issued and the storage devices containing the requested data. In this context, techniques to manage and mitigate non-uniform memory access (NUMA) effects consist in migrating threads, memory pages or both and are generally applied by the system software. Isaac Sánchez Barrera, Miquel Moretó, Eduard Ayguadé, Jesús Labarta, Mateo Valero, Marc Casas |
ICS | 4 |
| 2018 | Variable Batched DGEMMabstractMany scientific applications are in need to solve a high number of small-size independent problems. These individual problems do not provide enough parallelism and then, these must be computed as a batch. Today, vendors such as Intel and NVIDIA are developing their own suite of batch routines. Although most of the works focus on computing batches of fixed size, in real applications we can not assume a uniform size for all set of problems. We explore and analyze different strategies based on parallel for, task and taskloop OpenMP pragmas. Although these strategies are straightforward from a programmer's point of view, they have a different impact on performance. We also analyze a new prototype provided by Intel (MKL), which deals with batch operations (cblas_dgemm_batch). We propose a new approach called grouping. It basically groups a set of problems until filling a limit in terms of memory occupancy or number of operations. In this way, groups composed by different number of problems are distributed on cores, achieving a more balanced distribution in terms of computational cost. This strategy is able to be up to 6× faster than the Intel (MKL) batch routine. Pedro Valero-Lara, Ivan Martínez-Pérez, Sergi Mateo, Raül Sirvent, Vicenç Beltran 0001, Xavier Martorell, Jesús Labarta |
PDP | 7 |
| 2018 | Graph partitioning applied to DAG scheduling to reduce NUMA effectsabstractThe complexity of shared memory systems is becoming more relevant as the number of memory domains increases, with different access latencies and bandwidth rates depending on the proximity between the cores and the devices containing the data. In this context, techniques to manage and mitigate non-uniform memory access (NUMA) effects consist in migrating threads, memory pages or both and are typically applied by the system software. Isaac Sánchez Barrera, Marc Casas, Miquel Moretó, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
PPoPP | 5 |
| 2018 | Improving the Interoperability between MPI and Task-Based Programming ModelsabstractIn this paper we propose an API to pause and resume task execution depending on external events. We leverage this generic API to improve the interoperability between MPI synchronous communication primitives and tasks. When an MPI operation blocks, the task running is paused so that the runtime system can schedule a new task on the core that became idle. Once the MPI operation is completed, the paused task is put again on the runtime system's ready queue. We expose our proposal through a new MPI threading level which we implement through two approaches. Kevin Sala, Jorge Bellón, Pau Farré, Xavier Teruel, Josep M. Pérez, Antonio J. Peña, Daniel J. Holmes, Vicenç Beltran 0001, Jesús Labarta |
EuroMPI | 9 |
| 2018 | MPI+OpenMP Tasking Scalability for the Simulation of the Human Brain: Human Brain ProjectabstractThe simulation of the behavior of the Human Brain is one of the most ambitious challenges today with a non-end of important applications. We can find many different initiatives in the USA, Europe and Japan which attempt to achieve such a challenging target. In this work we focus on the most important European initiative (Human Brain Project) and on one of the tools (Arbor). This tool simulates the spikes triggered in a neuronal network by computing the voltage capacitance on the neurons' morphology, being one of the most precise simulators today. In the present work, we have evaluated the use of MPI+OpenMP tasking on top of the Arbor simulator. In this paper, we present the main characteristics of the Arbor tool and how these can be efficiently managed by using MPI+OpenMP tasking. We prove that this approach is able to achieve a good scaling even when computing a relatively low workload (number of neurons) per node using up to 32 nodes. Our target consists of achieving not only a highly scalable implementation based on MPI, but also to develop a tool with a high degree of abstraction without losing control and performance by using MPI+OpenMP tasking. Pedro Valero-Lara, Raül Sirvent, Antonio J. Peña, Xavier Martorell, Jesús Labarta |
EuroMPI | 5 |
| 2018 | On the Behavior of Convolutional Nets for Feature ExtractionabstractDeep neural networks are representation learning techniques. During training, a deep net is capable of generating a descriptive language of unprecedented size and detail in machine learning. Extracting the descriptive language coded within a trained CNN model (in the case of image data), and reusing it for other purposes is a field of interest, as it provides access to the visual descriptors previously learnt by the CNN after processing millions of images, without requiring an expensive training phase. Contributions to this field (commonly known as feature representation transfer or transfer learning) have been purely empirical so far, extracting all CNN features from a single layer close to the output and testing their performance by feeding them to a classifier. This approach has provided consistent results, although its relevance is limited to classification tasks. In a completely different approach, in this paper we statistically measure the discriminative power of every single feature found within a deep CNN, when used for characterizing every class of 11 datasets. We seek to provide new insights into the behavior of CNN features, particularly the ones from convolutional layers, as this can be relevant for their application to knowledge representation and reasoning. Our results confirm that low and middle level features may behave differently to high level features, but only under certain conditions. We find that all CNN features can be used for knowledge representation purposes both by their presence or by their absence, doubling the information a single CNN feature may provide. We also study how much noise these features may include, and propose a thresholding approach to discard most of it. All these insights have a direct application to the generation of CNN embedding spaces. Dario Garcia-Gasulla, Ferran Parés, Armand Vilalta, Jonathan Moreno, Eduard Ayguadé, Jesús Labarta, Ulises Cortés, Toyotaro Suzumura |
J. Artif. Intell. Res. | 6 |
| 2018 | Understanding memory access patterns using the BSC performance tools
Harald Servat, Jesús Labarta, Hans-Christian Hoppe, Judit Giménez, Antonio J. Peña |
Parallel Comput. | 2 |
| 2018 | Asynchronous and Exact Forward Recovery for Detected Errors in Iterative SolversabstractCurrent trends and projections show that faults in computer systems become increasingly common. Such errors may be detected, and possibly corrected transparently, e.g., by Error Correcting Codes (ECC). For a program to be fault-tolerant, it needs to also handle the Errors that are Detected and Uncorrected (DUE), such as an ECC encountering too many bit flips in a codeword. While correcting an error has an overhead in itself, it can also affect the progress of a program. The most generic technique, rolling back the program state to a previously taken checkpoint, sets back any progress done since then. Alternately, application specific techniques exist, such as restarting an iterative program with its latest iteration's values as initial guess. We introduce a novel error correction technique for iterative linear solvers, designed to preserve both the progress made and the solver's future convergence by recovering the program's state exactly. Leveraging the asynchrony of task-based programming models, we mask our technique's overhead by overlapping error correction with the solver's normal workload. Our technique relies on analysing solvers to find redundancy in the form of relations between data. We are then able to restore discarded or corrupted data by recomputing or inverting the appropriate relations. We demonstrate that this approach allows to recover any part of three widely used Krylov subspace methods: CG, GMRES and BiCGStab, and their pre-conditioned versions. We implement our technique for CG and recover lost data at the scale of a memory page, which is the granularity at which Operating Systems (OS) report memory errors on commodity hardware, and study the effect of varying the memory page size to address non-standard sizes and the possible use of huge pages in High Performance Computing (HPC). When compared to checkpointing and to the state-of-the-art algorithmic restart technique, on small (8 cores) to large scale (1024 cores), our methods show less overhead. A trade-off arises between our straightforward and asynchronous approaches, based on the rate at which faults happen. At the lowest considered rate and page size, overlapping recoveries decreases their average cost from 5.40 to 2.24 percent of the ideal faultless execution time. Our methods generally outperform the state-of-the-art even with increased overheads on big page sizes, and perform similarly on edge cases. These results also indicate that our techniques are increasingly efficient as the matrix size increases. Luc Jaulmes, Miquel Moretó, Eduard Ayguadé, Jesús Labarta, Mateo Valero, Marc Casas |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Designing and Modelling Selective Replication for Fault-tolerant HPC ApplicationsabstractFail-stop errors and Silent Data Corruptions (SDCs) are the most common failure modes for High Performance Computing (HPC) applications. There are studies that address fail-stop errors and studies that address SDCs. However few studies address both types of errors together. In this paper we propose a software-based selective replication technique for HPC applications for both fail-stop errors and SDCs. Since complete replication of applications can be costly in terms of resources, we develop a runtime-based technique for selective replication. Selective replication provides an opportunity to meet HPC reliability targets while decreasing resource costs. Our technique is low-overhead, automatic and completely transparent to the user. Omer Subasi, Gulay Yalcin, Ferad Zyulkyarov, Osman S. Unsal, Jesús Labarta |
CCGrid | 5 |
| 2017 | Automating the Application Data Placement in Hybrid Memory SystemsabstractMulti-tiered memory systems, such as those based on Intel®Xeon Phi™ processors, are equipped with several memory tiers with different characteristics including, among others, capacity, access latency, bandwidth, energy consumption, and volatility. The proper distribution of the application data objects into the available memory layers is key to shorten the time- to-solution, but the way developers and end-users determine the most appropriate memory tier to place the application data objects has not been properly addressed to date. In this paper we present a novel methodology to build an extensible framework to automatically identify and place the application's most relevant memory objects into the Intel Xeon Phi fast on-package memory. Our proposal works on top of inproduction binaries by first exploring the application behavior and then substituting the dynamic memory allocations. This makes this proposal valuable even for end-users who do not have the possibility of modifying the application source code. We demonstrate the value of a framework based in our methodology for several relevant HPC applications using different allocation strategies to help end-users improve performance with minimal intervention. The results of our evaluation reveal that our proposal is able to identify the key objects to be promoted into fast on-package memory in order to optimize performance, leading to even surpassing hardware-based solutions. Harald Servat, Antonio J. Peña, Germán Llort, Estanislao Mercadal, Hans-Christian Hoppe, Jesús Labarta |
CLUSTER | 6 |
| 2017 | MACORD: Online Adaptive Machine Learning Framework for Silent Error DetectionabstractFuture high-performance computing (HPC) systems with ever-increasing resource capacity (such as compute cores, memory and storage) may significantly increase the risks on reliability. Silent data corruptions (SDCs) or silent errors are among the major sources that corrupt HPC execution results. Unlike fail-stop errors, SDCs can be harmful and dangerous in that they cannot be detected by hardware. To remedy this, we propose an online MAchine-learning-based silent data CORruption Detection framework (abbreviated as MACORD) for detecting SDCs in HPC applications. In our study, we comprehensively investigate the prediction ability of a multitude of machine-learning algorithms and enable the detector to automatically select the best-fit algorithms at runtime to adapt to the data dynamics. Because it takes only spatial features (i.e., neighboring data values for each data point in the current time step) into the training data, our learning framework exhibits low memory overhead (less than 1%). Experiments based on real-world scientific applications/benchmarks show that our framework can elevate the detection sensitivity (i.e., recall) up to 99%. Meanwhile the false positive rate is limited to 0.1% in most cases, which is one order of magnitude improvement compared with the latest state-of-the-art spatial technique. Omer Subasi, Sheng Di, Prasanna Balaprakash, Osman S. Unsal, Jesús Labarta, Adrián Cristal, Sriram Krishnamoorthy, Franck Cappello |
CLUSTER | 5 |
| 2017 | Improving the Integration of Task Nesting and Dependencies in OpenMPabstractThe tasking model of OpenMP 4.0 supports both nesting and the definition of dependences between sibling tasks. A natural way to parallelize many codes with tasks is to first taskify the high-level functions and then to further refine these tasks with additional subtasks. However, this top-down approach has some drawbacks since combining nesting with dependencies usually requires additional measures to enforce the correct coordination of dependencies across nesting levels. For instance, most non-leaf tasks need to include a taskwait at the end of their code. While these measures enforce the correct order of execution, as a side effect, they also limit the discovery of parallelism. In this paper we extend the OpenMP tasking model to improve the integration of nesting and dependencies. Our proposal builds on both formulas, nesting and dependencies, and benefits from their individual strengths. On one hand, it encourages a top-down approach to parallelizing codes that also enables the parallel instantiation of tasks. On the other hand, it allows the runtime to control dependencies at a fine grain that until now was only possible using a single domain of dependencies. Our proposal is realized through additions to the OpenMP task directive that ensure backward compatibility with current codes. We have implemented a new runtime with these extensions and used it to evaluate the impact on several benchmarks. Our initial findings show that our extensions improve performance in three areas. First, they expose more parallelism. Second, they uncover dependencies across nesting levels, which allows the runtime to make better scheduling decisions. And third, they allow the parallel instantiation of tasks with dependencies between them. Josep M. Pérez, Vicenç Beltran 0001, Jesús Labarta, Eduard Ayguadé |
IPDPS | 3 |
| 2017 | Noise Inspector ToolabstractThe operating system noise can interfere with normal execution programs. This behavior is becoming especially important when scaling parallel programs and amplified with global synchronizations. This work presents a tool to detect in a non-intrusive way the alien programs that share resources with current running applications in a multicore cluster. Gladys Utrera, Jordi Fornes, Jesús Labarta |
PDP | 3 |
| 2017 | Task Scheduling Techniques for Asymmetric Multi-Core SystemsabstractAs performance and energy efficiency have become the main challenges for next-generation high-performance computing, asymmetric multi-core architectures can provide solutions to tackle these issues. Parallel programming models need to be able to suit the needs of such systems and keep on increasing the application’s portability and efficiency. This paper proposes two task scheduling approaches that target asymmetric systems. These dynamic scheduling policies reduce total execution time either by detecting the longest or the critical path of the dynamic task dependency graph of the application, or by finding the earliest executor of a task. They use dynamic scheduling and information discoverable during execution, fact that makes them implementable and functional without the need of off-line profiling. In our evaluation we compare these scheduling approaches with two existing state-of the art heterogeneous schedulers and we track their improvement over a FIFO baseline scheduler. We show that the heterogeneous schedulers improve the baseline by up to 1.45$\times$in a real 8-core asymmetric system and up to 2.1$\times$in a simulated 32-core asymmetric chip. Kallia Chronaki, Alejandro Rico, Marc Casas, Miquel Moretó, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2016 | Reducing Cache Coherence Traffic with Hierarchical Directory Cache and NUMA-Aware Runtime SchedulingabstractCache Coherent NUMA (ccNUMA) architectures are a widespread paradigm due to the benefits they provide for scaling core count and memory capacity. Also, the flat memory address space they offer considerably improves programmability. However, ccNUMA architectures require sophisticated and expensive cache coherence protocols to enforce correctness during parallel executions, which trigger a significant amount of on- and off-chip traffic in the system. Paul Caheny, Marc Casas, Miquel Moretó, Hervé Gloaguen, Maxime Saintes, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
PACT | 7 |
| 2016 | POSTER: Exploiting Asymmetric Multi-Core Processors with Flexible System SofwareabstractEnergy efficiency has become the main challenge for high performance computing (HPC). The use of mobile asymmetric multi-core architectures to build future multi-core systems is an approach towards energy savings while keeping high performance. However, it is not known yet whether such systems are ready to handle parallel applications. Kallia Chronaki, Miquel Moretó, Marc Casas, Alejandro Rico, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
PACT | 7 |
| 2016 | POSTER: Collective Dynamic Parallelism for Directive Based GPU Programming Languages and CompilersabstractEarly programs for GPU (Graphics Processing Units) acceleration were based on a flat, bulk parallel programming model, in which programs had to perform a sequence of kernel launches from the host CPU. In the latest releases of these devices, dynamic (or nested) parallelism is supported, making possible to launch kernels from threads running on the device, without host intervention. Unfortunately, the overhead of launching kernels from the device is higher compared to launching from the host CPU, making the exploitation of dynamic parallelism unprofitable. Guray Ozen, Eduard Ayguadé, Jesús Labarta |
PACT | 3 |
| 2016 | Spatial Support Vector Regression to Detect Silent Errors in the Exascale EraabstractAs the exascale era approaches, the increasing capacity of high-performance computing (HPC) systems with targeted power and energy budget goals introduces significant challenges in reliability. Silent data corruptions (SDCs) or silent errors are one of the major sources that corrupt the executionresults of HPC applications without being detected. In this work, we explore a low-memory-overhead SDC detector, by leveraging epsilon-insensitive support vector machine regression, to detect SDCs that occur in HPC applications that can be characterized by an impact error bound. The key contributions are three fold. (1) Our design takes spatialfeatures (i.e., neighbouring data values for each data point in a snapshot) into training data, such that little memory overhead (less than 1%) is introduced. (2) We provide an in-depth study on the detection ability and performance with different parameters, and we optimize the detection range carefully. (3) Experiments with eight real-world HPC applications show thatour detector can achieve the detection sensitivity (i.e., recall) up to 99% yet suffer a less than 1% of false positive rate for most cases. Our detector incurs low performance overhead, 5% on average, for all benchmarks studied in the paper. Compared with other state-of-the-art techniques, our detector exhibits the best tradeoff considering the detection ability and overheads. Omer Subasi, Sheng Di, Leonardo Arturo Bautista-Gomez, Prasanna Balaprakash, Osman S. Unsal, Jesús Labarta, Adrián Cristal, Franck Cappello |
CCGrid | 6 |
| 2016 | A Runtime Heuristic to Selectively Replicate Tasks for Application-Specific Reliability TargetsabstractIn this paper we propose a runtime-based selective task replication technique for task-parallel high performance computing applications. Our selective task replication technique is automatic and does not require modification/recompilation of OS, compiler or application code. Our heuristic, we call App_FIT, selects tasks to replicate such that the specified reliability target for an application is achieved. In our experimental evaluation, we show that App FIT selective replication heuristic is low-overhead and highly scalable. In addition, results indicate that complete task replication is overkill for achieving reliability targets. We show that with App FIT, we can tolerate pessimistic exascale error rates with only 53% of the tasks being replicated. Omer Subasi, Gulay Yalcin, Ferad Zyulkyarov, Osman S. Unsal, Jesús Labarta |
CLUSTER | 5 |
| 2016 | Runtime-Guided Mitigation of Manufacturing Variability in Power-Constrained Multi-Socket NUMA NodesabstractCurrent large scale systems show increasing power demands, to the point that it has become a huge strain on facilities and budgets. Researchers in academia, labs and industry are focusing on dealing with this "power wall", striving to find a balance between performance and power consumption. Some commodity processors enable power capping, which opens up new opportunities for applications to directly manage their power behavior at user level. However, while power capping ensures a system will never exceed a given power limit, it also leads to a new form of heterogeneity: natural manufacturing variability, which was previously hidden by varying power to achieve homogeneous performance, now results in heterogeneous performance caused by different CPU frequencies, potentially for each core, to enforce the power limit. Dimitrios Chasapis, Marc Casas, Miquel Moretó, Martin Schulz 0001, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
ICS | 6 |
| 2016 | CATA: Criticality Aware Task Acceleration for Multicore ProcessorsabstractManaging criticality in task-based programming models opens a wide range of performance and power optimization opportunities in future manycore systems. Criticality aware task schedulers can benefit from these opportunities by scheduling tasks to the most appropriate cores. However, these schedulers may suffer from priority inversion and static binding problems that limit their expected improvements. Based on the observation that task criticality information can be exploited to drive hardware reconfigurations, we propose a Criticality Aware Task Acceleration (CATA) mechanism that dynamically adapts the computational power of a task depending on its criticality. As a result, CATA achieves significant improvements over a baseline static scheduler, reaching average improvements up to 18.4% in execution time and 30.1% in Energy-Delay Product (EDP) on a simulated 32-core system. The cost of reconfiguring hardware by means of a software-only solution rises with the number of cores due to lock contention and reconfiguration overhead. Therefore, novel architectural support is proposed to eliminate these overheads on future manycore systems. This architectural support minimally extends hardware structures already present in current processors, which allows further improvements in performance with negligible overhead. As a consequence, average improvements of up to 20.4% in execution time and 34.0% in EDP are obtained, outperforming state-of-the-art acceleration proposals not aware of task criticality. Emilio Castillo, Miquel Moretó, Marc Casas, Lluc Alvarez, Enrique Vallejo 0001, Kallia Chronaki, Rosa M. Badia, José Luis Bosque, Ramón Beivide, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
IPDPS | 11 |
| 2016 | CRC-Based Memory Reliability for Task-Parallel HPC ApplicationsabstractMemory reliability will be one of the major concerns for future HPC and Exascale systems. This concern is mostly attributed to the expected massive increase in memory capacity and the number of memory devices in Exascale systems. For memory systems Error Correcting Codes (ECC) are the most commonly used mechanism. However state-of-the art hardware ECCs will not be sufficient in terms of error coverage for future computing systems and stronger hardware ECCs providing more coverage have prohibitive costs in terms of area, power and latency. Software-based solutions are needed to cooperate with hardware. In this work, we propose a Cyclic Redundancy Checks (CRCs) based software mechanism for task-parallel HPC applications. Our mechanism incurs only 1.7% performance overhead with hardware acceleration while being highly scalable at large scale. Our mathematical analysis demonstrates the effectiveness of our scheme and its error coverage. Results show that our CRC-based mechanism reduces the memory vulnerability by 87% on average with up to 32-bit burst (consecutive) and 5-bit arbitrary error correction capability. Omer Subasi, Osman S. Unsal, Jesús Labarta, Gulay Yalcin, Adrián Cristal |
IPDPS | 3 |
| 2016 | Bio-Inspired Call-Stack Reconstruction for Performance AnalysisabstractThe correlation of performance bottlenecks and their associated source code has become a cornerstone of performance analysis. It allows understanding why the efficiency of an application falls behind the computer's peak performance and enabling optimizations on the code ultimately. To this end, performance analysis tools collect the processor call-stack and then combine this information with measurements to allow the analyst comprehend the application behavior. Some tools modify the call-stack during run-time to diminish the collection expense but at the cost of resulting in non-portable solutions. In this paper, we present a novel portable approach to associate performance issues with their source code counterpart. To address it, we capture a reduced segment of the call-stack (up to three levels) and then process the segments using an algorithm inspired by multi-sequence alignment techniques. The results of our approach are easily mapped to detailed performance views, enabling the analyst to unveil the application behavior and its corresponding region of code. To demonstrate the usefulness of our approach, we have applied the algorithm to several first-time seen in-production applications to describe them finely, and optimize them by using tiny modifications based on the analyses. Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
PDP | 5 |
| 2016 | MUSA: a multi-level simulation approach for next-generation HPC machinesabstractThe complexity of High Performance Computing (HPC) systems is increasing in the number of components and their heterogeneity. Interactions between software and hardware involve many different aspects which are typically not transparent to scientific programmers and system architects. Therefore, predicting the behavior of current scientific applications on future HPC infrastructures is a challenging task. In this paper we present MUSA, an end-to-end methodology that employs a multi-level simulation infrastructure. By combining different levels of abstraction, MUSA is able to model the communication network, microarchitectural details and system software interactions, providing different trade-offs in terms of simulation cost and accuracy. We compare detailed MUSA simulations with native executions of up to 2,048 cores and find relative errors that are within 10% in the common case. In addition, we use MUSA to simulate up to 16,384 cores and successfully identify scalability bottlenecks due to different factors, e.g. memory contention or load imbalance. We also compare different system configurations, showing how MUSA can help system designers to assess the usefulness of future technologies in next-generation HPC machines. Thomas Grass, César Allande, Adrià Armejach, Alejandro Rico, Eduard Ayguadé, Jesús Labarta, Mateo Valero, Marc Casas, Miquel Moretó |
SC | 6 |
| 2016 | The mont-blanc prototype: an alternative approach for HPC systemsabstractHigh-performance computing (HPC) is recognized as one of the pillars for further progress in science, industry, medicine, and education. Current HPC systems are being developed to overcome emerging architectural challenges in order to reach Exascale level of performance, projected for the year 2020. The much larger embedded and mobile market allows for rapid development of intellectual property (IP) blocks and provides more flexibility in designing an application-specific system-on-chip (SoC), in turn providing the possibility in balancing performance, energy-efficiency, and cost. In the Mont-Blanc project, we advocate for HPC systems being built from such commodity IP blocks, currently used in embedded and mobile SoCs. As a first demonstrator of such an approach, we present the Mont-Blanc prototype; the first HPC system built with commodity SoCs, memories, and network interface cards (NICs) from the embedded and mobile domain, and off-the-shelf HPC networking, storage, cooling, and integration solutions. We present the system's architecture and evaluate both performance and energy efficiency. Further, we compare the system's abilities against a production level supercomputer. At the end, we discuss parallel scalability and estimate the maximum scalability point of this approach across a set of applications. Nikola Rajovic, Alejandro Rico, Filippo Mantovani, Daniel Ruiz 0003, Josep Oriol Vilarrubi, Constantino Gómez, Luna Backes, Diego Nieto, Harald Servat, Xavier Martorell, Jesús Labarta, Eduard Ayguadé, Chris Adeniyi-Jones, Said Derradji, Hervé Gloaguen, Piero Lanucara, Nico Sanna, Jean-François Méhaut, Kevin Pouget, Brice Videau, Eric Boyer, Momme Allalen, Axel Auweter, David Brayford, Daniele Tafani, Volker Weinberg, Dirk Brömmel, René Halver, Jan H. Meinke, Ramón Beivide, Mariano Benito, Enrique Vallejo 0001, Mateo Valero, Alex Ramírez |
SC | 11 |
| 2016 | Detailed and simultaneous power and performance analysisabstractSummary On the road to Exascale computing, both performance and power areas are meant to be tackled at different levels, from system to processor level. The processor itself is the main responsible for the serial node performance and also for the most of the energy consumed by the system. Thus, it is important to have tools to simultaneously analyze both performance and energy efficiency at processor level. Performance tools have allowed analysts to understand, and even improve, the performance of an application that runs in a system. With the advent of recent processor capabilities to measure its own power consumption, performance tools can increase their collection of metrics by adding those related to energy consumption and provide a correlation between the source code, its performance and its energy efficiency. In this paper, we present a performance tool that has been extended to gather such energy metrics. The results of this tool are passed to a mechanism called folding that produces detailed metrics and source code references by using coarse grain sampling. We have used the tool with multiple serial benchmarks as well as parallel applications to demonstrate its usefulness by locating hot spots in terms of performance and power drained. Copyright © 2013 John Wiley & Sons, Ltd. Harald Servat, Germán Llort, Judit Giménez, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | PARSECSs: Evaluating the Impact of Task Parallelism in the PARSEC Benchmark SuiteabstractIn this work, we show how parallel applications can be implemented efficiently using task parallelism. We also evaluate the benefits of such parallel paradigm with respect to other approaches. We use the PARSEC benchmark suite as our test bed, which includes applications representative of a wide range of domains from HPC to desktop and server applications. We adopt different parallelization techniques, tailored to the needs of each application, to fully exploit the task-based model. Our evaluation shows that task parallelism achieves better performance than thread-based parallelization models, such as Pthreads. Our experimental results show that we can obtain scalability improvements up to 42% on a 16-core system and code size reductions up to 81%. Such reductions are achieved by removing from the source code application specific schedulers or thread pooling systems and transferring these responsibilities to the runtime system software. Dimitrios Chasapis, Marc Casas, Miquel Moretó, Raul Vidal, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
ACM Trans. Archit. Code Optim. | 6 |
| 2015 | Runtime-Guided Management of Scratchpad Memories in Multicore ArchitecturesabstractThe increasing number of cores and the anticipated level of heterogeneity in upcoming multicore architectures cause important problems in traditional cache hierarchies. A good way to alleviate these problems is to add scratchpad memories alongside the cache hierarchy, forming a hybrid memory hierarchy. This memory organization has the potential to improve performance and to reduce the power consumption and the on-chip network traffic, but exposing such a complex memory model to the programmer has a very negative impact on the programmability of the architecture. Emerging task-based programming models are a promising alternative to program heterogeneous multicore architectures. In these models the runtime system manages the execution of the tasks on the architecture, allowing them to apply many optimizations in a generic way at the runtime system level. This paper proposes giving the runtime system the responsibility to manage the scratchpad memories of a hybrid memory hierarchy in multicore processors, transparently to the programmer. In the envisioned system, the runtime system takes advantage of the information found in the task dependences to map the inputs and outputs of a task to the scratchpad memory of the core that is going to execute it. In addition, the paper exploits two mechanisms to overlap the data transfers with computation and a locality-aware scheduler to reduce the data motion. In a 32-core multicore architecture, the hybrid memory hierarchy outperforms cache-only hierarchies by up to 16%, reduces on-chip network traffic by up to 31% and saves up to 22% of the consumed power. Lluc Alvarez, Miquel Moretó, Marc Casas, Emilio Castillo, Xavier Martorell, Jesús Labarta, Eduard Ayguadé, Mateo Valero |
PACT | 6 |
| 2015 | Spark deployment and performance evaluation on the MareNostrum supercomputerabstractIn this paper we present a framework to enable data-intensive Spark workloads on MareNostrum, a petascale supercomputer designed mainly for compute-intensive applications. As far as we know, this is the first attempt to investigate optimized deployment configurations of Spark on a petascale HPC setup. We detail the design of the framework and present some benchmark data to provide insights into the scalabilityof the system. We examine the impact of different configurations including parallelism, storage and networking alternatives, and we discuss several aspects in executing Big Data workloads on a computing system that is based on the compute-centric paradigm. Further, we derive conclusions aiming to pave the way towards systematic and optimized methodologies for fine-tuning data-intensive application on large clusters emphasizing on parallelism configurations. Rubén Tous, Anastasios Gounaris, Carlos Tripiana, Jordi Torres, Sergi Girona, Eduard Ayguadé, Jesús Labarta, Yolanda Becerra 0001, David Carrera 0001, Mateo Valero |
IEEE BigData | 7 |
| 2015 | Fault-Tolerant Protocol for Hybrid Task-Parallel Message-Passing ApplicationsabstractWe present a fault-tolerant protocol for task-parallel message-passing applications to mitigate transient errors. The protocol requires the restart only of the task that experienced the error and transparently handles any MPI calls inside the task. The protocol is implemented in Nanos -- a dataflow runtime for task-based OmpSs programming model -- and the PMPI profiling layer to fully support hybrid OmpSs+MPI applications. In our experiments we demonstrate that our fault-tolerant solution has a reasonable overhead, with a maximum observed overhead of 4.5%. We also show that fine-grained parallelization is important for hiding the overheads related to the protocol as well as the recovery of tasks. Tatiana V. Martsinkevich, Omer Subasi, Osman S. Unsal, Franck Cappello, Jesús Labarta |
CLUSTER | 5 |
| 2015 | Runtime-Aware Architectures
Marc Casas, Miquel Moretó, Lluc Alvarez, Emilio Castillo, Dimitrios Chasapis, Timothy Hayes 0001, Luc Jaulmes, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
Euro-Par | 12 |
| 2015 | Low-Overhead Detection of Memory Access Patterns and Their Time Evolution
Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
Euro-Par | 5 |
| 2015 | Collective Offload for Heterogeneous ClustersabstractExascale performance requires a level of energy efficiency only achievable with specialized hardware. Hence, for building a general purpose HPC system with Exascale performance different types of processors, memory technologies and interconnection networks will be necessary. Heterogeneous hardware is already present on some top supercomputer systems that are composed of different compute nodes, which at the same time, contain different types of processors and memories. Moreover, heterogeneous hardware is much harder to manage and exploit than homogeneous hardware, further increasing the complexity of applications that run on HPC systems. Most HPC applications use MPI to implement a rigid Single Program Multiple Data (SPMD) execution model that no longer fits the heterogeneous nature of the underlying hardware. However, MPI provides a powerful and flexible MPI_Comm_spawn API call that was designed to exploit heterogeneous hardware dynamically but at the expense of higher complexity, hindering a wider adoption of this API. In this paper, we have extended the OmpSs programming model to offload MPI kernels dynamically, replacing the low-level and more error-prone MPI_Comm_ spawn call with high-level and easier to use OmpSs pragmas. The evaluation shows that our proposal simplifies the dynamic offload of MPI kernels while keeping competitive performance and scaling to a high number of nodes. Florentino Sainz, Jorge Bellón, Vicenç Beltran 0001, Jesús Labarta |
HiPC | 4 |
| 2015 | Criticality-Aware Dynamic Task Scheduling for Heterogeneous ArchitecturesabstractCurrent and future parallel programming models need to be portable and efficient when moving to heterogeneous multi-core systems. OmpSs is a task-based programming model with dependency tracking and dynamic scheduling. This paper describes the OmpSs approach on scheduling dependent tasks onto the asymmetric cores of a heterogeneous system. The proposed scheduling policy improves performance by prioritizing the newly-created tasks at runtime, detecting the longest path of the dynamic task dependency graph, and assigning critical tasks to fast cores. While previous works use profiling information and are static, this dynamic scheduling approach uses information that is discoverable at runtime which makes it implementable and functional without the need of an oracle or profiling. The evaluation results show that our proposal outperforms a dynamic implementation of Heterogeneous Earliest Finish Time by up to 1.15x, and the default breadth-first OmpSs scheduler by up to 1.3x in an 8-core heterogeneous platform and up to 2.7x in a simulated 128-core chip. Kallia Chronaki, Alejandro Rico, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
ICS | 5 |
| 2015 | Quiet Neighborhoods: Key to Protect Job Performance PredictabilityabstractInterference of nearby jobs has been recently identified as the dominant reason for the high performance variability of parallel applications running on High Performance Computing (HPC) systems. Typically, HPC systems are dynamic with multiple jobs coming and leaving in an unpredictable fashion, sharing simultaneously the system interconnection network. In such environment contention for network resources is causing random stalls in the progress of application execution degrading application and system performance overall. Eliminating job interactions in their neighbourhoods is key for guaranteeing performance predictability of applications. In this paper we are proposing the concept of quiet neighbourhoods that significantly reduce job interactions. Quiet neighbourhoods are created by the system resource manager in two phases. First, multiple virtual network blocks are defined on the top of the physical network resources based on typical workload distributions. Second, newly arriving jobs are allocated in these virtual blocks based on their size. Ana Jokanovic, José Carlos Sancho, Alejandro Lucero, Cyriel Minkenberg, Jesús Labarta |
IPDPS | 6 |
| 2015 | NanoCheckpoints: A Task-Based Asynchronous Dataflow Framework for Efficient and Scalable Checkpoint/RestartabstractIn this paper, we present NanoCheckpoints which is a lightweight software-based checkpoint/restart scheme for task-parallel HPC applications. We leverage OmpSs, a task-based OpenMP derivative programming model (PM) and its Nanos asynchronous dataflow runtime. NanoCheckpoints achieves minimal overheads by check pointing only tasks' inputs which are available for free in the OmpSs PM. We evaluate NanoCheckpoints by both pure task-parallel shared memory benchmarks (up to 16 cores) and hybrid OmpSs+MPI applications (up to 1024 cores). The results indicate that NanoCheckpoints has on average overhead 3% for shared memory benchmarks. The dataflow semantics of Nanos, where both check pointing and error recovery are asynchronous, allows NanoCheckpoints to scale at large core counts even when high error rates are present. For hybrid OmpSs+MPI benchmarks, NanoCheckpoints has very low overhead, on average 2%, and high scalability. Javier Arias Moreno, Osman S. Unsal, Jesús Labarta, Adrián Cristal |
PDP | 3 |
| 2015 | Exploiting asynchrony from exact forward recovery for DUE in iterative solversabstractThis paper presents a method to protect iterative solvers from Detected and Uncorrected Errors (DUE) relying on error detection techniques already available in commodity hardware. Detection operates at the memory page level, which enables the use of simple algorithmic redundancies to correct errors. Such redundancies would be inapplicable under coarse grain error detection, but become very powerful when the hardware is able to precisely detect errors. Luc Jaulmes, Marc Casas, Miquel Moretó, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
SC | 5 |
| 2014 | ALOJA: A systematic study of Hadoop deployment variables to enable automated characterization of cost-effectivenessabstractThis article presents the ALOJA project, an initiative to produce mechanisms for an automated characterization of cost-effectiveness of Hadoop deployments and reports its initial results. ALOJA is the latest phase of a long-term collaborative engagement between BSC and Microsoft which, over the past 6 years has explored a range of different aspects of computing systems, software technologies and performance profiling. While during the last 5 years, Hadoop has become the de-facto platform for Big Data deployments, still little is understood of how the different layers of the software and hardware deployment options affects its performance. Early ALOJA results show that Hadoop's runtime performance, and therefore its price, are critically affected by relatively simple software and hardware configuration choices e.g., number of mappers, compression, or volume configuration. Project ALOJA presents a vendor-neutral repository featuring over 5000 Hadoop runs, a test bed, and tools to evaluate the cost-effectiveness of different hardware, parameter tuning, and Cloud services for Hadoop. As few organizations have the time or performance profiling expertise, we expect our growing repository will benefit Hadoop customers to meet their Big Data application needs. ALOJA seeks to provide both knowledge and an online service to with which users make better informed configuration choices for their Hadoop compute infrastructure whether this be on-premise or cloud-based. The initial version of ALOJA's Web application and sources are available at http://hadoop.bsc.es Nicolás Poggi, David Carrera 0001, Aaron Call, Sergio Mendoza, Yolanda Becerra 0001, Jordi Torres, Eduard Ayguadé, Fabrizio Gagliardi, Jesús Labarta, Rob Reinauer, Nikola Vujic, Daron Green, José A. Blakeley |
IEEE BigData | 9 |
| 2014 | Identifying Code Phases Using Piece-Wise Linear RegressionsabstractNode-level performance is one of the factors that may limit applications from reaching the supercomputers' peak performance. Studying node-level performance and attributing it to the source code results into valuable insight that can be used to improve the application efficiency, albeit performing such a study may be an intimidating task due to the complexity and size of the applications. We present in this paper a mechanism that takes advantage of combining piece-wise linear regressions, coarse-grain sampling, and minimal instrumentation to detect performance phases in the computation regions even if their granularity is very fine. This mechanism then maps the performance of each phase into the application syntactical structure displaying a correlation between performance and source code. We introduce a methodology on top of this mechanism to describe the node-level performance of parallel applications, even for first-time seen applications. Finally, we demonstrate the methodology describing optimized in-production applications and further improving their performance applying small transformations to the code based on the hints discovered. Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
IPDPS | 5 |
| 2014 | Hints to improve automatic load balancing with LeWI for hybrid applications
Marta Garcia-Gasulla, Jesús Labarta, Julita Corbalán |
J. Parallel Distributed Comput. | 2 |
| 2014 | Scheduling parallel jobs on multicore clusters using CPU oversubscription
Gladys Utrera, Julita Corbalán, Jesús Labarta |
J. Supercomput. | 3 |
| 2013 | Topic 1: Support Tools and Environments - (Introduction)
Bronis R. de Supinski, Bettina Krammer, Karl Fürlinger, Jesús Labarta, Dimitrios S. Nikolopoulos |
Euro-Par | 4 |
| 2013 | Implementing OmpSs support for regions of data in architectures with multiple address spacesabstractThe need for features for managing complex data accesses in modern programming models has increased due to the emerging hardware architectures. HPC hardware has moved towards clusters of accelerators and/or multicores, architectures with a complex memory hierarchy exposed to the programmer. Javier Bueno, Xavier Martorell, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta |
ICS | 5 |
| 2013 | Programmable and Scalable Reductions on ClustersabstractReductions matter and they are here to stay. Wide adoption of parallel processing hardware in a broad range of computer applications has encouraged recent research efforts on their efficient parallelization. Furthermore, trends towards high productivity languages in mainstream computing increases the demand for efficient programming support. In this paper we present a new approach on parallel reductions for distributed memory systems that provides both scalability and programmability. Using OmpSs, a task-based parallel programming model, the developer has the ability to express scalable reductions through a single pragma annotation. This pragma annotation is applicable for tasks as well as for work-sharing constructs (with implicit tasking) and instructs the compiler to generate the required runtime calls. The supporting runtime handles data and task distribution, parallel execution and data reduction. Scalability is achieved through a software cache that maximizes local and temporal data reuse and allows overlapped computation and communication. Results confirm scalability for up to 32 12-core cluster nodes. Jan Ciesko, Javier Bueno, Nikola Puzovic, Alex Ramírez, Rosa M. Badia, Jesús Labarta |
IPDPS | 6 |
| 2013 | Self-Adaptive OmpSs Tasks in Heterogeneous EnvironmentsabstractAs new heterogeneous systems and hardware accelerators appear, high performance computers can reach a higher level of computational power. Nevertheless, this does not come for free: the more heterogeneity the system presents, the more complex becomes the programming task in terms of resource management. OmpSs is a task-based programming model and framework focused on the runtime exploitation of parallelism from annotated sequential applications. This paper presents a set of extensions to this framework: we show how the application programmer can expose different specialized versions of tasks (i.e. pieces of specific code targeted and optimized for a particular architecture) and how the system can choose between these versions at runtime to obtain the best performance achievable for the given application. From the results obtained in a multi-GPU system, we prove that our proposal gives flexibility to application's source code and can potentially increase application's performance. Judit Planas, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta |
IPDPS | 4 |
| 2013 | Identifying Critical Code Sections in Dataflow Programming ModelsabstractThe years of practice in optimizing applications point that the major issue is focus - identifying the critical code section whose optimization would yield the highest overall speedup. While this issue is mainly solved for sequential applications, it remains a serious hurdle in the world of parallel computing. Furthermore, the newest dataflow parallel programming models expose very irregular parallelism, making the identification of the critical code section even harder. To address this issue, we designed an environment that identifies critical code sections in applications. The programmer can use this environment to estimate the potential benefits of the optimization for a specific parallel platform. This is very important because the programmer can anticipate the benefits of his optimization and assure that the optimization is worth the effort. Furthermore, we showed that in many applications, the choice of the critical code section decisively depends on the configuration of the target machine. For instance, in HP Linpack, optimizing a task that takes 0.49% of the total computation time yields the overall speedup of less than 0.25% on one machine, and at the same time, yields the overall speedup of more than 24% on a machine with different number of cores. Vladimir Subotic, José Carlos Sancho, Jesús Labarta, Mateo Valero |
PDP | 3 |
| 2013 | On the usefulness of object tracking techniques in performance analysisabstractUnderstanding the behavior of a parallel application is crucial if we are to tune it to achieve its maximum performance. Yet the behavior the application exhibits may change over time and depend on the actual execution scenario: particular inputs and program settings, the number of processes used, or hardware-specific problems. So beyond the details of a single experiment a far more interesting question arises: how does the application behavior respond to changes in the execution conditions? Germán Llort, Harald Servat, Juan Gonzalez, Judit Giménez, Jesús Labarta |
SC | 5 |
| 2013 | Framework for a productive performance optimization
Harald Servat, Germán Llort, Kevin A. Huck, Judit Giménez, Jesús Labarta |
Parallel Comput. | 5 |
| 2012 | A Job Scheduling Approach for Multi-core Clusters Based on Virtual Malleability
Gladys Utrera, Siham Tabik, Julita Corbalán, Jesús Labarta |
Euro-Par | 4 |
| 2012 | Tools for Power-Energy Modelling and Analysis of Parallel Scientific ApplicationsabstractUnderstanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core. Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ICPP | 3 |
| 2012 | On-the-Fly Adaptive Routing in High-Radix Hierarchical NetworksabstractDragonfly networks have been recently proposed for the interconnection network of forthcoming exascale supercomputers. Relying on large-radix routers, they build a topology with low diameter and high throughput, divided into multiple groups of routers. While minimal routing is appropriate for uniform traffic patterns, adversarial traffic patterns can saturate inter-group links and degrade the obtained performance. Such traffic patterns occur in typical communication patterns used by many HPC applications, such as neighbor data exchanges in multi-dimensional space decompositions. Non-minimal traffic routing is employed to handle such cases. Adaptive policies have been designed to select between minimal and nonminimal routing to handle variable traffic patterns. However, previous papers have not taken into account the effect of saturation of intra-group (local) links. This paper studies how local link saturation can be common in these networks, and shows that it can largely reduce the performance. The solution to this problem is to use nonminimal paths that avoid those saturated local links. However, this extends the maximum path length, and since all previous routing proposals prevent deadlock by relying on an ascending order of virtual channels, it would imply unaffordable cost and complexity in the network routers. In this paper we introduce a novel routing/flow-control scheme that decouples the routing and the deadlock avoidance mechanisms. Our model does not impose any dependencies between virtual channels, allowing for on-the-fly (in-transit) adaptive routing of packets. To prevent deadlock we employ a deadlock-free escape sub network based on injection restriction. Simulations show that our model obtains lower latency, higher throughput, and faster adaptation to transient traffic, because it dynamically exploits a higher path diversity to avoid saturated links. Notably, our proposal consumes traffic bursts 43% faster than previous ones. Marina García, Enrique Vallejo 0001, Ramón Beivide, Miguel Odriozola, Cristobal Camarero, Mateo Valero, Jesús Labarta, Cyriel Minkenberg |
ICPP | 8 |
| 2012 | Productive Programming of GPU Clusters with OmpSsabstractClusters of GPUs are emerging as a new computational scenario. Programming them requires the use of hybrid models that increase the complexity of the applications, reducing the productivity of programmers. We present the implementation of OmpSs for clusters of GPUs, which supports asynchrony and heterogeneity for task parallelism. It is based on annotating a serial application with directives that are translated by the compiler. With it, the same program that runs sequentially in a node with a single GPU can run in parallel in multiple GPUs either local (single node) or remote (cluster of GPUs). Besides performing a task-based parallelization, the runtime system moves the data as needed between the different nodes and GPUs minimizing the impact of communication by using affinity scheduling, caching, and by overlapping communication with the computational task. We show several applications programmed with OmpSs and their performance with multiple GPUs in a local node and in remote nodes. The results show good tradeoff between performance and effort from the programmer. Javier Bueno, Judit Planas, Alejandro Duran, Rosa M. Badia, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
IPDPS | 7 |
| 2012 | The Network Adapter: The Missing Link between MPI Applications and Network PerformanceabstractNetwork design aspects that influence cost and performance can be classified according to their distance from the applications, into issues concerning topology, switch technology, link technology, network adapter, and communication library. The network adapter has a privileged position to take decisions with more global information than any other component in the network. It receives feedback from the switches and requests from the communication libraries and applications. Also, compared to a network switch, an adapter has access to significantly more memory (host memory and on-chip memory) and memory bandwidth (which typically exceeds network bandwidth). The potential of the adapter to improve global network performance has not yet been fully exploited. In this work we show a series of noticeable performance improvements (of at least 10% to 15%) for medium-sized message exchanges in typical HPC communication patterns by optimizing message segmentation and packet injection policies, that can be implemented in an adapter's firmware inexpensively. We also show that implementing equivalent solutions in the switch (as opposed to the adapter) leads to only marginal performance improvements as the ones obtained by controlling the segmentation and injection policy at the adapter, while involving significantly more cost. In addition, enhancing the adapter will lead to less hardware complexity in the switches, thus reducing cost and energy consumption. Cyriel Minkenberg, Ronald P. Luijten, Ramón Beivide, Patrick Geoffray, Jesús Labarta, Mateo Valero, Stephen W. Poole |
SBAC-PAD | 6 |
| 2012 | A high-productivity task-based programming model for clustersabstractSUMMARY Programming for large‐scale, multicore‐based architectures requires adequate tools that offer ease of programming and do not hinder application performance. StarSs is a family of parallel programming models based on automatic function‐level parallelism that targets productivity. StarSs deploys a data‐flow model: it analyzes dependencies between tasks and manages their execution, exploiting their concurrency as much as possible. This paper introduces Cluster Superscalar (ClusterSs), a new StarSs member designed to execute on clusters of SMPs (Symmetric Multiprocessors). ClusterSs tasks are asynchronously created and assigned to the available resources with the support of the IBM APGAS runtime, which provides an efficient and portable communication layer based on one‐sided communication. We present the design of ClusterSs on top of APGAS, as well as the programming model and execution runtime for Java applications. Finally, we evaluate the productivity of ClusterSs, both in terms of programmability and performance and compare it to that of the IBM X10 language. Copyright © 2012 John Wiley & Sons, Ltd. Enric Tejedor, Montse Farreras, David Grove, Rosa M. Badia, Gheorghe Almási 0001, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 6 |
| 2012 | Understanding the future of energy-performance trade-off via DVFS in HPC environments
Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero |
J. Parallel Distributed Comput. | 3 |
| 2012 | Parallel job scheduling for power constrained HPC systems
Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero |
Parallel Comput. | 3 |
| 2011 | Parallel Implementation of the Integral Histogram
Pieter Bellens, Kannappan Palaniappan, Rosa M. Badia, Guna Seetharaman, Jesús Labarta |
ACIVS | 5 |
| 2011 | Productive Cluster Programming with OmpSs
Javier Bueno, Luis Martinell, Alejandro Duran, Montse Farreras, Xavier Martorell, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta |
Euro-Par (1) | 8 |
| 2011 | Quantifying the Potential Task-Based Dataflow Parallelism in MPI Applications
Vladimir Subotic, Roger Ferrer, José Carlos Sancho, Jesús Labarta, Mateo Valero |
Euro-Par (1) | 4 |
| 2011 | ClusterSs: a task-based programming model for clustersabstractProgramming for large-scale, multicore-based architectures requires adequate tools that offer ease of programming while not hindering application performance. StarSs is a family of parallel programming models based on automatic function level parallelism that targets productivity. StarSs deploys a data-flow model: it analyses dependencies between tasks and manages their execution, exploiting their concurrency as much as possible. We introduce Cluster Superscalar (ClusterSs), a new StarSs member designed to execute on clusters of SMPs. ClusterSs tasks are asynchronously created and assigned to the available resources with the support of the IBM APGAS runtime, which provides an efficient and portable communication layer based on one-sided communication.This short paper gives an overview of the ClusterSs design on top of APGAS, as well as the conclusions of a productivity study; in this study, ClusterSs was compared to the IBM X10 language, both in terms of programmability and performance. A technical report is available with the details. Enric Tejedor, Montse Farreras, David Grove, Rosa M. Badia, Gheorghe Almási 0001, Jesús Labarta |
HPDC | 6 |
| 2011 | Trace Spectral Analysis toward Dynamic Levels of DetailabstractThe emergence of Petascale systems has raised new challenges to performance analysis tools. Understanding every single detail of an execution is important to bridge the gap between the theoretical peak and the actual performance achieved. Tracing tools are the best option when it comes to providing detailed information about the application behavior, but not without liabilities. The amount of information that a single execution can generate grows so fast that it easily becomes unmanageable. An effective analysis in such scenarios necessitates the intelligent selection of information. In this paper we present an on-line performance tool based on spectral analysis of signals that automatically identifies the different computing phases of the application as it runs, selects a few representative periods and decides the granularity of the information gathered for these regions. As a result, the execution is completely characterized at different levels of detail, reducing the amount of data collected while maximizing the amount of useful information presented for the analysis. Germán Llort, Marc Casas, Harald Servat, Kevin A. Huck, Judit Giménez, Jesús Labarta |
ICPADS | 6 |
| 2011 | Unveiling Internal Evolution of Parallel Application Computation PhasesabstractAs access to supercomputing resources is becoming more and more commonplace, performance analysis tools are gaining importance in order to decrease the gap between the application performance and the supercomputers' peak performance. Performance analysis tools allow the analyst to understand the idiosyncrasies of an application in order to improve it. However, these tools require monitoring regions of the application to provide information to the analysts, leaving non-monitored regions of code unknown, which may result in lack of understanding of important regions of the application. In this paper we describe an automated methodology that reports very detailed application insights and improves the analysis experience of performance tools based on traces. We apply this methodology to three production applications and provide suggestions on how to improve their performance. Our methodology uses computation burst clustering and a mechanism called folding. While clustering automatically detects application structure, folding combines instrumentation and sampling to augment the performance analysis details. Folding provides fine grain performance information from coarse grain sampling on iterative applications. Folding results closely resemble the performance data gathered from fine grain sampling with an absolute mean difference less than 5% without overhead of fine grain. Harald Servat, Germán Llort, Judit Giménez, Kevin A. Huck, Jesús Labarta |
ICPP | 5 |
| 2011 | Poster: programming clusters of GPUs with OMPSsabstractOmpSs is a programming model that provides an environment to develop parallel applications for cluster environments with heterogeneous architectures. Based on OpenMP and StarSs, it offers a set of compiler directives that can be used to annotate a sequential code. Additional features have been added to support the use of accelerators like GPUs. This schema offers a high productivity environment due to its simplicity compared to other models like MPI. Our current implementation has shown a good performance when running different benchmarks. Javier Bueno, Alejandro Duran, Xavier Martorell, Eduard Ayguadé, Rosa M. Badia, Jesús Labarta |
ICS | 6 |
| 2011 | A Study of Speculative Distributed Scheduling on the Cell/B.EabstractStar Superscalar's (StarSs) programming model converts a sequential application in C or Fortran into an efficient parallel program. The resulting parallel code is highly dynamic in the sense that data analysis and task scheduling occur at run-time, while the application executes. In this paper we compare this approach to the strategy adopted by other multi-core programming environments. The prize to pay for dynamic scheduling and dependence tracking is higher runtime overhead. We propose a distributed scheduler for Task Dependence Graphs (TDGs) to attenuate the scheduling cost in heterogeneous multi-core architectures. This scheduler allows the cores to speculatively select tasks from a conservative estimate of the TDG. In case of conflicts or lack of tasks a lightweight centralized scheduler services the faulting core after which the latter resumes its participation in the distributed scheme. Experiments with Cell Super scalar (CellSs) on a representative set of benchmarks demonstrate the reduction in runtime overhead achieved by the distributed scheduler. This reduction in runtime overhead carries over directly to a performance improvement for a large fraction of the benchmarks. Pieter Bellens, Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
IPDPS | 4 |
| 2011 | The Impact of Application's Micro-Imbalance on the Communication-Computation OverlapabstractAlthough the community sees overlapping communication and computation as a perspective avenue for advancing parallel execution, it remains unclear what type of applications, under which conditions, and to which extent could benefit from this technique. To tackle this issue, we designed a simulation environment that allowed us to profoundly study overlap. We found out that overlapping potential in an application is determined by the application's parallel behavior and the pattern by which each process locally produces/consumes data involved in communication. We identified two behaviors that directly influence the application's overlapping potential - we name them microscopic imbalance of computation and microscopic imbalance of communication. In an application that expresses some of these two behaviors, a fine-grain overlapping technique can achieve a significant execution speedup, a speedup that can even be higher than 2. We believe that our findings can help a programmer estimate how much his application could benefit from overlap, and therefore decide whether implementing that technique is worth the effort. Vladimir Subotic, José Carlos Sancho, Jesús Labarta, Mateo Valero |
PDP | 3 |
| 2011 | Extracting the optimal sampling frequency of applications using spectral analysisabstractSUMMARY The research community have agreed on several applications as benchmarks to evaluate the adequateness of architectures and high performance computing infrastructures. The performance of these benchmarks is used to determine the weaknesses and strengths of novel designs. Therefore, the performance evaluation of benchmarks is a key factor in the process of designing new architectures. In this paper, we propose a new method based on spectral analysis that allows to perform an automatic analysis of benchmarks' executions. The output of the new method is a representative segment of the benchmarks' executions. Given the nature of the method, the optimal sampling interval length of applications is obtained. This method complements and improves existing techniques focused on the reduction of the application's instruction execution stream of sequential benchmarks and enables the extraction of significant performance information of parallel benchmarks without executing the whole application. The results obtained with the SPEC CPU2000 and the NAS Parallel Benchmarks demonstrate the efficiency and benefits of the approach. Copyright © 2011 John Wiley & Sons, Ltd. Marc Casas, Harald Servat, Rosa M. Badia, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | A Simulation Framework to Automatically Analyze the Communication-Computation Overlap in Scientific ApplicationsabstractOverlapping communication and computation has been devised as an attractive technique to alleviate the huge application's network requirements at large scale. Overlapping will allow to fully or partially hide the long communication delays suffered when transferring messages through the network. This will relax the application's network requirements, and hence allow to deploy more cost-effective network designs. However, today's scientific applications make little use of overlapping. In addition, there is no support to analyze how overlap could impact the performance of real scientific applications. In this paper we address this issue by presenting a simulation framework to automatically analyze the benefits of communication-computation overlap. The simulation framework consists of a binary translation tool (Valgrind), a distributed machine simulator (Dimemas), and a visualization tool (Paraver). Valgrind instruments the legacy MPI application and generates the execution traces, then Dimemas uses the obtained traces and reconstructs the application's time-behavior on a configurable parallel platform, and finally Paraver visualizes the obtained time-behaviors. Our simulation methodology brings two new features into the study of overlap: 1) automatic simulation of the overlapped execution - as there is no need for code restructuring in applications; and 2) visualization of simulated time behaviors, that further allows useful comparisons of the non-overlapped and the overlapped executions. Vladimir Subotic, José Carlos Sancho, Jesús Labarta, Mateo Valero |
CLUSTER | 3 |
| 2010 | Performance Data Extrapolation in Parallel CodesabstractMeasuring the performance of parallel codes is a compromise between lots of factors. The most important one is which data has to be analyzed. Current supercomputers are able to run applications in large number of processors as well as the analysis data that can be extracted is also large and varied. That implies a hard compromise between the potential problems one want to analyze and the information one is able to capture during the application execution. In this paper we present an extrapolation methodology to maximize the information extracted in a single application execution. It is based on a structural characterization of the applications, performed using clustering techniques, the ability to multiplex the read of performance hardware counters, plus a projection process. As a result, we obtain the approximated values of a large set of metrics for each phase of the application, with minimum error. Juan Gonzalez, Judit Giménez, Jesús Labarta |
ICPADS | 3 |
| 2010 | Detailed Load Balance Analysis of Large Scale Parallel ApplicationsabstractBalancing the workload in parallel applications is a difficult task, even in conventional cases. Many computing cycles are wasted when the load is not evenly balanced across processing nodes. Global load balance analysis may determine that an application is well balanced, when in fact the application has hidden inefficiencies. In this paper, we consider the load balance of parallel applications which present unique challenges in the analysis process. We have performed trace analysis and simulation to demonstrate the existence of otherwise undiscovered performance issues. We also demonstrate that by collecting dynamic phase profiles, we are able to approximate the analysis results of trace analysis and simulation, and more accurately represent the performance behavior of complex parallel applications than through flat or callpath profiles alone. Kevin A. Huck, Jesús Labarta |
ICPP | 2 |
| 2010 | Overlapping communication and computation by using a hybrid MPI/SMPSs approachabstractCommunication overhead is one of the dominant factors affecting performance in high-end computing systems. To reduce the negative impact of communication, programmers overlap communication and computation by using asynchronous communication primitives. This increases code complexity, requiring more development effort and making less readable programs. This paper presents the hybrid use of MPI and SMPSs (SMP superscalar, a task-based shared-memory programming model), allowing the programmer to easily introduce the asynchrony necessary to overlap communication and computation. We also describe implementation issues in the SMPSs run time that support its efficient interoperation with MPI. We demonstrate the hybrid use of MPI/SMPSs with four application kernels (matrix multiply, Jacobi, conjugate gradient and NAS BT) and with the high-performance LINPACK benchmark. For the application kernels, the hybrid MPI/SMPSs versions significantly improve the performance of the pure MPI counterparts. For LINPACK we get close to the asymptotic performance at relatively small problem sizes and still get significant benefits at large problem sizes. In addition, the hybrid MPI/SMPSs approach substantially reduces code complexity and is less sensitive to network bandwidth and operating system noise than the pure MPI versions. Vladimir Marjanovic, Jesús Labarta, Eduard Ayguadé, Mateo Valero |
ICS | 2 |
| 2010 | Handling task dependencies under strided and aliased referencesabstractThe emergence of multicore processors has increased the need for simple parallel programming models usable by nonexperts. The ability to specify subparts of a bigger data structure is an important trait of High Productivity Programming Languages. Such a concept can also be applied to dependency-aware task-parallel programming models. In those paradigms, tasks may have data dependencies, and those are used for scheduling them in parallel. Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
ICS | 3 |
| 2010 | On-line detection of large-scale parallel application's structureabstractWith larger and larger systems being constantly deployed, trace-based performance analysis of parallel applications has become a challenging task. Even if the amount of performance data gathered per single process is small, traces rapidly become unmanageable when merging together the information collected from all processes. In general, an efficient analysis of such a large volume of data is subject to a previous filtering step that directs the analyst's attention towards what is meaningful to understand the observed application behavior. Furthermore, the iterative nature of most scientific applications usually ends up producing repetitive information. Discarding irrelevant data aims at reducing both the size of traces, and the time required to perform the analysis and deliver results. In this paper, we present an on-line analysis framework that relies on clustering techniques to intelligently select the most relevant information to understand how the application behaves, while keeping the volume of performance data at a reasonable size. Germán Llort, Juan Gonzalez, Harald Servat, Judit Giménez, Jesús Labarta |
IPDPS | 5 |
| 2010 | Simulation environment for studying overlap of communication and computationabstractOverlapping communication and computation allows both processors and network to be utilized concurrently and leads to two clear benefits: overall speedup and a reduction in network performance requirements. Still, it remains unclear how much overlap can be actually achieved in practice - in real-world applications. This work designs a precise simulation environment that measures how much a scientific MPI application can profit from overlapping communication and computation. The simulation takes into account a wide range of application properties and allows to study overlap on the configurable platform. Additionally, the environment can visualize the simulated time-behaviors, so the non-overlapped and overlapped executions can be compared both quantitatively and qualitatively, providing new insights into the mechanism and potential of overlap. We found that the overlapping potential is very limited by pattern by which an application computes on the communicated data. Finally, we identified as the the biggest benefit of overlap the fact that it can highly relax network constraints without consequently degrading performance. Vladimir Subotic, Jesús Labarta, Mateo Valero |
ISPASS | 2 |
| 2010 | Task Superscalar: An Out-of-Order Task PipelineabstractWe present \emph{Task Super scalar}, an abstraction of instruction-level out-of-order pipeline that operates at the task-level. Like ILP pipelines, which uncover parallelism in a sequential instruction stream, task super scalar uncovers task-level parallelism among tasks generated by a sequential thread. Utilizing intuitive programmer annotations of task inputs and outputs, the task super scalar pipeline dynamically detects inter-task data dependencies, identifies task-level parallelism, and executes tasks out-of-order. Furthermore, we propose a design for a distributed task super scalar pipeline front end, that can be embedded into any many core fabric, and manages cores as functional units. We show that our proposed mechanism is capable of driving hundreds of cores simultaneously with non-speculative tasks, which allows our pipeline to sustain work windows consisting of tens of thousands of tasks. We further show that our pipeline can maintain a decode rate faster than 60ns per task and dynamically uncover data dependencies among as many as ~50,000 in-flight tasks, using 7MB of on-chip eDRAM storage. This configuration achieves speedups of 95-255x (average 183x) over sequential execution for nine scientific benchmarks, running on a simulated CMP with 256 cores. Task super scalar thus enables programmers to exploit many core systems effectively, while simultaneously simplifying their programming model. Yoav Etsion, Felipe Cabarcas, Alejandro Rico, Alex Ramírez, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
MICRO | 7 |
| 2010 | Effective communication and computation overlap with hybrid MPI/SMPSsabstractCommunication overhead is one of the dominant factors affecting performance in high-performance computing systems. To reduce the negative impact of communication, programmers overlap communication and computation by using asynchronous communication primitives. This increases code complexity, requiring more development effort and making less readable programs. This paper presents the hybrid use of MPI and SMPSs (SMP superscalar, a task-based shared-memory programming model) that allows the programmer to easily introduce the asynchrony necessary to overlap communication and computation. We demonstrate the hybrid use of MPI/SMPSs with the high-performance LINPACK benchmark (HPL), and compare it to the pure MPI implementation, which uses the look-ahead technique to overlap communication and computation. The hybrid MPI/SMPSs version significantly improves the performance of the pure MPI version, getting close to the asymptotic performance at medium problem sizes and still getting significant benefits at small/large problem sizes. Vladimir Marjanovic, Jesús Labarta, Eduard Ayguadé, Mateo Valero |
PPoPP | 2 |
| 2009 | Oblivious routing schemes in extended generalized Fat Tree networksabstractA family of oblivious routing schemes for fat trees and their slimmed versions is presented in this work. First, two popular oblivious routing algorithms, which we refer to as S-mod-k and D-mod-k, are analyzed in detail. S-mod-k is the default routing algorithm given as an example in the first works formally describing fat tree networks. D-mod-k has been independently proposed and investigated by several authors, who conclude in their evaluations that it achieves better performance than a random or adaptive routing approach. First, we identify the reasons why these algorithms perform well. Using this insight we extend these algorithms, originally intended for full bisection networks, to slimmed networks. Based on the lessons learned we propose a new generalized family of algorithms that provides a better oblivious solution than the existing ones for this class of networks. Moreover, this family extends the previous work from k-ary n-trees to the more general class of extended generalized fat trees. Cyriel Minkenberg, Ramón Beivide, Ronald P. Luijten, Jesús Labarta, Mateo Valero |
CLUSTER | 5 |
| 2009 | An Extension of the StarSs Programming Model for Platforms with Multiple GPUs
Eduard Ayguadé, Rosa M. Badia, Francisco D. Igual, Jesús Labarta, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 4 |
| 2009 | Graph-Based Task Replication for Workflow ApplicationsabstractThe Grid is an heterogeneous and dynamic environment which enables distributed computation. This makes it a technology prone to failures. Some related work uses replication to overcome failures in a set of independent tasks, and in workflow applications, but they do not consider possible resource limitations when scheduling the replicas. In this paper, we focus on the use of task replication techniques for workflow applications, trying to achieve not only tolerance to the possible failures in an execution, but also to speed up the computation without demanding the user to implement an application-level checkpoint, which may be a difficult task depending on the application. Moreover, we also study what to do when there are not enough resources for replicating all running tasks. We establish different priorities of replication depending on the graph of the workflow application, giving more priority to tasks with a higher output degree. We have implemented our proposed policy in the GRID superscalar system, and we have run the fastDNAml as an experiment to prove our objectives are reached. Finally, we have identified and studied a problem which may arise due to the use of replication in workflow applications: the replication wait time. Raül Sirvent, Rosa M. Badia, Jesús Labarta |
HPCC | 3 |
| 2009 | LeWI: A Runtime Balancing Algorithm for Nested ParallelismabstractWe present LeWI: a novel load balancing algorithm, that can balance applications with very different patterns of imbalance. Our algorithm can balance fine grain imbalances, non iterative applications and applications with irregular imbalance. To achieve this LeWI reassigns the computational resources of blocked processes to other processes more loaded. We have implemented LeWI within DLB a Dynamic Load Balancing Library developed by us. DLB helps parallel programming models to make the most of the computational power available with the minimum effort. It solves the imbalance among processes in applications with two levels of parallelism using the malleability of the inner level. The performance evaluation shows that LeWI, the novel balancing algorithm we are presenting in this paper, together with DLB is able to improve the performance of a different range of unbalanced applications and when applied to well balanced applications it does not introduce significant overhead. Therefore we present a mechanism that can be used with any hybrid application without needing a programmer to analyze the application nor modify it. Marta Garcia-Gasulla, Julita Corbalán, Jesús Labarta |
ICPP | 3 |
| 2009 | Exploring pattern-aware routing in generalized fat tree networksabstractNew static source routing algorithms for High Performance Computing (HPC) are presented in this work. The target parallel architectures are based on the commonly used fat-tree networks and their slimmed versions. The evaluation of such proposals and their comparison against currently used routing mechanisms have been driven by realistic traffic generated by HPC applications. Our experimental framework is based on the integration of two existing simulators, one replaying an MPI application and another simulating the network details. The resulting simulation platform has been fed with traces from real executions. Ramón Beivide, Cyriel Minkenberg, Jesús Labarta, Mateo Valero |
ICS | 4 |
| 2009 | Tools for scalable performance analysis on Petascale systemsabstractTools are becoming increasingly important to efficiently utilize the computing power available in contemporary large scale systems. The drastic increase in the size and the complexity of systems require tools to be scalable while producing meaning full and easily digestible information that may help the user pin-point problems at scale. The goal of this tutorial is to introduce some state-of-the-art performance tools from three different organizations to a diverse audience group. Together these tools provide a broad spectrum of capabilities necessary to analyze the performance of scientific and engineering applications on a variety of large and small scale systems. These tools include: • IBM High Performance Computing Toolkit: The IBM High Performance Computing Toolkit is a suite of performance-related tools and libraries to assist in application tuning. This toolkit is an integrated environment for performance analysis of sequential and parallel applications using the MPI and OpenMP paradigms. Scientists can collect rich performance data from selected parts of an execution, digest the data at a very high level, and plan for improvements within a single unified interface. It provides a common framework for IBM's mid-range server offerings, including pSeries and eSeries servers and Blue Gene systems, on both AIX and Linux. More information cab be found here: http://domino.research.ibm.com/comm/research_projects.nsf/pages/hpct.index.html • Scalable Performance Analysis of Large-Scale Applications (SCALASCA) Toolset: Scalasca is an open-source toolset that can be used to analyze the performance behavior of parallel applications and to identify opportunities for optimization. It has been specifically designed for use on large-scale systems including BlueGene and Cray XT, but is also well-suited for small- and medium-scale HPC platforms. Scalasca supports an incremental performance-analysis procedure that integrates runtime summaries with in-depth studies of concurrent behavior via event tracing, adopting a strategy of successively refined measurement configurations. A distinctive feature is the ability to identify wait states that occur, for example, as a result of unevenly distributed workloads. Especially when trying to scale communication-intensive applications to large processor counts, such wait states can present severe challenges to achieving good performance. Scalasca is developed by the Julich Supercomputing Centre and available under the New BSD open-source license. More information can be found here: http://www.scalasca.org • CEPBA Toolkit: The CEPBA-tools environment is a trace based analysis environment consisting with two major components, Paraver, a browser for traces obtained from a parallel run and Dimemas , a simulator to rebuild the time behavior of a parallel program from a trace. More information can be found here: http://www.bsc.es/plantillaF.php?cat_id=52 I-Hsin Chung, Seetharami R. Seelam, Bernd Mohr, Jesús Labarta |
IPDPS | 4 |
| 2009 | Power-aware load balancing of large scale MPI applicationsabstractPower consumption is a very important issue for HPC community, both at the level of one application or at the level of whole workload. Load imbalance of a MPI application can be exploited to save CPU energy without penalizing the execution time. An application is load imbalanced when some nodes are assigned more computation than others. The nodes with less computation can be run at lower frequency since otherwise they have to wait for the nodes with more computation blocked in MPI calls. A technique that can be used to reduce the speed is Dynamic Voltage Frequency Scaling (DVFS). Dynamic power dissipation is proportional to the product of the frequency and the square of the supply voltage, while static power is proportional to the supply voltage. Thus decreasing voltage and/or frequency results in power reduction. Furthermore, over-clocking can be applied in some CPUs to reduce overall execution time. This paper investigates the impact of using different gear sets, over-clocking, and application and platform properties to reduce CPU power. A new algorithm applying DVFS and CPU over-clocking is proposed that reduces execution time while achieving power savings comparable to prior work. The results show that it is possible to save up to 60% of CPU energy in applications with high load imbalance. Our results show that six gear sets achieve, on average, results close to the continuous frequency set that has been used as a baseline. Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero, Alexander V. Veidenbaum |
IPDPS | 3 |
| 2009 | Automatic detection of parallel applications computation phasesabstractAnalyzing parallel programs has become increasingly difficult due to the immense amount of information collected on large systems. The use of clustering techniques has been proposed to analyze applications. However, while the objective of previous works is focused on identifying groups of processes with similar characteristics, we target a much finer granularity in the application behavior. In this paper, we present a tool that automatically characterizes the different computation regions between communication primitives in message-passing applications. This study shows how some of the clustering algorithms which may be applicable at a coarse grain are no longer adequate at this level. Density-based clustering algorithms applied to the performance counters offered by modern processors are more appropriate in this context. This tool automatically generates accurate displays of the structure of the application as well as detailed reports on a broad range of metrics for each individual region detected. Juan Gonzalez, Judit Giménez, Jesús Labarta |
IPDPS | 3 |
| 2009 | Automatic Evaluation of the Computation Structure of Parallel ApplicationsabstractMany data mining techniques have been proposed for parallel applications performance analysis, the most interesting being clustering analysis. Most cases have been used to detect processors with similar behavior. In previous work, we presented a different approach: clustering was used to detect the computation structure of the applications and how these different computation phases behave. In this paper, we present a method to evaluate the accuracy of this structure detection. This new method is based on the Single Program Multiple Data (SPMD) paradigm exhibited by real parallel programs. Assuming an SPMD structure, we expect that all tasks of a parallel application execute the same operation sequence. Using a Multiple Sequence Alignment (MSA) algorithm, we check the sequence ordering of the detected clusters to evaluate the quality of the clustering results. Juan Gonzalez, Judit Giménez, Jesús Labarta |
PDCAT | 3 |
| 2009 | Impact of the Memory Hierarchy on Shared Memory Architectures in Multicore Programming ModelsabstractMany and multicore architectures put a big pressure in parallel programming but gives a unique opportunity to propose new programming models that automatically exploit the parallelism of these architectures. Open MP is a very well known standard that exploits parallelism in shared memory architectures. SMPSs has recently been proposed as a task based programming model that exploits the parallelism at the task level and takes into account data dependencies between tasks. However, besides parallelism in the programming, the memory hierarchy impact in many/multi core architectures is a feature of large importance. This paper presents an evaluation of these two programming models with regard to the impact of different levels of the memory hierarchy in the duration of the application. The evaluation is based on trace-files with hardware counters on the execution of a memory intensive benchmark in both programming models. Rosa M. Badia, Josep M. Pérez, Eduard Ayguadé, Jesús Labarta |
PDP | 4 |
| 2009 | Parallelizing dense and banded linear algebra libraries using SMPSsabstractAbstract The promise of future many‐core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP‐like pragmas and a run‐time system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS‐level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column‐major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run‐time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms‐by‐blocks in the FLAME argot) for most operations and storage‐by‐blocks (or block data layout) is already in place. Copyright © 2009 John Wiley & Sons, Ltd. Rosa M. Badia, José R. Herrero 0001, Jesús Labarta, Josep M. Pérez, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | Prediction of behavior of MPI applicationsabstractScalability and performance of applications is a very important issue today. As more complex have become high performance architectures, it is more complex to predict the behavior of a given application running on them. In this paper, we propose a methodology which automatically and quickly predicts, from a very limited number of runs using very few processors, the scalability and performance of a given application in a wide range of supercomputers taking into account details of the architecture and the network of the machines. Marc Casas, Rosa M. Badia, Jesús Labarta |
CLUSTER | 3 |
| 2008 | A dependency-aware task-based programming environment for multi-core architecturesabstractParallel programming on SMP and multi-core architectures is hard. In this paper we present a programming model for those environments based on automatic function level parallelism that strives to be easy, flexible, portable, and performant. Its main trait is its ability to exploit task level parallelism by analyzing task dependencies at run time. We present the programming environment in the context of algorithms from several domains and pinpoint its benefits compared to other approaches. We discuss its execution model and its scheduler. Finally we analyze its performance and demonstrate that it offers reasonable performance without tuning, and that it can rival highly tuned libraries with minimal tuning effort. Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
CLUSTER | 3 |
| 2008 | Supercomputing for the Future, Supercomputing from the Past (Keynote)
Mateo Valero, Jesús Labarta |
HiPEAC | 2 |
| 2008 | Automatic analysis of speedup of MPI applicationsabstractThe intricacy of high performance computing applications has been growing veryfast in the last years. Only skilled analysts are able to determine the factors that are undermining the performance of up-to-date applications. Analyst time is a very expensive resource and, for that reason, a strong effort to develop automatic performance analysis methodologies has been made by the scientific community. In this paper, we propose a methodology that is able to automatically detect the main performance problems of applications. This methodology is based on, first, a size reduction of the performance data obtained from the executions and, second, an analytical model obtained from this performance data which fits the speedup of the applications in terms of several parameters related to several performance issues. The paper also shows results obtained from real up-to-date applications and validates the conclusions automatically derived from the methodology. Marc Casas, Rosa M. Badia, Jesús Labarta |
ICS | 3 |
| 2008 | Balancing HPC applications through smart allocation of resources in MT processorsabstractMany studies have shown that load imbalancing causes significant performance degradation in high performance computing (HPC) applications. Nowadays, multi-threaded (MT1) processors are widely used in HPC for their good performance/energy consumption and performance/cost ratios achieved sharing internal resources, like the instruction window or the physical register. Some of these processors provide the software hardware mechanisms for controlling the allocation of processor's internal resources. In this paper, we show, for the first time, that by appropriately using these mechanisms, we are able to control the tasks speed, reducing the imbalance in parallel applications transparently to the user and, hence, reducing the total execution time. Our results show that our proposal leads to a performance improvement up to 18% for one of the NAS benchmark. For a real HPC application (much more dynamic than the benchmark) the performance improvement is 8.1%. Our results also show that, if resource allocation is not used properly, the imbalance of applications is worsened causing performance loss. Carlos Boneti, Roberto Gioiosa, Francisco J. Cazorla, Julita Corbalán, Jesús Labarta, Mateo Valero |
IPDPS | 5 |
| 2007 | Automatic Structure Extraction from MPI Applications Tracefiles
Marc Casas, Rosa M. Badia, Jesús Labarta |
Euro-Par | 3 |
| 2007 | Modeling the Impact of Resource Sharing in Backfilling Policies using the Alvio SimulatorabstractJob scheduling policies for HPC centers have been extensively studied during these last years, specially backfilling based policies. Almost all of these studies have been done using simulation tools. These tools evaluate the performance of scheduling policies using the workloads and the resource definition as an input. To the best of our knowledge, all the existent simulators use the runtime (either requested or real) provided in the workload as a basis of their simulations. However, the runtime of a job, even executed with a fixed number of processors, depends on runtime issues such as the specific resource selection policy used for allocate the jobs or the resource jobs requirements. This paper is the first part of a more complex research project that analyzes the impact in the system performance of considering the resource sharing of running jobs. With this purpose we have included in our job scheduler simulator (the Alvio simulator) a performance model that estimates the penalty introduced in the application runtime when sharing the memory bandwidth. Experiments have been conducted with two resource selection policies and we present both the impact from the point of view of global performance metrics, such as average slowdown, and per job impact such as percentage of penalized runtime. Francesc Guim 0001, Julita Corbalán, Jesús Labarta |
MASCOTS | 3 |
| 2007 | Prediction f Based Models for Evaluating Backfilling Scheduling PoliciesabstractThe research on the usage of prediction techniques in HPC scheduling policies rather than user estimates has increased it relevance these recent years. In the coming scheduling architectures, like grids and very heterogeneous computational resources, such techniques are having a crucial relevance due to users in most of the cases will not have enough information or enough skills for specify for how long will their jobs run. Many studies have analyzed the impact of the user runtime estimates accuracy in the performance of the scheduling policies. Using user runtime estimation models, such as the f-model, researchers have evaluated how the accuracy of the runtime estimates provided by the user at the job submission can affect the performance of the backfilling policies and its variants. However, these traditional estimation models can not applied to backfilling scheduling policies that use runtime predictions rather than user estimates. Clearly, predictions can not be characterized with these models. For instance because the underestimation of the runtime is not considered by them and obviously it can occurs. In this paper we describe and evaluate a set of f-model based prediction models that characterize the behavior that prediction techniques have shown in HPC centers. They have been designed for evaluate scheduling policies that use predictions rather than user estimates. Francesc Guim 0001, Julita Corbalán, Jesús Labarta |
PDCAT | 3 |
| 2007 | Monitoring and Analysis Framework for Grid MiddlewareabstractAs the use of complex grid middleware becomes widespread and more facilites are offered by these pieces of software, distributed grid applications are becoming more and more popular. But as grid middleware grows in size and offers more advanced features, they become more complex and heavier, as well as harder to tune. Since the performance of a distributed grid application can be strongly influenced by the operation of the underlying grid middleware, it becomes of extreme importance to study and analyse its behaviour and performance. In this paper we present the eDragon monitoring framework (eDMF), a set of tools that can be used for the instrumentation and analysis of grid middleware, and which provides a unique environment to study the performance of grid applications. The eDMF is composed of a set of specialised monitoring tools as well as by a flexible and powerful performance analysis platform. Additionally we also provide a practical application of the eDMF to the Globus toolkit 4 (GT4), one of the most extended and popular grid middleware, showing how it helped us in the detection and resolution of several job management problems observed in the GT4 middleware Ramon Nou, Ferran Julià, David Carrera 0001, Kevin Hogan, Jordi Caubet, Jesús Labarta, Jordi Torres |
PDP | 6 |
| 2006 | Uniform Job Monitoring using the HPC-Europa Single Point of Access
Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Jesús Labarta, Ariel Oleksiak, Tomasz Kuczynski, Dawid Szejnfeld, Jarek Nabrzyski |
CCGRID | 4 |
| 2006 | How the JSDL can Exploit the Parallelism?abstractThe description of the jobs is a very important issue for the scheduling and management of grid jobs. Since there are a lot of different languages for describing grid jobs, the GGF have presented the Job Submission Description Language (JSDL) to standardize the job submission language. We believe that the JSDL is a good solution but it has some deficiencies regarding the parallelism issues. In this paper, we propose an extension of the JSDL to specify the parallelism details of grid jobs. This extension is proposed in general terms for supporting current multilevel parallel applications and incoming approaches in parallel programming models. We also discus the suitability of the multilevel parallel programming models for grids, in particular the MPI+OpenMP since our project, the eNANOS project, is based on this hybrid programming model. Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Jesús Labarta |
CCGRID | 4 |
| 2006 | Including SMP in Grids as Execution Platform and Other Extensions in GRID SuperscalarabstractGRID superscalar provides a very easy to use programming environment for enabling applications on the grid. Although the system already has many features, there are some areas that we wanted to enhance. In this paper we present a new version of GRID superscalar based on code annotations that includes full renaming support for scalar, array and structure parameters. We also present a tracing mechanism that allows fine tuning GRID superscalar applications, and improved support for running on SMP hosts. Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
e-Science | 3 |
| 2006 | Topic 2: Performance Prediction and Evaluation
Jesús Labarta, Bernd Mohr, Allan Snavely, Jeffrey S. Vetter |
Euro-Par | 1 |
| 2006 | Scaling MPI to short-memory MPPs such as BG/LabstractScalability to large number of processes is one of the weaknesses of current MPI implementations. Standard implementations are able to scale to hundreds of nodes, but not beyond. The main problem in these implementations is that they assume some resources (for both data and control-data) will always be available to receive/process unexpected messages. As we will show, this is not always true, especially in short-memory machines like the BG/L that has 64K nodes but each node only has 512Mbytes of memory.The objective of this paper is to present one algorithm that improves the robustness of MPI implementations for short-memory MPPs, taking care of data and control-data reception, the system will scale up to any number of nodes. The proposed solution achieves this goal without any observable overhead when there are no memory problems. Furthermore, in the worst case, when memory resources are extremely scarce, the overhead will never double the execution time (and we should never forget that in this extreme situation, traditional MPI implementations would fail to execute). Montse Farreras, Toni Cortes, Jesús Labarta, Gheorghe Almási 0001 |
ICS | 3 |
| 2006 | Techniques supporting threadprivate in OpenMPabstractThis paper presents the alternatives available to support threadprivate data in OpenMP and evaluates them. We show how current compilation systems rely on custom techniques for implementing thread-local data. But in fact the ELF binary specification currently supports data sections that become threadprivate by default. ELF naming for such areas is thread-local storage (TLS). Our experiments demonstrate that implementing threadprivate based on the TLS support is very easy, and more efficient. This proposal goes in the same line as the future implementation of OpenMP on the GNU compiler collection. In addition, our experience with the use of threadprivate in OpenMP applications shows that usually it is better to avoid it. This is because threadprivate variables reside in common blocks and they impede the compiler to fully optimize the code. So it is better to keep threadprivate as a temporary technique only to ease porting MPI codes to OpenMP. Xavier Martorell, Marc González 0001, Alejandro Duran, Jairo Balart, Roger Ferrer, Eduard Ayguadé, Jesús Labarta |
IPDPS | 7 |
| 2006 | Memory - CellSs: a programming model for the cell BE architectureabstractIn this work we present Cell superscalar (CellSs) which addresses the automatic exploitation of the functional parallelism of a sequential program through the different processing elements of the Cell BE architecture. The focus in on the simplicity and flexibility of the programming model. Based on a simple annotation of the source code, a source to source compiler generates the necessary code and a runtime library exploits the existing parallelism by building at runtime a task dependency graph. The runtime takes care of the task scheduling and data handling between the different processors of this heterogeneous architecture. Besides, a locality-aware task scheduling has been implemented to reduce the overhead of data transfers. The approach has been implemented and tested with a set of examples and the results obtained since now are promising. Pieter Bellens, Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
SC | 4 |
| 2006 | Automatic Grid workflow based on imperative programming languagesabstractAbstract GRID superscalar is a Grid programming environment that enables one to parallelize the execution of sequential applications in computational Grids. The run‐time library automatically builds a task data‐dependence graph of the application and it can be seen as an implicit workflow system. The current interface supports C/C++ and Perl applications. The run‐time library is based on Globus Toolkit 2.x using GRAM and GSIFTP services. In this document we describe the GRID superscalar basics emphasizing those aspects related to Grid workflow, in particular the flexibility of using an imperative language to describe the application. Copyright © 2005 John Wiley & Sons, Ltd. Raül Sirvent, Josep M. Pérez, Rosa M. Badia, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 4 |
| 2006 | Running OpenMP applications efficiently on an everything-shared SDSM
Juan José Costa, Toni Cortes, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
J. Parallel Distributed Comput. | 5 |
| 2005 | Implementing phylogenetic inference with GRID superscalarabstractThe grid has appeared recently as a new computing paradigm. However, to make the use of the grid available to the scientific community, frameworks that enable to easily write applications and to run them efficiently in the grid should be provided. GRID superscalar has been specially designed to satisfy the two requirements mentioned above. This paper presents an implementation of a biological application, fastDNAml, using GRID superscalar. The objective is not only to demonstrate the performance that can be achieved, but the programmability of the framework. The description contains details of the fastDNAml implementation, new features of GRID superscalar and summary of results obtained. Vasilis Dialinos, Rosa M. Badia, Raül Sirvent, Josep M. Pérez, Jesús Labarta |
CCGRID | 5 |
| 2005 | Data Distribution Strategies for Domain Decomposition Applications in Grid Environments
Beatriz Otero, José María Cela, Rosa M. Badia, Jesús Labarta |
ICA3PP | 4 |
| 2005 | Another approach to backfilled jobs: applying virtual malleability to expired windowsabstractAn efficient job scheduling must ensure high throughput and good performance. Moreover in highly parallel systems where processors are a critical resource, high machine utilization becomes an essential aspect.Backfilling consists on moving jobs ahead in the queue, given that they do not delay certain previously submitted jobs. When the execution time of a backfilled job was underestimated, some action has to be taken with it: abort, suspend/resume, checkpoint/restart, remain executing.In this paper we propose an alternative choice for that situation which consists on apply Virtual Malleability to the backfilled job. This means that its processors partition will be reduced, and as MPI jobs aren't really malleable, we make the job contend with itself for the use of processors by applying Co-scheduling. In this way resources are freed and the job at the head of the queue have a chance to start executing. In addition to this, as MPI parallel jobs can be Moldable, we add this possibility to the scheme.We obtained better performance than traditional backfilling in about 25 %, especially in high machine utilization. We claim also for the portability of our technique which does not requires special support from the operating system as checkpointing does. Gladys Utrera, Julita Corbalán, Jesús Labarta |
ICS | 3 |
| 2005 | Performance-Driven Processor AllocationabstractIn current multiprogrammed multiprocessor systems, to take into account the performance of parallel applications is critical to decide an efficient processor allocation. In this paper, we present the performance-driven processor allocation policy (PDPA). PDPA is a new scheduling policy that implements a processor allocation policy and a multiprogramming-level policy, in a coordinated way, based on the measured application performance. With regard to the processor allocation, PDPA is a dynamic policy that allocates to applications the maximum number of processors to reach a given target efficiency. With regard to the multiprogramming level, PDPA allows the execution of a new application when free processors are available and the allocation of all the running applications is stable, or if some applications show bad performance. Results demonstrate that PDPA automatically adjusts the processor allocation of parallel applications to reach the specified target efficiency, and that it adjusts the multiprogramming level to the workload characteristics. PDPA is able to adjust the processor allocation and the multiprogramming level without human intervention, which is a desirable property for self-configurable systems, resulting in a better individual application response time. Julita Corbalán, Xavier Martorell, Jesús Labarta |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2004 | Generation of Simple Analytical Models for Message Passing Applications
Rosa M. Badia, Jesús Labarta |
Euro-Par | 3 |
| 2004 | Scheduling of MPI Applications: Self-co-scheduling
Gladys Utrera, Julita Corbalán, Jesús Labarta |
Euro-Par | 3 |
| 2004 | Dynamic Load Balancing of MPI+OpenMP ApplicationsabstractThe hybrid programming model MPI+OpenMP are useful to solve the problems of load balancing of parallel applications independently of the architecture. Typical approaches to balance parallel applications using two levels of parallelism or only MPI consist of including complex codes that dynamically detect which data domains are more computational intensive and either manually redistribute the allocated processors or manually redistribute data. This approach has two drawbacks: it is time consuming and it requires an expert in application analysis. In this paper we present an automatic and dynamic approach for load balancing MPI+OpenMP applications. The system calculates the percentage of load imbalance and decides a processor distribution for the MPI processes that eliminates the computational load imbalance. Results show that this method can balance effectively applications without analyzing nor modifying them and that in the cases that the application was well balanced does not incur in a great overhead for the dynamic instrumentation and analysis realized. Julita Corbalán, Alejandro Duran, Jesús Labarta |
ICPP | 3 |
| 2004 | Running OpenMP Applications Efficiently on an Everything-Shared SDSMabstractSummary form only given. Traditional software distributed shared memory (SDSM) systems modify the semantics of a real hardware shared memory system by relaxing the coherence semantic and by limiting the memory regions that are actually shared. These semantic modifications are done to improve performance of the applications using it. We show that a SDSM system that behaves like a real shared memory system (without the afore mentioned relaxations) can also be used to execute OpenMP applications and achieve similar speedups as the ones obtained by traditional SDSM systems. This performance can be achieved by encouraging the cooperation between the SDSM and the OpenMP runtime instead of relaxing the semantics of the shared memory. In addition, techniques like boundaries alignment and page presend are demonstrated as very useful to overcome the limitations of the current SDSM systems. Juan José Costa, Toni Cortes, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
IPDPS | 5 |
| 2003 | Performance Evaluation and Prediction
Jeffrey K. Hollingsworth, Allen D. Malony, Jesús Labarta, Thomas Fahringer |
Euro-Par | 3 |
| 2003 | Evaluation of the memory page migration influence in the system performance: the case of the SGI O2000abstractCurrent shared-memory multiprocessor CC-NUMA architectures provide a global address space to applications by hardware. However, even though the memory is virtually shared, it is actually physically distributed. Since memory nodes are distributed across the system, the cost of the memory accesses depends on the distance between the node that accesses the data and the node that physically contains the data. To reduce the impact of a bad initial memory placement, some operating systems offer a dynamic memory migration mechanism.In this paper, we want to demonstrate that memory migration mechanisms are a useful approach, but that their performance depends more on related issues, such as the processor scheduling, than on the mechanism itself. To show that, we evaluate the case of the automatic memory migration mechanism provided by IRIX, in Origin systems.We have evaluated several workloads of OpenMP applications under different system conditions such as the processor scheduling policy or the system load. In particular, we have focused on the effects of the page migration mechanism on the CPU time consumed by each application, the processor allocation received, and the speedup, when applying performance-driven scheduling policies.Results show that, if the scheduler is memory conscious, that is, it maintains as much as possible the system stable, the automatic memory page migration mechanism provided by IRIX will improve the execution time of OpenMPapplications. Experiments also show that the combination of performance-driven policies and the memory migration mechanism results in a system that can be automatically self-evaluated and self-configured. Julita Corbalán, Xavier Martorell, Jesús Labarta |
ICS | 3 |
| 2003 | Complete instrumentation requirements for performance analysis of Web based technologiesabstractIn this paper we present the eDragon environment, a research platform created to perform complete performance analysis of new Web-based technologies. eDragon enables the understanding of how application servers work in both sequential and parallel platforms offering a new insight in the usage of system resources. The environment is composed of a set of instrumentation modules, a performance analysis and visualization tool and a set of experimental methodologies to perform complete performance analysis of Web-based technologies. This paper describes the design and implementation of this research platform and highlights some of its main functionalities. We will also show how a detailed analytical view can be obtained through the application of a bottom-up strategy, starting with a group of system events and advancing to more complex performance metrics using a continuous derivation process. David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé, Jesús Labarta |
ISPASS | 5 |
| 2003 | Programming Grid Applications with GRID Superscalar
Rosa M. Badia, Jesús Labarta, Raül Sirvent, Josep M. Pérez, José María Cela, Rogeli Grima |
J. Grid Comput. | 2 |
| 2003 | Taking advantage of heterogeneity in disk arrays
Toni Cortes, Jesús Labarta |
J. Parallel Distributed Comput. | 2 |
| 2002 | On the Scalability of Tracing Mechanisms
Felix Freitag, Jordi Caubet, Jesús Labarta |
Euro-Par | 3 |
| 2002 | Performance Evaluation, Analysis and Optimization
Barton P. Miller, Jesús Labarta, Florian Schintke, Jens Simon |
Euro-Par | 2 |
| 2002 | A Trace-Scaling Agent for Parallel Application TracingabstractTracing and performance analysis tools are an important component in the development of high performance applications. Tracing parallel programs with current tracing tools, however, easily leads to large trace files with hundreds of Megabytes. The storage, visualization, and analysis of such trace files is often difficult. We propose a trace-scaling agent for tracing parallel applications, which learns the application behavior in runtime and achieves a small, easy to handle trace. The agent dynamically identifies the amount of information needed to capture the application behavior. This knowledge acquired at runtime allows recording only the non-iterative trace information, which drastically reduces the size of the trace file. Felix Freitag, Jordi Caubet, Jesús Labarta |
ICTAI | 3 |
| 2002 | A framework for performance modeling and predictionabstractCycle-accurate simulation is far too slow for modeling the expected performance of full parallel applications on large HPC systems. And just running an application on a system and observing wallclock time tells you nothing about why the application performs as it does (and is anyway impossible on yet-to-be-built systems). Here we present a framework for performance modeling and prediction that is faster than cycle-accurate simulation, more informative than simple benchmarking, and is shown useful for performance investigations in several dimensions. Allan Snavely, Laura Carrington, Nicole Wolter, Jesús Labarta, Rosa M. Badia, Avi Purkayastha |
SC | 4 |
| 2002 | Scheduler-Activated Dynamic Page Migration for Multiprogrammed DSM Multiprocessors
Dimitrios S. Nikolopoulos, Constantine D. Polychronopoulos, Theodore S. Papatheodorou, Jesús Labarta, Eduard Ayguadé |
J. Parallel Distributed Comput. | 4 |
| 2001 | Complex Pipelined Executions in OpenMP Parallel ApplicationsabstractThis paper proposes a set of extensions to the OpenMP programming model to express complex pipelined computations. This is accomplished by defining, in the form of directives, precedence relations among the tasks originated from work-sharing constructs. The proposal is based on the definition of a name space that identifies the work parceled out by these work-sharing constructs. Then the programmer defines the precedence relations using this name space. This relieves the programmer from the burden of defining complex synchronization data structures and the insertion of explicit synchronization actions in the program that make the program difficult to understand and maintain. This work is transparently done by the compiler with the support of the OpenMP runtime library. The proposal is motivated and evaluated with a synthetic multi-block example. The paper also includes a description of the compiler and runtime support in the framework of the NanosCompiler for OpenMP. Marc González 0001, Eduard Ayguadé, Xavier Martorell, Jesús Labarta |
ICPP | 4 |
| 2001 | Improving Gang Scheduling through job performance analysis and malleabilityabstractThe OpenMP programming model provides parallel applications a very important feature: job malleability. Job malleability is the capacity of an application to dynamically adapt its parallelism to the number of processors allocated to it. We believe that job malleability provides to applications the flexibility that a system needs to achieve its maximum performance. We also defend that a system has to take its decisions not only based on user requirements but also based on run-time performance measurements to ensure the efficient use of resources. Job malleability is the application characteristic that makes possible the run-time performance analysis. Without malleability applications would not be able to adapt their parallelism to the system decisions. To support these ideas, we present two new approaches to attack the two main problems of Gang Scheduling: the excessive number of time slots and the fragmentation. Our first proposal is to apply a scheduling policy inside each time slot of Gang Scheduling to distribute processors among applications considering their efficiency, calculated based on run-time measurements. We call this policy Performance-Driven Gang Scheduling. Our second approach is a new re-packing algorithm, Compress&Join, that exploits the job malleability. This algorithm modifies the processor allocation of running applications to adapt it to the system necessities and minimize the fragmentation and number of time slots. These proposals have been implemented in a SGI Origin 2000 with 64 processors. Results show the validity and convenience of both, to consider the job performance analysis calculated at run-time to decide the processor allocation, and to use a flexible programming model that adapts applications to system decisions. Julita Corbalán, Xavier Martorell, Jesús Labarta |
ICS | 3 |
| 2001 | The trade-off between implicit and explicit data distribution in shared-memory programming paradigmsabstractThis paper explores previously established and novel methods for scaling the performance of OpenMP on NUMA architectures. The spectrum of methods under investigation includes OS-level automatic page placement algorithms, dynamic page migrationd manual data distribution. The trade-off that these methods face lies between performance and programming effort. Automatic page placement algorithms are transparent to the programmer, but may compromise memory access locality. Dynamic page migration is also transparent, but requires careful engineering of online algorithms to be effective. Manual data distribution on the other requires substantial programming effort and architecture-specific extensions to OpenMP, but may localize memory accesses in a nearly optimal manner. Dimitrios S. Nikolopoulos, Eduard Ayguadé, Theodore S. Papatheodorou, Constantine D. Polychronopoulos, Jesús Labarta |
ICS | 5 |
| 2001 | Improving Processor Allocation through Run-Time Measured EfficiencyabstractIn a multiprocessor architecture it is very important to allocate processors to applications in a proportional way to the performance that applications are achieving, not considering this performance can result in an under-utilization of the multiprocessor and also it can slowdown the execution time of parallel applications. However the performance of parallel applications is not known before their execution. In this work, we propose to use dynamically measured application efficiency of OpenMP applications to improve the performance of two scheduling policies proposed so far, the equipartition and the equal efficiency. The modified scheduling policies will request parallel applications to achieve a target efficiency to receive more processors. We refer to the modified equipartition and equal efficiency as equip++ and equal eff++. We also propose to use a dynamic multiprogramming level to avoid the under-utilization of the machine introduced by these new scheduling policies when using a static multiprogramming level. We have evaluated this work by executing several workloads in an SGI Origin2000 with 64 processors. Results show that the combination of (target efficiency+dynamic multiprogramming level) achieves, in the worst case, the same performance as the equipartition and the equal efficiency, and in the best case it achieves a speedup of up to 1.3 in individual applications and in specific workloads a speedup of up to 2.5, with respect to the original algorithms. Julita Corbalán, Jesús Labarta |
IPDPS | 2 |
| 2001 | A Dynamic Periodicity Detector: Application to Speedup ComputationabstractWe propose a dynamic periodicity detector (DPD) for the estimation of periodicities in data series obtained from the execution of applications. We analyze the algorithm used by the periodicity detector and its performance on a number of data streams. It is shown how the periodicity detector is used for the segmentation and prediction of data streams. In an application case we describe how the periodicity detector is applied to the dynamic detection of iterations in parallel applications, where the detected segments are evaluated by a speedup computation tool. We test the performance of the periodicity detector on a number of parallelized benchmarks. The periodicity detector correctly identifies the iterations of parallel structures also in the case where the application has nested parallelism. In our implementation we measure only a negligible overhead produced by the periodicity detector. We find the DPD to be useful and suitable for the incorporation in dynamic optimization tools. Felix Freitag, Julita Corbalán, Jesús Labarta |
IPDPS | 3 |
| 2001 | Extending Heterogeneity to RAID Level 5
Toni Cortes, Jesús Labarta |
USENIX ATC, General Track | 2 |
| 2001 | A Framework for Integrating Data Alignment, Distribution, and Redistribution in Distributed Memory MultiprocessorsabstractParallel architectures with physically distributed memory provide a cost-effective scalability to solve many large scale scientific problems. However, these systems are very difficult to program and tune. In these systems, the choice of a good data mapping and parallelization strategy can dramatically improve the efficiency of the resulting program. In this paper, we present a framework for automatic data mapping in the context of distributed memory multiprocessor systems. The framework is based on a new approach that allows the alignment, distribution, and redistribution problems to be solved together using a single graph representation. The Communication Parallelism Graph (CPG) is the structure that holds symbolic information about the potential data movement and parallelism inherent to the whole program. The CPG is then particularized for a given problem size and target system and used to find a minimal cost path through the graph using a general purpose linear 0-1 integer programming solver. The data layout strategy generated is optimal according to our current cost and compilation models. Jordi Garcia 0001, Eduard Ayguadé, Jesús Labarta |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2000 | A Case for Heterogeneous Disk ArraysabstractHeterogeneous disk arrays are becoming a common configuration in many sites and especially in storage area networks (SAN). As new disks have different characteristics than old ones, adding new disks or replacing old ones ends up in a heterogeneous disk array. Current solutions to these kinds of arrays do not take advantage of the improved characteristics of the new disks. The authors present a block-distribution algorithm that takes advantage of these new characteristics and thus improves the performance and capacity of heterogeneous disk arrays compared to current solutions. Toni Cortes, Jesús Labarta |
CLUSTER | 2 |
| 2000 | Sparse Matrix Structure for Dynamic Parallelisation Efficiency
Markus Ast, Cristina Barrado, José María Cela, Jesús Labarta, Óscar Laborda, Hartmut Manz, Uwe Schulz |
Euro-Par | 5 |
| 2000 | User-Level Dynamic Page Migration for Multiprogrammed Shared-Memory MultiprocessorsabstractThis paper presents algorithms for improving the performance of parallel programs on multiprogrammed shared-memory NUMA multiprocessors, via the use of user-level dynamic page migration. The idea that drives the algorithms is that a page migration engine can perform accurate and timely page migrations in a multiprogrammed system if it can correlate page reference information with scheduling information obtained from the operating system. The necessary page migrations can be performed as a response to scheduling events that break the implicit association between threads and their memory affinity sets. We present two algorithms that use feedback from the kernel scheduler to aggressively migrate pages upon thread migrations. The first algorithm exploits the iterative nature of parallel programs, while the second targets generic codes without making assumptions on their structure. Performance evaluation on an SGI Origin2000 shows that our page migration algorithms provide substantial improvements in throughput of up to 264% compared to the native IRIX 6.5.5 page placement and migration schemes. Dimitrios S. Nikolopoulos, Theodore S. Papatheodorou, Constantine D. Polychronopoulos, Jesús Labarta, Eduard Ayguadé |
ICPP | 4 |
| 2000 | A case for use-level dynamic page migrationabstractThis paper presents user-level dynamic page migration, a runtime technique which transparently enables parallel pro-grams to tune their memory performance on distributed shared memory multiprocessors, with feedback obtained from dynamic monitoring of memory activity. Our technique exploits the iterative nature of parallel programs and information available to the program both at compile time and at runtime in order to improve the accuracy and the timeliness of page migrations, as well as amortize better the overhead, compared to page migration engines implemented in the operating system. We present an adaptive page migration algorithm based on a competitive and a predictive criterion. The competitive criterion is used to correct poor page placement decisions of the operating system, while the predictive criterion makes the algorithm respon-sive to scheduling events that necessitate immediate page migrations, such as preemptions and migrations of threads. We also present a new technique for preventing page ping-pong and a mechanism for monitoring the performance of page migration algorithms at runtime and tuning their sen-sitive parameters accordingly. Our experimental evidence on a SGI Origin2000 shows that unmodified OpenMP codes linked with our runtime system for dynamic page migration are effectively immune to the page placement strategy of the operating system and the associated problems with data locality. Furthermore, our runtime system achieves solid performance improvements compared to the IRIX 6.5.5 page migration engine, for single parallel OpenMP codes and multiprogrammed workloads. Dimitrios S. Nikolopoulos, Theodore S. Papatheodorou, Constantine D. Polychronopoulos, Jesús Labarta, Eduard Ayguadé |
ICS | 4 |
| 2000 | Applying Interposition Techniques for Performance Analysis of OpenMP Parallel ApplicationsabstractTuning parallel applications requires the use of effective tools for detecting performance bottlenecks. Along a parallel program execution, many individual situations of performance degradation may arise. We believe that an exhaustive and time-aware tracing at a fine-grain level is essential to capture this kind of situations. This paper presents a tracing mechanism based on dynamic code interposition, and compares it with the usual compiler-directed code injection. Dynamic code interposition adds monitoring code at run-time to unmodified binaries and shared libraries, making it suitable for environments in which the compiler or the available tools do not offer instrumentation facilities. Static injection and dynamic interposition techniques are used to collect detailed traces that feed an analysis tool. Both environments meet the accuracy and performance goals required to profile and analyze parallel applications and runtime libraries. Marc González 0001, Albert Serra, Xavier Martorell, José Oliver 0002, Eduard Ayguadé, Jesús Labarta, Nacho Navarro |
IPDPS | 6 |
| 2000 | A Tool to Schedule Parallel Applications on Multiprocessors: The NANOS CPU MANAGER
Xavier Martorell, Julita Corbalán, Dimitrios S. Nikolopoulos, Nacho Navarro, Eleftherios D. Polychronopoulos, Theodore S. Papatheodorou, Jesús Labarta |
JSSPP | 7 |
| 2000 | Performance-Driven Processor Allocation
Julita Corbalán, Xavier Martorell, Jesús Labarta |
OSDI | 3 |
| 2000 | Is Data Distribution Necessary in OpenMP?abstractThis paper investigates the performance implications of data placement in OpenMP programs running on modern ccNUMA multiprocessors. Data locality and minimization of the rate of remote memory accesses are critical for sustaining high performance on these systems. We show that due to the low remote-to-local memory access latency ratio of state-of-the-art ccNUMA architectures, reasonably balanced page placement schemes, such as round-robin or random distribution of pages incur modest performance losses. We also show that performance leaks stemming from suboptimal page placement schemes can be remedied with a smart user-level page migration engine. The main body of the paper describes how the OpenMP runtime environment can use page migration for implementing implicit data distribution and redistribution schemes without programmer intervention. Our experimental results support the effectiveness of these mechanisms and provide a proof of concept that there is no need to introduce data distribution directives in OpenMP and warrant the portability of the programming model. Dimitrios S. Nikolopoulos, Theodore S. Papatheodorou, Constantine D. Polychronopoulos, Jesús Labarta, Eduard Ayguadé |
SC | 4 |
| 2000 | NanosCompiler: supporting flexible multilevel parallelism exploitation in OpenMPabstractThis paper describes the support provided by the NanosCompiler to nested parallelism in OpenMP. The NanosCompiler is a source-to-source parallelizing compiler implemented around a hierarchical internal program representation that captures the parallelism expressed by the user (through OpenMP directives and extensions) and the parallelism automatically discovered by the compiler through a detailed analysis of data and control dependences. The compiler is finally responsible for encapsulating work into threads, establishing their execution precedences and selecting the mechanisms to execute them in parallel. The NanosCompiler enables the experimentation with different work allocation strategies for nested parallel constructs. Some OpenMP extensions are proposed to allow the specification of thread groups and precedence relations among them. Copyright © 2000 John Wiley & Sons, Ltd. Marc González 0001, Eduard Ayguadé, Xavier Martorell, Jesús Labarta, Nacho Navarro, José Oliver 0002 |
Concurr. Pract. Exp. | 4 |
| 2000 | Sensitivity of Performance Prediction of Message Passing Programs
Sergi Girona, Jesús Labarta |
J. Supercomput. | 2 |
| 1999 | Influence of Variable Time Operations in Static Instruction Scheduling
Patricia Borensztejn, Cristina Barrado, Jesús Labarta |
Euro-Par | 3 |
| 1999 | Exploiting Multiple Levels of Parallelism in OpenMP: A Case StudyabstractMost current shared-memory parallel programming environments are based on thread packages that allow the exploitation of a single level of parallelism. These thread packages do not enable the spawning of new parallelism from a previously activated parallel region. Current initiatives (like OpenMP) include in their definition the exploitation of multiple levels of parallelism through the nesting of parallel constructs. This paper analyzes the requirements towards an efficient multi-level parallelization and reports some conclusions gathered from the experience in the parallelization of two benchmark applications. The underlying system is based on: i) an OpenMP compiler which accepts some extensions to the original definition and ii) a user-level threads library that supports the exploitation of both fine-grain and multi-level parallelism. Eduard Ayguadé, Xavier Martorell, Jesús Labarta, Marc González 0001, Nacho Navarro |
ICPP | 3 |
| 1999 | Thread fork/join techniques for multi-level parallelism exploitation in NUMA multiprocessorsabstractThis paper presents some techniques for efficient thread forking and joining in parallel execution environments, taking into consideration the physical structure of NUMA machines and the support for multi-level parallelization and processor grouping.Two work generation schemes and one join mechanism are designed, implemented, evaluated and compared with the ones used in the IFUX MP library, an efficient implementation which supports a single level of parallelism.Supporting multiple levels of parallelism is a current research goal, both in shared and distributed memory machines.Our proposals include a first work generation scheme (GWD, or global work descriptor) which supports multiple levels of parallelism, but not processor grouping.The second work generation scheme (LWD, or local work descriptor) has been designed to support multiple levels of parallelism and processor grouping.Processor grouping is needed to distribute processors among different parts of the computation and maintain the working set of each processor across different parallel constructs.The mechanisms are evaluated using synthetic benchmarks, two SPEC95fp applications and one NAS application.The performance evaluation concludes that: i) the overhead of the proposed mechanisms is similar to the overhead of the existing ones when exploiting a single level of parallelism, and ii) a remarkable improvement in performance is obtained for applications that have multiple levels of parallelism.The comparison with the traditional single-level parallelism exploitation gives an improvement in the range of 3065% for these applications. Xavier Martorell, Eduard Ayguadé, Nacho Navarro, Julita Corbalán, Marc González 0001, Jesús Labarta |
International Conference on Supercomputing | 6 |
| 1998 | Kernel-level Scheduling for the Nano-threads Programming ModelabstractMultiprocessor systems are increasingly becoming the systems of choice for low and high-end servers, running such diverse tasks as number crunching, large-scale simulations, data base engines and world wide web server applications.With such diverse workloads, system utilization and throughpuf as well as execution time become important performance metrics.In this paper we present efficient kernel scheduling policies and propose a new kernel-user interface aiming at supporting efficient parallel execution in diverse workload environments.Our approach relies on support for user level threads which are used to exploit parallelism within applications, and a two-level scheduling policy which coordinates the number of resources allocated by the kernel with the number of threads generated by each application.We compare our scheduling policies with the native gang scheduling policy of the IRIX 6.4 operating system on a Silicon Graphics Ori-gin2000.Our experimental results show substantial performance gains in terms of overall workload execution times, individual application execution times, and cache performance. Eleftherios D. Polychronopoulos, Xavier Martorell, Dimitrios S. Nikolopoulos, Jesús Labarta, Theodore S. Papatheodorou, Nacho Navarro |
International Conference on Supercomputing | 4 |
| 1997 | Hamiltonian Recurrence for ILP
Cristina Barrado, Jesús Labarta |
Euro-Par | 2 |
| 1997 | Analyzing Scheduling Policies Using Dimemas
Jesús Labarta, Sergi Girona, Toni Cortes |
Parallel Comput. | 1 |
| 1995 | Automatic generation of loop scheduling for VLIW
Cristina Barrado, Jesús Labarta, Eduard Ayguadé, Mateo Valero |
PACT | 2 |
| 1995 | A Novel Approach Towards Automatic Data DistributionabstractData distribution is one of the key aspects that a parallelizing compiler for a distributed memory architecture should consider, in order to get efficiency from the system. The cost of accessing local and remote data can be one or several orders of magnitude different, and this can dramatically affect performance. In this paper, we present a novel approach to automatically perform static data distribution. All the constraints related to parallelism and data movement are contained in a single data structure, the Communication-Parallelism Graph (CPG). The problem is solved using a linear 0-1 integer programming model and solver. In this paper we present the solution for one-dimensional array distributions, although its extension to multi-dimensional array distributions is also outlined. The solution is static in the sense that the layout of the arrays does not change during the execution of the program. We also show the feasibility of using this approach to solve the problem in terms of compilation time and quality of the solutions generated. Jordi Garcia 0001, Eduard Ayguadé, Jesús Labarta |
SC | 3 |
| 1989 | GTS: parallelization and vectorization of tight recurrencesabstractIn this paper we present a new method for extracting the maximum parallelism or vector operations out of DO loops with tight recurrences using sequential programming languages. We have named the method Graph Traverse Scheduling (GTS). It is devised to produce code for shared memory multiprocessors or vector machines. When parallelizing, hardware support for fast synchronization is assumed. Eduard Ayguadé, Jesús Labarta, Jordi Torres, Patricia Borensztejn |
SC | 2 |
| 1987 | Optimized Mesh-Connected Networks for SIMD and MIMD ArchitecturesabstractA class of mesh networks with wrap-around links is obtained from a class of circulant graphs by means of a graph isomorphism. We demonstrate how to obtain, from the adjacency pattern of the graph, simple parameters that serve to construct a planar design of the network. Several performance parameters are evaluated: in particular, we show that diameter and average distance are simultaneously minimized. This implies a minimization of the network communication delays. Due to its easy implementation and good behavior characteristics, the proposed interconnection scheme is appropriate in several architectural environments. Specifically, this topology is suitable as an interconnection subsystem for message passing MIMD architectures, as well as for SIMD machines with a static interconnection scheme. In the particular case of SIMD machines, a comparison is made with the ILLIAC IV-type networks. As a consequence, we propose still another topology, when the number of processing elements is an even power of 2; for this solution, we show that a reduction in the network distances is achieved, without losing speed in performing arbitrary permutations. Ramón Beivide, Enrique Herrada, José L. Balcázar, Jesús Labarta |
ISCA | 4 |
| 1985 | Analysis and Simulation of Multiplexed Single-Bus Networks With and Without BufferingabstractPerformance issues of a single-bus interconnection network for multiprocessor systems, operating in a multiplexed way, are presented in this paper. Several models are developed and used to allow system performance evaluation. Comparisons with equivalent crossbar systems are provided. It is shown how crossbar EBW values can be reached and exceeded when appropriate operation parameters are chosen in a multiplexed single-bus system. Another architectural feature is considered, concerning the utilization of buffers at the memory modules. With the buffering scheme, memory interference can be reduced so that the system performance is practically improved. José María Llabería, Mateo Valero, Enrique Herrada Lillo, Jesús Labarta |
ISCA | 4 |
| 1983 | A performance evaluation of the multiple bus network for multiprocessor systemsabstractIn this paper we present a mathematical model to compute the bandwidth of the multiple bus interconnection network. Due to the computational complexity associated with the exact solution, the processors are removed from the queues at the end of each memory cycle to facilitate the analysis. This leads to approximate solutions which are both easier to obtain and very accurate. Mateo Valero, José María Llabería, Jesús Labarta, Emilio Sanvicente, Tomás Lang |
SIGMETRICS | 3 |